AIs have a bad reputation for truth, so three important findings in this paper:
1) "LLM agents can achieve superhuman rating performance" on fact checking when given access to Google!
2) Bigger models are more factual
3) LLMs are 20x cheaper than humans https://
arxiv.org/pdf/2403.18802
.pdf
…
LLMS
-
LLM Agents Achieve Superhuman Fact-Checking Performance With Google Access
By
–
-

Hybrid Mamba-Transformer Model Addresses Architecture Limitations
By
–
This one looks like a big deal. It's not a pure transformer model – it's a combination of both Mamba and transformer, which appears to address different limitations present in each architecture
-
Databricks LLM Training Cost: Enterprise Model ROI Analysis
By
–
The Databricks LLM apparently cost $10M to train, which is a price I have seen given for a GPT-3.5 class model multiple times. Given that pricing, training your own corporate model versus open source needs to deliver a lot of benefit. I haven't seen evidence for that value, yet.
-

DBRX Becomes Number One Trending Model on Hugging Face
By
–
Not a surprise but DBRX is already #1 trending on HF!
-
Elastic integrates Cohere Embed v3 for RAG systems
By
–
We're excited to announce that @elastic now supports Cohere's Embed v3 model series in their Inference API, enabling developers to easily build accurate and scalable data retrieval pipelines for their RAG systems:
-

TheProfessor: Advanced LLM Merging with Conversational and Scientific AI
By
–
TheProfessor is a merge of #LLMs, created with help from @AbacusAI
's #GenerativeAI team. It can be used for many things with its conversational #AI, reasoning, scientific, medical, & mathematical skills interactive brainstorming: https://
huggingface.co/abacusai/ThePr
ofessor-155b
… -

Kirk Borne Updates Third Edition of Transformers NLP and Computer Vision Book
By
–
WOW! I have updated 3rd Edition of this: http://
amzn.to/3TkBknI
"Transformers for Natural Language Processing and Computer Vision: Explore #GenerativeAI and Large Language Models #LLMs with Hugging Face, ChatGPT, GPT-4V, DALL-E 3" Huge! 728 pages of #AI #DeepLearning -
AI21 Labs Introduces Jamba: New AI Architecture Model
By
–
Meet Jamba http://
ai21.com/jamba Build with Jamba https://
huggingface.co/ai21labs/Jamba
-v0.1
… Read more http://
ai21.com/blog/announcin
g-jamba
… #Jamba #Mamba #Transformer #AIarchitecture #SSM -

Jamba: Revolutionary SSM-Transformer Open Model with 3X Throughput
By
–
Introducing Jamba, our groundbreaking SSM-Transformer open model! As the first production-grade model based on Mamba architecture, Jamba achieves an unprecedented 3X throughput and fits 140K context on a single GPU. Meet Jamba http://
ai21.com/jamba Build on @huggingface -

Train Mixtral 8x7b Models on DGX Without PEFT
By
–
Now you can train models like Mixtral 8x7b without PEFT on 8xH100s with brand new "Train on DGX" feature on Hugging Face