LLM hallucinations sound scary. But what if they could actually help you build more resilient AI solutions? Join us for a live session on April 10 at #GoogleCloudNext to learn how to take advantage of hallucinations and build your AI stack with the agility to proactively
LLMS
-
Stanford Launches Center for Research on Foundation Models
By
–
August 2021: Recognizing a paradigm shift in AI, we launched the Center for Research on Foundation Models (
@StanfordCRFM
) led by @percyliang and published a groundbreaking report on the opportunities and risks of foundation models. #MondayMilestones 11/n -

Scaling Mamba: Addressing Training Instability with Layer Adjustments
By
–
We found that scaling Mamba to substantial sizes isn’t trivial and causes unstable training. We had to make some adjustments to the Mamba layers to stabilize them, as can be seen in the graph below (training loss before & after the change) 6/6
-

Jamba combines Transformer, Mamba, MoE for efficient scaling
By
–
Combining Transformer, Mamba & MoE allows flexibility in balancing low memory usage, high throughput, and high quality. Jamba’s KV cache – which becomes a limiting factor when scaling context in pure Transformers – is 8x smaller compared to a pure Transformer. 5/6
-

Jamba’s Impressive Long-Context Performance with Minimal Attention Layers
By
–
Jamba was trained to handle contexts of up to 256K. Jamba has excellent performance in the needle-in-a-haystack evaluation, which is especially interesting given its use of only 4 attention layers. It also outperforms Mixtral on most long-context benchmarks. 4/6
-

Mamba Models Struggle with In-Context Learning Compared to Attention
By
–
We noticed that pure Mamba models struggle to develop in-context learning capabilities. E.g., they performed substantially worse than the pure attention model in 3 common benchmarks while the attention–Mamba exhibits similar results to just Transformers. 3/6
-

Jamba Hybrid Model Outperforms Pure Attention and Mamba Architectures
By
–
We see that the hybrid Jamba model outperforms both pure attention and pure Mamba models. The ratio of attention-to-Mamba layers of either 1:3 or 1:7 performs comparably, but given that a 1:7 ratio is more compute-efficient, we opt for it in our model. 2/6
-

Jamba Whitepaper Released: Hybrid SSM-Transformer Architecture Details
By
–
Jamba whitepaper is out!
The whitepaper details our in-depth ablations on this novel hybrid SSM-Transformer architecture, and how we chose to interleave Mamba, Transformer and MoE. https://
arxiv.org/abs/2403.19887 Here are some highlights from the paper 1/6 -

Manage LLM Deployments with Predibase New Deployments Page
By
–
View and manage your dedicated and #serverless #LLM deployments in one place with our brand new Deployments page. Quickly spin up new #GPUs or simply use our serverless #finetuned endpoints for per-token pricing. https://t.co/AZautnDFGu pic.twitter.com/rKmWNzPhJu
— Predibase by Rubrik (@predibase) 1 avril 2024View and manage your dedicated and #serverless #LLM deployments in one place with our brand new Deployments page. Quickly spin up new #GPUs or simply use our serverless #finetuned endpoints for per-token pricing. https://
pbase.ai/3TGTijv -
Python, Machine Learning, MLOps, CV/NLP, LLMs tutorials daily
By
–
If you are interested in: – Python – Machine Learning – MLOps – CV/NLP – LLMs Find me → @Sumanth_077 Everyday, I share tutorials on above topics! Like/RT the first tweet to help this reach more people!
