DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads https://
arxiv.org/abs/2410.10819 https://
github.com/mit-han-lab/du
o-attention
… #MIT @songhan_mit
LLMS
-
DuoAttention: Efficient Long-Context LLM Inference with Retrieval
By
–
-
MIT KV Cache Framework Reduces LLM Decoding Latency
By
–
MIT HAN Lab introduce a framework that only applies a full KV cache to retrieval heads while using a light-weight, constant-length KV cache for streaming heads, which reduces both LLM's decoding and pre-filling memory and latency without compromising its long-context abilities.
-

Anthropic Launches Computer Use API for Claude
By
–
Major Innovations from Anthropic: Unveiling the Future of AI! Anthropic has been hard at work, and they have some game-changing updates to share: Computer Use API They’ve developed an advanced API that empowers Claude to perceive and interact with computer interfaces.
-

Google Open-Sources SynthID-Text for AI-Generated Content Detection
By
–
We created SynthID, a robust digital watermarking technology to tag & identify AI-generated content. Now we’re open-sourcing SynthID-Text so developers can use it to embed & detect watermarks in text outputs from their own LLMs. Published today in @Nature https://
nature.com/articles/s4158
6-024-08025-4
… -

Neural and Non-Neural AI: Reasoning, Transformers, and LSTMs
By
–
Neural and Non-Neural AI, Reasoning, Transformers, and LSTMs https://
bit.ly/3XPO9aB
#AI #MachineLearning #DeepLearning #LLMs #DataScience -

10000 SLMs Fine-tuned on Predibase Platform
By
–
Over 10,000 #SLMs have been fine-tuned on Predibase! 🎉
— Predibase by Rubrik (@predibase) 23 octobre 2024
Why do leading AI companies like #Checkr, #Nubank and #Upstage fine-tune and #serve models on Predibase?
🎯 Better Accuracy: Fine-tuned models on Predibase beat hashtag#GPT4 by 5-20% (see our leaderboard:… pic.twitter.com/KR6blwhg9QOver 10,000 #SLMs have been fine-tuned on Predibase! Why do leading AI companies like #Checkr, #Nubank and #Upstage fine-tune and #serve models on Predibase? Better Accuracy: Fine-tuned models on Predibase beat hashtag#GPT4 by 5-20% (see our leaderboard:
-
Memento Wins CalHacks Prize Powered by Groq
By
–
(12/15) Memento – https://devpost.com/software/memento-1p0jel … – @CalHacks prize winner powered by Groq
-

Extending Fineweb filtering to 1000+ languages for pretraining
By
–
We want to extend up to 1000+ languages the data-driven filtering approach we used to create the *Fineweb* and *Fineweb-edu* large scale pretraining datasets The first step –which proved surprisingly difficult– was to find reliable high-early-signal evaluations in many languages
-

SambaNova Cloud Achieves Fast Llama 3.2 Inference Performance
By
–
Hit the ground running with fast #AI inference on @AIatMeta
's Llama 3.2! SambaNova Cloud allows #devs to achieve 2470 tokens per sec on 1B and 1566 tokens per sec on 3B Start developing now -
LangSmith Adds Side-by-Side Trace Comparison with Prompt Highlighting
By
–
no more opening up two LangSmith windows side by side!
— Harrison Chase (@hwchase17) 23 octobre 2024
easily compare traces to see where the differ… including some nice highlighting of prompt differences https://t.co/xSbh7J4k4Wno more opening up two LangSmith windows side by side! easily compare traces to see where the differ… including some nice highlighting of prompt differences