I've been playing with @SambaNovaAI
's API serving fast Llama 3.1 405B tokens. Really cool to see leading model running at speed. Congrats to Samba Nova for hitting a 114 tokens/sec speed record (and also thanks @KunleOlukotun for getting me an API key!)
LLMS
-
SambaNova Achieves 114 Tokens Per Second Record with Llama
By
–
-

Research: Humans Outsmart Automated LLM Defenses
By
–
Humans, noted virtuosi of adversarial yap, remain #1 at trolling LLMs! New research from @scale_AI
's SEAL team shows human red teamers achieve 70%+ success rates against LLM defenses that stump automated attacks, exploiting their susceptibility to multi-turn jailbreaks. -

Coding Improvement of Gemini 1.5 Pro
By
–
Gemini 1.5 Pro got improved in coding in complex prompts and this is how it compares with a previous variant
-
AI Regulation Replacing Jury Trials with LLM Judges
By
–
…and also an AI regulation that removes jury trials and replaces them with LLMs as a judge — that's good too?
-
AI Regulation and LLM-Based Education: Ethical Concerns
By
–
Cool. So you must be in favor then of an AI regulation that requires all children be taught entirely using LLM-created content at schools, since all AI regulations are good?
-
Cerebras Inference Achieves Record Throughput for Llama 3.1 Models
By
–
Verified by @ArtificialAnlys, @CerebrasSystems Inference is capable of serving Llama 3.1 70B at 450 tokens/sec and Llama 3.1 8B at 1,850 tokens/sec! https://t.co/hCb9MmSvOo
— AI at Meta (@AIatMeta) 27 août 2024Verified by @ArtificialAnlys
, @cerebras Inference is capable of serving Llama 3.1 70B at 450 tokens/sec and Llama 3.1 8B at 1,850 tokens/sec! -

Jamba 1.5: Balancing Speed and Quality Efficiency
By
–
But our cost, efficiency and speed don't come at the expense of quality. In the following chart, Jamba 1.5 Large and Mini both show a great balance between speed and quality (QI is where you want to be). >>
-

Jamba Architecture Optimizes Speed-Cost-Quality Triangle
By
–
High throughput itself is never enough; it's all about optimizing the speed-cost-quality triangle. Thanks to Jamba’s architecture, we offer high speed at a very competitive price (in the image below, QII is where you want to be). >>