"The biggest problem: developers don’t want GPUs. They want LLMs…. developers thinking about performance in terms of “tokens per second” aren’t counting milliseconds." — @mrkurt with a very clearheaded analysis of the gpu neocloud market. api calls are enough for most.
LLMS
-
Voice AI Future: Real-time Research Capabilities and Information Looping
By
–
Looking forward to a future where your realtime voice AI does deep research in the background and loops back information to you.
-
System Prompt Generator Tool for AI Agents and Chatbots
By
–
Free AI Agents System Prompt Generator? Generate powerful system prompts Using Chain-of-Thought For building AI Agents, Custom GPTs, & Chatbots! https://
godofprompt.ai/system-prompt-
generator
… -
DeepSeek-R1 671B Achieves 198 Tokens Per Second on SambaNova
By
–
⚡️ We gave our friends @AiBlckbx early access to try how fast DeepSeek-R1 671B runs on SambaNova Cloud.
— SambaNova (@SambaNovaAI) 14 février 2025
🎥 Huge speeds we’re seeing (198 t/s ⚡️⚡️⚡️) with DeepSeek-R1 671B on our cloud — we finished the whole output before GPUs could even finish reasoning.#AIWe gave our friends @blackboxai early access to try how fast DeepSeek-R1 671B runs on SambaNova Cloud. Huge speeds we’re seeing (198 t/s ) with DeepSeek-R1 671B on our cloud — we finished the whole output before GPUs could even finish reasoning. #AI
-

SambaNova Launches DeepSeek-R1 671B on Cloud Platform
By
–
Happy Valentine's Day from the SambaNova team! If you want to dive deep and analyze your relationship this Valentine’s Day, check out the fastest DeepSeek-R1 671B running on SambaNova Cloud. #DeepSeekR1 #HappyValentinesDay #AI
-

Grok-3 Praised as Scary Good with Superior Reasoning Abilities
By
–
And he says, Grok-3 is ‘scary good,’ and better at reasoning than all other models.
-

Chai Research Beats Character AI With Strong 2025 Roadmap
By
–
I feel like people slept on our @chai_research ep they straight up beat @character_ai at their own game*. like deepseek, former hedge fund guys going into the consumer LLM game mostly bootstrapped and crushing very hard. they just released their 2025 roadmap and… wow 25% x.com/latentspacepod…
-

Perplexity Deep Research Achieves 21.1% on Humanity’s Last Exam
By
–
Deep Research on Perplexity scores 21.1% on Humanity’s Last Exam, outperforming Gemini Thinking, o3-mini, o1, DeepSeek-R1, and other top models. We also have optimized Deep Research for speed.
-

Perplexity Deep Research Achieves 93.9% Accuracy on SimpleQA
By
–
Perplexity Deep Research surpasses leading models in performance—scoring 93.9% accuracy on the SimpleQA benchmark.
-
GPU Precision Evolution: From 32-bit to FP8 in Transformers
By
–
Yes. Originally GPUs were designed for 32 bit precision, but in transformer contexts, you get away with even lower precision (16 bit is the standard usually). In recent months, people pushed that even further to FP8. One of the most recent architectures that successfully did FP8
