The R1 model wasn't free, and the DeepSeek R1 paper doesn't talk about costs at all. The $5.567M cost for DeepSeek is for when you have all the ingredients in front of you – running that recipe costs $5.567M. It doesn't consider the cost of procuring those ingredients (in this
LLMS
-
DeepSeek Disruption Favors Application Layer Over Foundation Models
By
–
Today's "DeepSeek selloff" in the stock market — attributed to DeepSeek V3/R1 disrupting the tech ecosystem — is another sign that the application layer is a great place to be. The foundation model layer being hyper-competitive is great for people building applications.
-
llama.cpp enables easy local model execution
By
–
llama.cpp is the best! It’s wild that one can run a model of this size at the click of a button
-
LLM Merry-Go-Round: Predictable Cycle of Model Obsolescence
By
–
By the time you're back there will be something that unseats it (and comes in rose gold). This LLM merry-go-round is dizzying and sort of predictable.
-
DeepSeek Janus-Pro and R1 Model Integration Unveiled
By
–
Lots of talk about @deepseek_ai since last week. What's everyone think? This morning, DeepSeek debuted a new Janus-Pro image creation model. Meanwhile, @perplexity_ai just announced DeepSeek's R1 model is now available for deep web research.
-
DeepSeek challenges American AI dominance with superior reasoning
By
–
On est d’accord Frédéric mais ChatGPT n’est pas exempt de censure et d’erreur ! La menace est la pour l’hégémonie américaine et les marchés n’ont pas tardé à réagir. L’adoption chez les Américains est fulgurante et les performances sont la en terme de raisonnement. DeepSeek
-
Pika Labs lance Pika 2.1 pour la génération vidéo
By
–
BREAKING 🚨: Pika Labs released Pika 2.1, a new video generation model 🔥
— 🚨 AI News | TestingCatalog (@testingcatalog) 27 janvier 2025
"Our latest model with better motion and ingredients support" https://t.co/uWslj6VUzG pic.twitter.com/VZF76BgnpYBREAKING : Pika Labs released Pika 2.1, a new video generation model "Our latest model with better motion and ingredients support"
-
Lower-precision formats and distilled models enabling local long-context LLMs
By
–
Me neither :). I think with new lower-precision formats, and better distilled models, we'll maybe increasingly adopt long-context LLMs run locally.
-
Long-context needle-in-the-haystack improvements driven by cost optimization
By
–
I think that's exactly what's happening but for cost reasons rather than accuracy reasons. I think long-context needle-in-the-haystack issues have improved a lot last year since all major LLMs now have a dedicated long-context finetuning stage in pre/post-training.
-
Cheap Tokens: The Abundant Future of AI Computing Power
By
–
Worrying about cheap tokens is like wondering what you’d do with a 500gb hard drive in 1999. Trust me, we’ll figure it out.