The real question: What is the DeepSeek-R1 training cost? The $5.567M DeepSeek cost is missing the cost of training the R1 model to get distilled data. Table 9 in the DeepSeek v3 paper shows that the R1 distillation step is critical for quality. The R1 paper doesn't talk about
LLMS
-
Infinite demand for AI tokens: o1 model usage scaling
By
–
If I could run o1 1,000 times on each question I ask I would do it in a heartbeat. There’s infinite demand for tokens.
-
Pre-training remains superior to expensive knowledge injection methods
By
–
Interesting! But knowledge is still best instilled via pre-training. "Relatively" cheap still remains very expensive and cumbersome, even for 7B models, compared to just adding some more embeddings to a database.
-

2025: The Year of LLM Specialization and Multimodal Models
By
–
Remember when we were arguing about the term "foundation models"? It feels like ages ago! With those base models culminating in DeepSeek v3 in Dec 2024, 2025 will likely be the year of LLM specialization! 1. Multimodal LLMs: It was my big prediction for 2024. Most products
-
DeepSeek’s Cost-Efficient Training Methods Explained
By
–
One explanation of how DeepSeek was able to train its model so cheaply..
-
Document Databases and AI Workloads: Compatibility Myths Debunked
By
–
"Equivalent" . It might be good (and I hope it *is* good in its own right – JSON-native, document-style databases are critical for AI workloads, in particular), but let's not play make believe on compatibility. (Note F's 2.0 blog where they ack perf/etc has been relative)
-

DeepSeek R1 now available on Perplexity Pro Search
By
–
DeepSeek R1 is now available on Perplexity to support deep web research. There's a new Pro Search reasoning mode selector, along with OpenAI o1, with transparent chain of thought into model's reasoning. We're increasing the number of daily uses for both free and paid as add more
-

DeepSeek’s Rise Won’t End GPU Demand, Wall Street Says
By
–
Wall Street’s Best Ideas: No, DeepSeek does not mean the end of buying tons of GPUs The Chinese AI innovation DeepSeek gives the impression no on will buy Nvidia GPUs anymore. But the program scores lower on some tests, suggesting the race for progress will continue.
-
DeepSeek’s 10X Efficiency Breakthrough Reshapes AI Resource Demands
By
–
Take your intelligence estimates for Project Stargate and multiply them by 10. Every AI lab now has the means to make their models 10X+ more efficient because of DeekSeek. This doesn’t mean AI and chips will be used 10X less: the 10X cost decrease means it gets used 10X more.
-
Perplexity adds US-hosted DeepSeek R1 reasoning
By
–
BREAKING 🚨: Perplexity released US-hosted DeepSeek R1 reasoning option 👀 https://t.co/Wm5KT4bmpO pic.twitter.com/yKnzCQCt7n
— 🚨 AI News | TestingCatalog (@testingcatalog) 27 janvier 2025BREAKING : Perplexity released US-hosted DeepSeek R1 reasoning option
