At @hf0 last week I saw a startup that put an LLM into a tracking pixel. Every site that sells stuff has seen an increase in sales by using it. If I were an investor I would have written a big check.
LLMS
-
Sora 2 generations slow, heavy testing on 5.2 suggests imminent release
By
–
Sora 2 generations are taking twice as long as normal today OpenAI is doing heavy testing on 5.2 Release is imminent
-
Engineering Emergent Personalities in LLM Reward Optimization
By
–
There is definitely work going into engineering the "you" simulation – the personality that gets all the rewards in verifiable problems, or all the upvotes from users/judge LLMs, or mimics the responses of SFT, and there is an emergent composite personality from that. My point is
-
LLMs as Simulators: Rethinking Prompt Engineering Approaches
By
–
Don't think of LLMs as entities but as simulators. For example, when exploring a topic, don't ask: "What do you think about xyz"? There is no "you". Next time try: "What would be a good group of people to explore xyz? What would they say?" The LLM can channel/simulate many
-

Collaboration, Not AI, Is the Real Bottleneck
By
–
here's what broke my brain about this research. we've been optimizing the wrong side of the equation this entire time. better prompts. stronger models. higher benchmarks. longer context windows. more parameters. but the bottleneck isn't the AI. it's our ability to collaborate
-
Context Length Generalisation and Bandit Training in Language Models
By
–
Some really interesting comments below, but the question is still open and requires investigation. I hope a few students pick it up. I liked the discussions on context length generalisation, the fact that we typically train these models as bandits (even when we do RL, which is
-
Testing LLM Attention Mechanisms with Experimental Validation
By
–
Sure. But it attention truly worked, the LLM would no pay attention to the previous topic. Right? We should test this with experiments.
-

Benchmarking agentic AI APIs and token usage issues
By
–
Hello @MistralAI , Je fais un benchmark agentic en ce moment, dont Large 3, et votre API est la seule (parmi toutes les autres) qui oblige à avoir un message assistant en réponse obligatoirement après un tool call. Mais souvent c’est inutile et ça consomme des tokens en plus
-
10 Powerful Claude Prompts for a Million-Dollar Business
By
–
Claude Sonnet 4.5 is the closest thing to an economic cheat code we’ve ever touched but only if you ask it the prompts that make it uncomfortable. Here are 10 Powerful Claude prompts that will help you build a million dollar business (steal them):