A pipeline for creating VQA data using knowledge bases (wikidata). We demonstrated the method on world/cultural knowledge and could be extended to other domains. (there are rephrasing/filtration steps that require LLMs/VLMs, but the released data does not).
LLMS
-

Sonnet 5 tests ongoing, excellent release ahead
By
–
It seems that the first tests with Sonnet 5 are already in progress. If this is confirmed, we are in for an excellent release!
-

LLM Engineering Projects Roadmap 2026 Edition
By
–
The Ultimate Step-By-Step LLM Engineering Projects Roadmap (2026 Edition) – Build a tokenizer
– Learn embeddings
– Implement RoPE / ALiBi
– Hand-wire attention
– Build MHA
– Build a Transformer block
– Train a mini-former
– Compare objectives
– Build sampling
– Speculative -

First time seeing a warning screen in ChatGPT
By
–
It is the first time I am seeing this warning screen in ChatGPT
-

LangSmith LLM Gateway: Anonymizes Data and Controls Costs
By
–
LangSmith LLM Gateway sits between your agents and LLM providers. It applies spending limits and anonymizes personal data (PII) before requests reach the model, stopping problems at the source rather than simply
-
The release of GPT-5.6 paves the way for GPT-5.7
By
–
The sooner GPT-5.6 is released, the sooner we can start talking about GPT-5.7.
-

VibeThinker-3B compresses top-tier reasoning into small language models
By
–
/2 VibeThinker-3B is a 3-billion-parameter model proving that top-tier verifiable reasoning can be efficiently compressed into small language models. It uses a "Spectrum-to-Signal" pipeline with supervised fine-tuning, reinforcement learning, and self-distillation. By optimizing
-
Sending the prompt is better than an LLM-written email
By
–
Read this somewhere and thought this was spot on: "If you’re going to use an LLM to write me an email, I’d much rather you just send me the prompt; at least then I’d have an idea of what you actually meant to say.”
-
Large model not for home, but enables price competition among suppliers
By
–
Due to its size, the model is not one we celebrate because we can run it comfortably at home, but rather because many suppliers will be able to compete to offer it at a more competitive price than if it were a model from a single supplier.
