1/5 Releasing Jamba Reasoning 3B under Apache 2.0: Hybrid SSM-Transformer architecture that tops accuracy & speed across record context lengths. e.g. 3-5X faster than Llama 3.2 3B and Qwen3 4B at 32K tokens.
LLMS
-
GPT-Realtime and E2E Models in Production Projects
By
–
little survey: are you using gpt-realtime/ mini or any other E2E model in projects? if yes, what for?
-

Sam Altman, Gemini 2.5, and AI tools roundup
By
–
Top stories in AI today: – Sam Altman on Dev Day, AGI, and more
– Google releases Gemini 2.5 Computer Use
– Create LinkedIn carousels in ChatGPT with Canva
– Duke’s AI for smarter drug delivery
– 4 new AI tools, community workflows, and more Read more: https://
therundown.ai/p/exclusive-in
terview-sam-altman-on-dev-day-and-ais-future
… -
Ollama Alternatives: LMStudio, Llama.cpp, vLLM and More
By
–
ollama alternatives > lmstudio
> llama.cpp
> exllamav2/v3
> vllm
> sglang among many others like literally anything is better than ollama lmao -

QuestA: Expanding LLM Reasoning via Question Augmentation
By
–
#PapersAccepted by Jiqizhixin
Our report: https://
mp.weixin.qq.com/s/Mhy9fWM8KVnu
3mTy7I3osQ
… QuestA: Expanding Reasoning Capacity in LLMs via Question Augmentation Tsinghua University, Shanghai Qi Zhi Institute, and others
Paper: https://
arxiv.org/abs/2507.13266
Models: https://
huggingface.co/foreverlasting
1202/QuestA-Nemotron-1.5B
…
Code: -

QuestA: Reinforcement Learning Improves Language Model Reasoning
By
–
Can reinforcement learning really make language models better reasoners? This study says yes — with a twist. Introducing QuestA, a Question Augmentation strategy that feeds models partial solutions during RL training to ease difficulty and deliver richer feedback. Applied to
-

Nvidia Fast-dLLM v2: Efficient Block-Diffusion LLM
By
–
Nvidia presents Fast-dLLM v2
— AK (@_akhaliq) 8 octobre 2025
Efficient Block-Diffusion LLM pic.twitter.com/wjqpvi1LoANvidia presents Fast-dLLM v2 Efficient Block-Diffusion LLM
-
Build Your Own Byte-Pair Encoder: LLM Engineering Fundamentals
By
–
step-by-step LLM Engineering Projects each project = one concept learned the hard (i.e. real) way Tokenization & Embeddings > build byte-pair encoder + train your own subword vocab
> write a “token visualizer” to map words/chunks to IDs
> one-hot vs learned-embedding: plot -

Leading AI Models Likely Highly Profitable at Scale
By
–
By the way, this is equivalent to $120m in a month (at $1.25/$10 per million tokens & assuming 80% input / 20% output). I know people don't literally pay that, but leading models are likely quite profitable

