Thanks! That's the LFM2-8B-A1B available on Hugging Face with vLLM backend
CODE
-
Fast Execution with Dependency-Based Environment Caching
By
–
It creates a dedicated environment somewhere which I think it reuses for future executions, provided none of the dependencies have changed Whatever it's doing it is FAST after the first run
-
UV Recipe for Python Testing with PyProject
By
–
I figured out a uv recipe for running tests for any project with pyproject.toml or setuppy using any Python version: uv run –python 3.14 –isolated –with-editable '.[test]' pytest I've wrapped it in a uv-test script: uv-test -p 3.11 Full details:
-
MIT Professor Explores How AI Is Transforming Code Writing
By
–
MIT professor on how AI is reshaping what it means to "write code."
— MIT CSAIL (@MIT_CSAIL) 8 octobre 2025
Part 1 of the full video: https://t.co/MH0brjWKno pic.twitter.com/2TuSPcggxKMIT professor on how AI is reshaping what it means to "write code." Part 1 of the full video: https://
shorturl.at/DdmSF -

Sort Models by Creation Date in Replicate HTTP API
By
–
You can now sort models by creation date in our HTTP API. This makes it easier to find the hottest new models programmatically. https://
replicate.com/changelog/2025
-10-08-models-api-sorting
… -

Ant Ling releases 1T-params open-source coding model
By
–
Ant Ling introduced a new 1T-params, non thinking open source model with a good performance on coding tasks. 1T
-
Ollama’s bloated wrapper fails to match ggml’s efficiency
By
–
do not use Ollama ggerganov wrote blazing-fast
C++ inference (ggml, llama.cpp) then Ollama wrapped it
in a bloated binary and is now somehow the face of local LLMs
soaking up VC hype and it's not even a good wrapper lol -

Run Claude Code Locally on Your Own GPU Setup
By
–
tired of Anthropic’s weekly limits and nerfed models? with one command and a few GPUs,
you can route Claude Code to your own local LLM Buy a GPU p.s. full video tutorial pinned at the top of my profile -

Jamba Reasoning 3B: Hybrid SSM-Transformer Apache 2.0 Release
By
–
1/5 Releasing Jamba Reasoning 3B under Apache 2.0: Hybrid SSM-Transformer architecture that tops accuracy & speed across record context lengths. e.g. 3-5X faster than Llama 3.2 3B and Qwen3 4B at 32K tokens.
-
GPT-Realtime and E2E Models in Production Projects
By
–
little survey: are you using gpt-realtime/ mini or any other E2E model in projects? if yes, what for?