DeepSeek just released Tile Kernels! Tile Kernels are optimized GPU kernels for LLM operations, built with TileLang. DeepSeek claims that “most kernels in this project approach the limits of hardware performance in terms of compute intensity and memory bandwidth. Some of them
OPEN SOURCE
-
Google Open-Sources DESIGN.md Countering Anthropic Claude Limits
By
–
Google’s level of disrespect is OFF THE CHARTS right now.
— Charly Wargnier (@DataChaz) 23 avril 2026
Anthropic really thought they had us locked down with Claude Design’s ridiculous rate limits…
…and now Google has literally countered it straight away by open-sourcing DESIGN.md 🤯pic.twitter.com/9Hq4nLX2GK https://t.co/7KfGst99uxGoogle’s level of disrespect is OFF THE CHARTS right now. Anthropic really thought they had us locked down with Claude Design’s ridiculous rate limits… …and now Google has literally countered it straight away by open-sourcing DESIGN.md
-

Open-source community advances AI memory solutions
By
–
AI memory is one of the most important unsolved problems in the space. Glad to see the open-source community stepping up where the incumbents are building walls. 🙂
-
Security Fix: Dependencies Reorganized to Reduce Supply Chain Risk
By
–
I posted a workaround in Discord, or you remove/reinstall or wait for next release. Moved dependencies around to reduce supply chain risk. Security is hard.
-

Moonshot Releases Kimi K2.6 Open-Weight Model
By
–
Kimi K2.6, the new state-of-the-art open-weight model from Moonshot, is now available for Pro and Max subscribers.
-

Xiaomi unveils new MiMo-V2.5 open-source models
By
–

Xiaomi released new MiMo-V2.5-Pro and MiMo-V2.5 open-source models. > MiMo-V2.5-Pro, the strongest model yet. > MiMo-V2.5, native omnimodal with strong agentic capabilities.
-
Why Open-Weight Escape Hatches Matter for Embedding Models
By
–
This is why I won't use proprietary hosted embedding models myself – I am more than happy to pay for a hosted solution (cheaper, faster and more convenient than self-hosting) but I want an open weight escape hatch for if they ever stop serving it
-

NVIDIA NeMo RL Accelerates Agentic Performance with FP8
By
–
Improve agentic performance with accurate RL post-training on low-precision FP8. NVIDIA NeMo RL, an open-source library within NVIDIA NeMo, supports FP8 to speed up RL workloads by 1.48x on Qwen3-8B-Base—enabling faster iterations for agentic tool use and multi-step
-

AWQ Quantization Optimization for Agentic AI Workloads
By
–
For real agentic workloads (North), short-context calibration wasn't enough. We calibrated AWQ on long internal agentic traces (up to 64k tokens) and added token masking in llm-compressor to exclude repetitive chat templates/tool descriptions from calibration stats. Plus QAD
-
First Notebook Competition with marimo to Implement AI Papers
By
–
Looking for a weekend project that is both challenging and rewarding? Enter our first notebook competition with @marimo_io to implement AI papers. You also got 5 days left if you want to win a Mac Mini + over $1K in prizes! Full details found below
