No issues here. Nemotron is still my main LLM for most simple tasks and it taps Claude and GPT API models for anything more complex. Speed isn’t really an issue. Usually get a response within 20-seconds or so. Not the fastest in the world but not slow enough to be an issue.
GENERATIVE AI
-
Engramme’s Memory API Launched in Beta
By
–
Engramme's memory API is now live in public beta and built on an entirely new AI architecture, not Transformers, purpose-built to give apps persistent human memory without any search or prompting from the user.
— 🚨 AI News | TestingCatalog (@testingcatalog) 9 avril 2026
Engramme uses Large Memory Models, promising near-zero… https://t.co/M0Wpgav6uS pic.twitter.com/Dsw2u9MnF8Engramme's memory API is now live in public beta and built on an entirely new AI architecture, not Transformers, purpose-built to give apps persistent human memory without any search or prompting from the user. Engramme uses Large Memory Models, promising near-zero
-
New SGLang Course: Efficient LLM and Image Generation Inference
By
–
New course: Efficient Inference with SGLang: Text and Image Generation, built in partnership with LMSys @lmsysorg and RadixArk @radixark, and taught by Richard Chen @richardczl, a Member of Technical Staff at RadixArk.
— Andrew Ng (@AndrewYNg) 9 avril 2026
Running LLMs in production is expensive, and much of that… pic.twitter.com/baiT6LKDYYNew course: Efficient Inference with SGLang: Text and Image Generation, built in partnership with LMSys @lmsysorg and RadixArk @radixark, and taught by Richard Chen @richardczl, a Member of Technical Staff at RadixArk. Running LLMs in production is expensive, and much of that cost comes from redundant computation. This short course teaches you to eliminate that waste using SGLang, an open-source inference framework that caches computation already done and reuses it across future requests. When ten users share the same system prompt, SGLang processes it once, not ten times. The speedups compound quickly, especially when there's a lot of shared context across requests. Skills you'll gain: – Implement a KV cache from scratch to eliminate redundant computation within a single request – Scale caching across users and requests with RadixAttention, so shared context is only processed once – Accelerate image generation with diffusion models using SGLang's caching and multi-GPU parallelism Join and learn to make LLM inference faster and more cost-efficient at scale! deeplearning.ai/short-course…
→ View original post on X — @andrewyng, 2026-04-09 17:11 UTC
-

Gemini Adds Interactive Visualization in Chat
By
–
Gemini can now help visualize complex topics through interactive experiences directly in chat. "Show me the visualization" button will appear under certain questions, which could trigger this new experience. Testing time
-

Free Gemma 4 Fine-tuning in Google Colab No Coding Required
By
–
Fine-tune and run Gemma 4 and 500+ open source AI models in a free Google Colab.
— Shubham Saboo (@Saboo_Shubham_) 9 avril 2026
Just choose an AI model and hit start training. No need to write a single line of code.
100% free and Open Source.pic.twitter.com/u37nmZB2iXFine-tune and run Gemma 4 and 500+ open source AI models in a free Google Colab. Just choose an AI model and hit start training. No need to write a single line of code. 100% free and Open Source.
-

Meta’s AI Comeback Begins – New Superintelligence Newsletter
By
–
Today's Newsletter on Superintelligence has just been sent! Today's main article is: "Meta's AI Comeback Begins" In addition: – Hot AI news – Infographs – and much more Subscribe for free – link down below!
→ View original post on X — @kimmonismus, 2026-04-09 17:05 UTC
-

Anthropic Mythos announcement critique: Sandboxing disabled limitations
By
–
Yesterday’s Mythos announcement from Anthropic was overblown. • Sandboxing was turned off, so test didn’t show much about the real world. • Cheap open-weight models can (already) do some similar stuff • No evidence that Mythos itself is a major qualitative jump. In short,
-
AI Newsmakers Lists: See All Key Players at Once
By
–
That's why I built my lists so you can see everyone in one shot. Start with AI Newsmakers, even though the rest are awesome: https://
x.com/scobleizer/lis
ts
… -
ElevenLabs Cloud API: Fast Production Deployment with Multiple Voice Models
By
–
For everyone else, our cloud API is the fastest path to production. You get access to all our models and voices, and automatic scaling – managed entirely by ElevenLabs. We support data residency in the US, EU and India; Zero Retention Mode for enhanced privacy; and all of the