No issues here. Nemotron is still my main LLM for most simple tasks and it taps Claude and GPT API models for anything more complex. Speed isn’t really an issue. Usually get a response within 20-seconds or so. Not the fastest in the world but not slow enough to be an issue.
AI
-
AI Industry Thrives on X Platform Community
By
–
You can see the AI industry is here on X: https://
alignednews.com/ai and I couldn't build this site out of any of those other places. I hope you reconsider. Your fans and potential fans are here. -

2026: The Year AI Gets Real – Article Overview
By
–
2026: The Year #AI Gets Real by @Khulood_Almani #ArtificialIntelligence #MachineLearning #ML
→ View original post on X — @ronald_vanloon, 2026-04-09 17:23 UTC
-
PDF Conversion Challenges for Large Language Models
By
–
I just tried it this morning on the 245-page Mythos pdf and it failed badly and the outputs were all mangled. Converting pdfs is really hard, I think it has to probably be a Skill not a program, for a SOTA LLM for it to work properly.
-
AI Monitoring Agent Built to Track AI Community on X
By
–
Nice list. Mine are far far far more complete: https://
x.com/scobleizer/lis
ts
… And I built an AI to watch everyone in AI here on X: -
Optimal timing for maximum AI intervention impact
By
–
How do you determine the optimal intervention point for maximum benefit?
-
𝕏 Chat Messaging Platform Emphasizes Privacy Features
By
–
Use 𝕏 Chat for messaging and voice/video calls. Comes with this great benefit of actual privacy.
-
Engramme’s Memory API Launched in Beta
By
–
Engramme's memory API is now live in public beta and built on an entirely new AI architecture, not Transformers, purpose-built to give apps persistent human memory without any search or prompting from the user.
— 🚨 AI News | TestingCatalog (@testingcatalog) 9 avril 2026
Engramme uses Large Memory Models, promising near-zero… https://t.co/M0Wpgav6uS pic.twitter.com/Dsw2u9MnF8Engramme's memory API is now live in public beta and built on an entirely new AI architecture, not Transformers, purpose-built to give apps persistent human memory without any search or prompting from the user. Engramme uses Large Memory Models, promising near-zero
-
New SGLang Course: Efficient LLM and Image Generation Inference
By
–
New course: Efficient Inference with SGLang: Text and Image Generation, built in partnership with LMSys @lmsysorg and RadixArk @radixark, and taught by Richard Chen @richardczl, a Member of Technical Staff at RadixArk.
— Andrew Ng (@AndrewYNg) 9 avril 2026
Running LLMs in production is expensive, and much of that… pic.twitter.com/baiT6LKDYYNew course: Efficient Inference with SGLang: Text and Image Generation, built in partnership with LMSys @lmsysorg and RadixArk @radixark, and taught by Richard Chen @richardczl, a Member of Technical Staff at RadixArk. Running LLMs in production is expensive, and much of that cost comes from redundant computation. This short course teaches you to eliminate that waste using SGLang, an open-source inference framework that caches computation already done and reuses it across future requests. When ten users share the same system prompt, SGLang processes it once, not ten times. The speedups compound quickly, especially when there's a lot of shared context across requests. Skills you'll gain: – Implement a KV cache from scratch to eliminate redundant computation within a single request – Scale caching across users and requests with RadixAttention, so shared context is only processed once – Accelerate image generation with diffusion models using SGLang's caching and multi-GPU parallelism Join and learn to make LLM inference faster and more cost-efficient at scale! deeplearning.ai/short-course…
→ View original post on X — @andrewyng, 2026-04-09 17:11 UTC