I think the story that was shared in the Mythos System Card still has the signs of flawed LLM writing (which looks like good writing at first glance): A story that doesn't really hold together logically, but sounds like it should. The back-and-forth banter. Lack of characters.
GENERATIVE AI
-
LangSmith Enables Agent Tracing and Optimization Loop
By
–
as always, it's an exciting time to be working at LangChain! https://t.co/JrgR2YMT4Q
— Sam Crowder (@samecrowder) 8 avril 2026as always, it's an exciting time to be working at LangChain! LangChain (@LangChain) LangSmith 🤝 San Francisco You don't know what your agents will do until you actually run them. What works in demos can break in the real world. Without tracing and evals, you're just guessing at why. Track what your agent actually does. Optimize and fix your agents. Then measure whether your fixes work. That loop is how agents get better, and LangSmith is built to power that workflow. — https://nitter.net/LangChain/status/2041656189860393383#m
-
MoE Architecture and Streaming Experts Integration
By
–
It's a MoE though so maybe someone will get streaming experts to work for it
-
Summarize 0.13 Released with Video Slides and GPT-5.4 Support
By
–
📝Summarize 0.13 is out! 🎞️ Local video slides (–slides) 🤖 More model backends (GitHub Copilot) 🧠 Better GPT-5.4 support 📺 Better media handling (HLS detection.m3u8) It graduated from my tap to official homebrew formula! 🍺 brew install summarize github.com/steipete/summariz…
→ View original post on X — @scobleizer, 2026-04-08 00:08 UTC
-
OpenAI Anthropic IPO Benchmark Hacking Financial Concerns
By
–
What’s approaching the singularity is the benchmark hacking and financial legerdemain as OpenAI and Anthropic careen toward their IPOs.
-
Codex Feedback: User Experience Insights Shared
By
–
love itt! keen to hear your feedback and how your experience with codex has been!
-
Codex Improves Daily for OpenClaw Code Generation
By
–
What option do they have – pay for Claude API? Just kidding – Codex is getting better everyday for OpenClaw.
-

New Models Added to BullshitBench: Qwen Performance Analysis
By
–
I did a big clean up of some new models to add to the BullshitBench – none of them are particularly interesting tbh. Qwen scored relatively well, but below Qwen 3.5
-
Codex hits 3M weekly users, rate limits reset for builders
By
–
Insane that 3,000,000 of you use codex every week! Thanks for all your support, feature requests and bug reports! To celebrate Tibo has reset the rate limits, time to build is NOWWW!
-
Codex Resets Usage Limits at 3 Million Weekly Users
By
–
To celebrate 3 million weekly codex users, we are resetting usage limits. We will do this every million users up to 10 million. Happy building!
