shoutouts:
• why multi-agent LLM systems fail? (arXiv:2503.13657) — @mertcemri @melissapan + @istoica05 @matei_zaharia @profjoeyg @adityagp & team • DSPy (arXiv:2310.03714) — @lateinteraction + @hazyresearch lab & co-authors
• GRASP (arXiv:2605.29668) — Jonas Moll,
SOFTWARE
-
Shoutouts to multi-agent LLM systems, DSPy, and GRASP papers
By
–
-

Gated approach to agent self-modification via forking and testing
By
–
less novel, but still very interesting impo is the gated approach to self-modification the agent basically forks itself, propose a patch, run through multiple tests (static/sandbox/diff), and something called a binding held out gate before modificaiton lands
-
PettiChat uses AI to turn pet behavior into human conversations
By
–
PettiChat Uses #AI to Turn Pet Behavior Into Human Conversations
— Ronald van Loon (@Ronald_vanLoon) 10 juin 2026
by @IntEngineering#EmergingTech #Technology #Innovation #Tech #FutureTech pic.twitter.com/aiezKTRaSxPettiChat Uses #AI to Turn Pet Behavior Into Human Conversations
by @IntEngineering #EmergingTech #Technology #Innovation #Tech #FutureTech -
Cursor’s design mode enhances MagicPath creations
By
–
There are also a few Cursor-only things that are really cool. For example, you can use Cursor's very own design mode to work with MagicPath creations.
— Pietro Schirano (@skirano) 10 juin 2026
It's pretty cool! pic.twitter.com/smffWLajpkThere are also a few Cursor-only things that are really cool. For example, you can use Cursor's very own design mode to work with MagicPath creations. It's pretty cool!
-

Cohere Transcribe ranks #1 in enterprise speech recognition tests
By
–
These tests measure performance in varying signal-to-noise conditions: the kinds of audio found in meeting rooms, contact centres, & phone calls. In other words, environments where enterprise speech applications actually operate. Cohere Transcribe ranked #1 across every metric:
-
Transcribe surpasses IBM and NVIDIA speech models with 17.9 WER
By
–
Transcribe achieved a 17.9 WER – nearly 2 points ahead of IBM Granite Speech and 3.6 points ahead of NVIDIA’s Parakeet. Still Apache 2.0 and runs on your laptop. Enterprise performance developer ergonomics. Full results:
-
Transcribe tops OpenASR and far-field speech benchmarks
By
–
In March, Transcribe topped the OpenASR leaderboard for general-purpose speech recognition. Today, it leads a benchmark designed to go beyond and test robustness in real-world, far-field audio environments. Give it a try and share back what you build:
-

Cohere Transcribe, open-source ASR model, #1 on Hugging Face benchmark
By
–
Cohere Transcribe, our open-source speech recognition model, is #1 on the new @huggingface Far-Field ASR benchmark.
-

LangSmith Fleet template: Software Engineer coding agent from Slack, Linear, GitHub
By
–
LangSmith Fleet template spotlight: Software Engineer Ships code from Slack, Linear, and GitHub in a sandbox A coding agent that takes issues from @Linear
, writes and verifies the code, and opens a PR. Triggered directly from Slack. -

LLM Gateway integrates observability and application into a single tool
By
–
Obtaining both observability and application once meant assembling: A separate gateway, A guardrails platform, An observability stack… then correlating signals between the three when a problem occurred. LLM Gateway integrates both into