
Kimi K2.6 from @Kimi_Moonshot is a new open-source SOTA on HLE with tools, SWE Bench Pro, and other benchmarks!
– HLE w/ tools – 54.0
– SWE-Bench Pro – 58.6
– SWE-bench Multilingual – 76.7 Looks like it is testing time now

By
–

Kimi K2.6 from @Kimi_Moonshot is a new open-source SOTA on HLE with tools, SWE Bench Pro, and other benchmarks!
– HLE w/ tools – 54.0
– SWE-Bench Pro – 58.6
– SWE-bench Multilingual – 76.7 Looks like it is testing time now
By
–
Getting LLMs to simulate “true” randomness or generate diverse outputs is surprisingly difficult. We found a simple prompting trick that solves this by having the model generate and manipulate a random string. To be presented at #ICLR2026 this week!
— hardmaru (@hardmaru) 20 avril 2026
Blog: https://t.co/CyevqqJ5ej https://t.co/dN0yZZ5Mij
Getting LLMs to simulate “true” randomness or generate diverse outputs is surprisingly difficult. We found a simple prompting trick that solves this by having the model generate and manipulate a random string. To be presented at #ICLR2026 this week! Blog: https://
pub.sakana.ai/ssot

By
–
Moonshot AI is rolling out Kimi K2.6 on Kimi Chat and APIs. All models got upgraded. – Kimi K2.6 Instant
– Kimi K2.6 Thinking – Kimi K2.6 Agent – Kimi K2.6 Agent Swarm Did you get it already?

By
–
Can LLMs flip coins in their heads?
— Sakana AI (@SakanaAILabs) 20 avril 2026
When prompted to “Flip a fair coin” 100 times, the heads to tails ratio drifts far from 50:50. LLMs can understand what the target probability should be, but generating outputs that faithfully follow a given distribution is a separate problem.… pic.twitter.com/XyF7Xnj8Ll
Can LLMs flip coins in their heads? When prompted to “Flip a fair coin” 100 times, the heads to tails ratio drifts far from 50:50. LLMs can understand what the target probability should be, but generating outputs that faithfully follow a given distribution is a separate problem.
By
–
Devs are taking full advantage of the 262K limit, absolutely stuffing the context window without worrying about the bill Test it in your own workflows right now:
→ https://
openrouter.ai/openrouter/ele
phant-alpha
…

By
–
Top autonomous agents are already feasting on it! → @OpenClaw has pushed 107 Billion tokens
→ @Kilocode and @Claudeai routed 54 Billion combined
→ The Hermes Agent is nearing 26 Billion tokens It maintains 100% provider uptime. It is fast… and it won't cost you a dime.

By
–
Elephant Alpha is built for raw reliability. In AI BENCHY's "Anti-AI Tricks" tests, it hit a perfect 10.0 for consistency, completely ignoring the flaky behavior of other open models. It also skips the reasoning token trap: You get direct answers with zero reasoning bloat,

By
–
This is fascinating. A mystery stealth provider just flipped the entire market upside down. They dropped a massive 100B parameter model on @OpenRouter called Elephant Alpha. Best part? It's absolutely FREE Here's the tech stack you get: → 100B parameter intelligence.
→

By
–
Qwen-3.6 Max preview shows some impressive results. However, i prefer their benchmarks against opus 4.7 instead of opus 4.5..

By
–
Google researchers built a memory system that makes Transformers look wasteful. The tech giant's Titans model could process text cheaply, but its memory was fixed. Once full, old information got overwritten. Transformers have the opposite issue. They remember everything but