Anyhow, its not bad. Just not the vibe level that the benchmarks might indicate. And, for a first re-entry into the frontier model space, given the engineering efficiencies they achieved, it feels like a solid attempt. I am sure we will see better from Meta in the future.
@emollick
-
Meta’s Muse Spark Thinking lags behind Big Three, seems weird
By
–
After playing with it a bit, Meta's Muse Spark Thinking is fine so far, but really doesn't match the current Big Three models. It also is a bit… weird. Like some strange language & tone, a little loose with facts, etc.
— Ethan Mollick (@emollick) 9 avril 2026
And here is how it does on the neo-gothic shader test. https://t.co/hxsUTOVTO3 pic.twitter.com/aDlcslfdSuAfter playing with it a bit, Meta's Muse Spark Thinking is fine so far, but really doesn't match the current Big Three models. It also is a bit… weird. Like some strange language & tone, a little loose with facts, etc. And here is how it does on the neo-gothic shader test.
-
Methodologies for Reducing Errors in AI Systems
By
–
Some approaches:
1) Multiple reviews. Some papers already show having many AI team members review a problem reduces errors
2) Building in tests and checkpoints 3) Multiple independent answers that are cross-checked
4) Clear processes that escalate to smarter systems for review -
Applying organizational risk management structures to mitigate AI hallucinations
By
–
Hallucinations remain in LLMs, but note that over centuries we have developed complicated, successful machines that take uncertain output from unreliable sources & reduce the risk of errors. We call those machines organizational structures & we can apply similar approaches to AI
-

Discussion on Meta’s latest model and open-weights strategy
By
–
Seems like a good model from Meta that is still trailing the current series of releases. The most important thing to note is that it is not open weights. That was the main reason that Meta's models were so important. Without that, it is a lot harder to predict the value of Spark
-
CISO offices face urgent AI security risks as capabilities diffuse
By
–
Curious how many large organization CISO offices have taken the Mythos red team reports as the red alert that it is. (I suspect very few) Based on historical trends in AI they have, at most, about six to nine months until those capabilities become widely diffused to bad actors.
-

Frontier AI Model Capabilities and Potential Misuse Risks
By
–

In different hands, Mythos would be an unprecedented cyberweapon I am not sure how we deal with this, except to note a narrow window where we know only 3 companies could be at this level of capability. But it may be Chinese models (maybe open weights ones?) get there in 9 months
-
LLMs struggle with creative fiction and require new human-led benchmarks
By
–
Writing fiction seems to be a genuine weak spot for LLMs that is not improving as rapidly as almost every other area. There may be a lot of reasons why this is happening. It would be a really interesting benchmark (but you would need human judges, AI judges love AI fiction).
-

Critique of LLM writing quality and logical coherence in system cards
By
–
I think the story that was shared in the Mythos System Card still has the signs of flawed LLM writing (which looks like good writing at first glance): A story that doesn't really hold together logically, but sounds like it should. The back-and-forth banter. Lack of characters.
-

Technical analysis of Claude model behavior and personality traits
By
–

SuperClaude (Mythos) still seems irreducibly Claude-y given the transcripts in the system card. Here two versions of Mythos are forced to talk to each other across multiple rounds. They are less philosophical than Opus 4.6 or spiritual than Opus 4.1, but still very Claude-like.