The DTV paper[1] from 2024 uses a 62B model specialized in math and reasoning, building up from 48.6% baseline by putting more emphasis on the symbolic side. What if we did the opposite, matching the baseline with specialized system on CPU and adding more parameters to improve?
@alexjc
-

2.6K Parameter RL System Matches 62B Specialized Model
By
–
Oh My! I just built a new RL-native system with 2.6k parameters that matches a specialized 62B model from a 2024 paper, specifically on GSM8k. That's over 20,000,000x smaller than the baseline — but the approach is so different the traditional comparison doesn't make sense? 😉
-
Hypocrisy in AI: Silent until personal impact
By
–
As it stands, he just looks like the guy "When they came for others I said nothing", but suddenly "they came for me…"
-
AI Impact on Artists: Ethics and Moral Clarity Debate
By
–
Watching Julian get destroyed by the Machiavellians for making a questionable point badly is bittersweet. He could have made same point about AI & artists (no impact on cancer cures) and been morally justified, but he didn't have the clarity nor conviction at the time / on this.
-
Dense Models Beat MoE: Western AI Companies Strategic Shift
By
–
It's a big macro trend shift as they revert back to old dense models, while still comparable or better than MoE baselines. "Leak" because the big Western companies figured it out a while ago, letting the OSS community & China go down the MoE path — notably more complex + ~worse
-
Valve’s Strategic Direction Versus Tim’s Approach in Tech
By
–
Tim has been going and still is going in the wrong direction, Valve is right — as usual for a multi-decade market leader. But they need much more detailed tags, then it's extremely useful!
-
Do Tech Leaders Know and Accept AI Impact Consequences?
By
–
Have you considered that they know exactly what they are doing and the impact you foresee is known to them?
-

Model Benchmarks: 7% Performance Gap Versus Marketing Claims
By
–
If the initial benchmarks scores (and graphs used for PR) showcased a 3x reduction in size for the same performance, I think the broader public reception would have been less tepid. Just looking at this, it just seems 7% behind other existing models in the 4.x series…
-
Latent Process Planning: 6K Parameters Matching 100M Model Capability
By
–
A latent process that operates over plans? I've been working on this recently! What's most fascinating to me: my 6k parameter system can match aspects of models 100,000x bigger. Scaling laws apply very differently too…
-
Gemini 3 Cursor AI bugs and Claude Opus workarounds
By
–
For Gemini 3, I don't rule out bugs in @cursor_ai — as many new features don't work, worktrees getting trashed or even renamed (!) mid-way through agent working. But since Opus 4.5 manages around those bugs, it can't be entirely on the Cursor side.
