every ML conference I've been to since 2019 has had no fewer than three papers proposing new techniques for subquadratic attention
@jxmnop
-

New sub-quadratic attention technique makes long-context LLMs 10x cheaper
By
–

"Introducing a breakthrough new technique for sub-quadratic attention, making long-context LLMs 10x cheaper without sacrificing performance" Me:
-

Subquadratic Attention and Data Quality: Challenges for Large Context AI Models
By
–
people on here are dumb. the latest subquadratic attention trick might produce a model that *processes* 1M tokens (or 12M..) without going insane, but that doesn't make it good the real problem isn't the architecture, it's the data. humans haven't generated many contiguous
-
Codex Boosts Experiments but Only 15% of Results Are Trustworthy
By
–
with Codex, i can run 10x the experiments out of these experiments, i can trust about 15% of the results conclusion: i am 50% more productive with codex
-
Why 1M-Context Models Still Don’t Work Beyond 200K Tokens
By
–
it is endlessly fascinating to me that we still don't have a true 1M-context model it's an unusual case where the infra is far ahead of the science. Claude discontinued 1M+ context bc it didn't really work past ~200k we don't have the right data? training techniques? not sure
-
Claude Code Unusable After Excessive Vibecoding Dogfooding
By
–
very interesting that Claude Code is the ultimate product for vibecoding, and Claude Code's engineers vibecoded Claude Code so hard it became unusable an entire company overdosing on Dogfood
-
Distillation Challenges Widen as AI Agents Complexity Increases
By
–
I think the world of agents makes distillation much harder, so the gap is starting to widen
-
Growing Gap Between Open-Source and Frontier AI Models
By
–
this is why there is a massive (growing) gap between open-source and the frontier
-
American AI Labs Stockpile Data While China Awakens to Reality
By
–
over the last few years ~four american AI labs have stockpiling many petabytes of Good Data. no one else has this, it costs billions of dollars once china wakes up to the reality of cold-hard Data (purchased, stockpiled, compounding), the landscape will undergo a giant shift
