Our official implementation of DiffusionBlocks on image classification using Vision Transformers (ViT)
RESEARCH
-
Nostalgia for early 2024 AI pipelines and progress
By
–
In the early days you had to build so many things just to make this work.
— Pietro Schirano (@skirano) 31 mai 2026
Computer use wasn’t a thing, models were slow, and you had to pipeline so many models. I do miss those times but it’s amazing to see how far we’ve come.
This was 2024.
https://t.co/Gd8NI4b5vB https://t.co/WB7f7ivZxiIn the early days you had to build so many things just to make this work. Computer use wasn’t a thing, models were slow, and you had to pipeline so many models. I do miss those times but it’s amazing to see how far we’ve come. This was 2024. https://
x.com/skirano/status
/1804219134999192048?s=20
… -
Comparison of Seedance 2.0 and Gemini Omni Flash: Omni’s hidden advantage
By
–
Everyone is comparing Seedance 2.0 and Gemini Omni Flash side by side.
— God of Prompt (@godofprompt) 31 mai 2026
They're missing the point.
Sure, Omni might not be as good as Seedance right now, but it has one advantage… pic.twitter.com/Z4dOnnRZjjEveryone is comparing Seedance 2.0 and Gemini Omni Flash side by side. They're missing the point. Sure, Omni might not be as good as Seedance right now, but it has one advantage…
-

Hands-on Mathematics for Deep Learning: Big Data, AI, Machine Learning
By
–
Hands-on #Mathematics for Deep Learning. #BigData #Analytics #DataScience #AI #MachineLearning #IoT #IIoT #Python #RStats #TensorFlow #JavaScript #ReactJS #CloudComputing #Serverless #DataScientist #Linux #Books #Programming #Coding #100DaysofCode https://
geni.us/Math-Deep-Lear
ning
… -

Stunning AI Solution for 80-Year-Old Problem from GPT Shocks Mathematicians
By
–

Stunning AI Solution For 80-Year-Old Problem from GPT Shocks Mathematicians! #BigData #Analytics #DataScience #AI #MachineLearning #NLProc #LLM #IoT #IIoT #PyTorch #Python #RStats #TensorFlow #Java #JavaScript #ReactJS #Mathematics #GoLang #CloudComputing #Serverless
-
Introducing DiffusionBlocks for independent block-wise neural network training
By
–
DiffusionBlocks: Training Neural Networks One Block at a Time https://
pub.sakana.ai/diffusionblock
s/
… tl;dr We introduce DiffusionBlocks, a principled framework that partitions a residual network into blocks and trains each one independently by reinterpreting block-wise updates as the reverse -

World Action Models let robots imagine future before acting
By
–
What if robots could imagine the future before acting? Researchers from Fudan, NUS, and Shanghai Innovation Institute unveil World Action Models (WAMs). They merge predictive world models with action generation, so robots plan by simulating how the environment evolves, not
-
Opus 4.8 improves, but GPT-5.5 xhigh beats it for cheaper cost
By
–
Opus 4.8 is a solid jump over Opus 4.7 on DeepSWE, while also lowering the average cost per task.
— Chubby♨️ (@kimmonismus) 31 mai 2026
However, GPT-5.5 xhigh still beats it by a pretty clear margin while being cheaper.
OpenAI has been cooking insanely hard with its models lately. Really excited to see what GPT-5.6… https://t.co/UC7Rl2cX6lOpus 4.8 is a solid jump over Opus 4.7 on DeepSWE, while also lowering the average cost per task. However, GPT-5.5 xhigh still beats it by a pretty clear margin while being cheaper. OpenAI has been cooking insanely hard with its models lately. Really excited to see what GPT-5.6
-

GPT-5.5 beats Claude Opus 4.8 on DeepSWE with faster, cheaper performance
By
–


GPT-5.5 is #1 on DeepSWE, a hard long-horizon coding benchmark 70% pass@1 vs 58% for Claude Opus 4.8. And GPT-5.5 gets there with:
~2x faster runs
~1/2 the cost
~1/3 the output tokens Literally, better intelligence per dollar, per minute, per task. -

Open weight models lag behind closed models but shrinking
By
–
Open weight models have lagged the state of the art closed models by four months The lag is shrinking not expanding unlike what paid influencers would like you to believe