more info on the benchmarks here: blog.google/innovation-and-a…
→ View original post on X — @demishassabis, 2026-03-26 18:54 UTC
By
–
more info on the benchmarks here: blog.google/innovation-and-a…
→ View original post on X — @demishassabis, 2026-03-26 18:54 UTC
By
–
We are already working on this with tools like AlphaFold and the work we are doing at @IsomorphicLabs
By
–
Introducing MiniMax M2.7 for understanding research papers 🚀
— alphaXiv (@askalphaxiv) 26 mars 2026
Highlight any section of a paper to ask questions and “@” other papers for quick context, comparisons, and benchmark references
Congrats @MiniMax_AI on the new model release! pic.twitter.com/GGctGYHpQG
Introducing MiniMax M2.7 for understanding research papers Highlight any section of a paper to ask questions and “@” other papers for quick context, comparisons, and benchmark references Congrats @MiniMax_AI on the new model release!
By
–
Now available in addition to Gemini 3 Flash and GPT 5.4 mini. Check out http://
alphaXiv.org!
By
–
Thanks for the feedback, Daily Papers submission eligibility now uses the arXiv announcement date instead of the submission date, so moderation delayed papers aren't unfairly excluded

By
–
To me this mostly illustrates the futility of robust jailbreaking prevention

By
–
Could your AI coding agent be secretly working against you? Researchers from Nanyang Technological University, University of Oxford, and collaborators introduce SkillJect. They've developed the first automated framework that weaponizes an AI agent's "skills". It uses a

By
–
CUA-Suite
— AK (@_akhaliq) 26 mars 2026
Massive Human-annotated Video Demonstrations for Computer-Use Agents
paper: https://t.co/hi1WnffQSY pic.twitter.com/ACjYOPcDzH
CUA-Suite Massive Human-annotated Demonstrations for Computer-Use Agents paper: https://
huggingface.co/papers/2603.24
440
…

By
–
Qworld
— AK (@_akhaliq) 26 mars 2026
Question-Specific Evaluation Criteria for LLMs
paper: https://t.co/UJvFr6xdpD pic.twitter.com/n7RfxIIyB9
Qworld Question-Specific Evaluation Criteria for LLMs paper: https://
huggingface.co/papers/2603.23
522
…
By
–
Small AI models and specialized vertical AI models are very brittle. Any unusual situation or out-of-distribution issue and they break down. You also won’t get emergent leaps or good problem solving. They still have uses, but benchmarks don’t do a good job of showing weaknesses