Are LLMs truly capable of “PhD level” reasoning? FormulaOne: Measuring the Depth of Algorithmic Reasoning Beyond Competitive Programming Evaluated through PhD-level Dynamic Programming problems, and spoiler alert, AI code-solvers are only able to get under 1% right (for now…?)
@askalphaxiv
-

LLMs Solve Only 2% of Unsolved Scientific Questions
By
–
UQ: Assessing Language Models on Unsolved Questions This research assessed LLMs on actual unsolved questions & not benchmarks, with today’s best model only able to solve 10/500 questions!
-

Token Order Prediction Improves Language Model Pre-training
By
–
A better auxiliary LLM pre-training objective just dropped! Predicting the Order of Upcoming Tokens Improves Language Modeling Teach the model to rank which words are likely to come sooner vs later instead of guessing future words makes LLMs better with minimal overhead
-

Model-Driven FLOP Allocation Achieves 80% Sparsity, 4x CPU Speedups
By
–
What happens when you let a model decide how to allocate FLOPs? Introducing Compute Where It Counts, a new trainable sparsity paradigm released by @crystalAIorg that beats SOTA methods, enabling 80% sparsity and 4x+ speedups on CPU.
-
V-JEPA 2 Learns World Models from Million Hours Video
By
–
internet videos to robot actions? @AIatMeta
's V-JEPA 2 learns a world model from over 1 million hours of video, enabling it to perform complex, real-world robotic tasks zero-shot. Join author Nicolas Ballas for a talk on V-JEPA 2 this Friday! -

Top LLMs Fail on Hard Coding Problems: LiveCodeBench Pro
By
–
Why do top LLMs fail completely on hard coding problems? We'll have author Peiyao Shang from @SentientAGI discussing their work on the LiveCodeBench Pro benchmark. Join us this Friday! https://
lu.ma/v45b9ltc -
arXiv Preprint Server Support and PDF Upload Features
By
–
Currently, arXiv is the only preprint server we support. However, you can upload any PDF as a private paper in your library. We recommend installing our Chrome extension which allows you to do this while viewing any PDF on your browser. Hope this helps!
-
AI-Driven Collider Analysis Explores Dark Matter Detection
By
–
Dark matter remains one of the most compelling mysteries in high-energy physics. Tomorrow we welcome Noah Bray-Ali to discuss using AI-driven collider analysis to explore a new 'volume frontier' for dark matter detection Signup:
-
Paper Chat Interface Redesign with Resizable Windows
By
–
Thanks for the suggestions, we're actually doing an overhaul of the paper page so that the chat can be resize-able among other things. Also the UX might be a bit unclear right now, but if you click "switch chat" it will show previous chats (should probably rename to "history")

