AI-assisted programming was happening, but the models weren't yet strong enough to support true don't-even-look-at-the-code vibe coding
LLMS
-
Apple Launches Multilingual Foundation Models for On-Device AI
By
–
10. Apple Intelligence Foundation Language Models Apple introduces two multilingual, multimodal foundation models: a 3B-parameter on-device model optimized for Apple silicon and a scalable server model using a novel PT-MoE transformer architecture.
-

Deep Researcher: Test-Time Diffusion for Report Generation
By
–
8. Deep Researcher with Test-Time Diffusion Rather than relying on static inference strategies like CoT or best-of-n sampling, this work frames the report generation process as a diffusion process.
-
MCPEval: Open-Source LLM Agent Evaluation Framework
By
–
9. MCPEval MCPEval is an open-source framework that automates end-to-end evaluation of LLM agents using a standardized Model Context Protocol, eliminating manual benchmarking.
-

Inverse Scaling in Large Reasoning Models Test-Time Compute
By
–
6. Inverse Scaling in Test-Time Compute Presents a systematic study of inverse scaling in large reasoning models, where increasing the test-time compute (i.e., reasoning length) harms rather than helps model performance.
-
Compute-Optimal Strategies for Many-Shot In-Context Learning
By
–
7. Towards Compute-Optimal Many-Shot In-Context Learning Proposes practical strategies for reducing the cost of many-shot in-context learning while preserving or improving performance. https://
arxiv.org/abs/2507.16217 -

In-Context Learning Without Weight Updates in LLMs
By
–
5. Learning without Training This paper provides a theoretical and empirical explanation for how LLMs exhibit in-context learning, the ability to learn from examples in a prompt without weight updates.
-

Routine: Structured Planning Format for LLM Agent Tool-Calling
By
–
4. Structural Planning for LLM Agents This paper introduces Routine, a structured planning format designed to improve the stability and accuracy of LLM agents executing multi-step tool-calling tasks in enterprise settings.
-
Model Convergence Through Distillation and Subliminal Learning Transfer
By
–
This model convergence is quite perplexing. Possibly related to recent results on subliminal learning? Basically deeper knowledge correlations transfer when training via distillation. As the amount of data online from LLMs increases, it’s possible this makes them converge to some
-
Best AI Models Capabilities Six Months Ago
By
–
The best models could do 6 months ago. pic.twitter.com/Kg5iAczkaS
— Ethan Mollick (@emollick) 27 juillet 2025The best models could do 6 months ago.