MathArena: Evaluating LLMs on Uncontaminated Math Competitions Balunović et al.: https://
arxiv.org/abs/2505.23281 #ArtificialIntelligence #DeepLearning #MachineLearning
GENERATIVE AI
-

MathArena: Evaluating LLMs on Uncontaminated Math Competitions
By
–
-
Attention Mechanisms and Core Deep Learning Components Explained
By
–
Attention is all you need. Oh, and also MLPs, layer norm, resnets, positional encoding, tokenization, Adam, fine-tuning, GPUs, etc.
-
LLMs in AIOps: Survey of 183 Papers and Methods
By
–
10. A Survey of AIOps This survey analyzes 183 papers to evaluate how LLMs are being used in AIOps, focusing on data sources, task evolution, applied methods, and evaluation practices.
-
Scaling Reinforcement Learning for Enhanced Reasoning in Small Models
By
–
6. Scaling up RL This paper investigates how prolonged RL can enhance reasoning abilities in small language models across diverse domains.
-

Machine Bullshit: LLMs and Indifference to Truth
By
–
7. Machine Bullshit This paper introduces the concept of machine bullshit, extending Harry Frankfurt’s definition, discourse made with indifference to truth, to LLMs.
-

Chain-of-Thought Monitorability for AI Safety Oversight
By
–
4. Chain-of-Thought Monitorability Proposes that language-based CoT reasoning in LLMs offers an opportunity for AI safety by enabling automated oversight of models’ internal reasoning processes.
-

Agentic-R1: 7B Language Model with Dynamic Tool-Based Reasoning
By
–
3. Agentic-R1 This paper introduces Agentic-R1, a 7B language model trained to dynamically switch between tool-based execution and pure text reasoning using a novel fine-tuning framework called DualDistill.
-
Prompts and Use Cases for Grok 4
By
–
People need to know about these prompts and use cases for Grok 4
-
GPT-5 Auto-Switching Between o3 and 4o Models
By
–
Even if GPT-5 did nothing besides switching people between o3 and 4o automatically, it would really transform most people’s view of AI. Very few people, even paying users, know that they should often switch to a more capable model, and when you show them o3, they are impressed.
-

OpenAI Agents: Testing ChatGPT’s New Features
By
–
NEW VIDEO in the LAB! This week has been intense for OpenAI with the release of the new AGENTS in ChatGPT and their big achievement in the math and programming competitions. Today I'm sharing my opinion on the agents after several days of testing them! Link below