7. Machine Bullshit This paper introduces the concept of machine bullshit, extending Harry Frankfurt’s definition, discourse made with indifference to truth, to LLMs.
LLMS
-

REST Benchmark Framework Evaluates Large Reasoning Models Robustness
By
–
5. Stress Testing Large Reasoning Models Proposes a new benchmark framework called REST to evaluate the robustness of Large Reasoning Models (LRMs) under multi-question stress.
-

Agentic-R1: 7B Language Model with Dynamic Tool-Based Reasoning
By
–
3. Agentic-R1 This paper introduces Agentic-R1, a 7B language model trained to dynamically switch between tool-based execution and pure text reasoning using a novel fine-tuning framework called DualDistill.
-

Chain-of-Thought Monitorability for AI Safety Oversight
By
–
4. Chain-of-Thought Monitorability Proposes that language-based CoT reasoning in LLMs offers an opportunity for AI safety by enabling automated oversight of models’ internal reasoning processes.
-

Study Evaluates LLM Performance Across Extended Context Lengths
By
–
2. Context Rot This comprehensive study by Chroma evaluates how state-of-the-art LLMs perform as input context length increases, challenging the common assumption that longer contexts are uniformly handled.
-
Prompts and Use Cases for Grok 4
By
–
People need to know about these prompts and use cases for Grok 4
-

LLMs Performance on 2025 International Math Olympiad Evaluation
By
–
Not Even Bronze: Evaluating LLMs on 2025 International Math Olympiad https://
matharena.ai/imo/ Nice blog post from the team behind MathArena: Evaluating LLMs on Uncontaminated Math Competitions (
https://
arxiv.org/abs/2505.23281) providing independent analysis of LLM performance on IMO. -

Agentic Doc: Python Library for Document Data Extraction
By
–
Turn documents into LLM-ready data! Agentic Doc is a Python library for agentic document extraction. It pulls structured data from visually complex documents like tables, images, and charts and returns a hierarchical JSON with exact element locations. 100% Open Source
-

LLM Intelligence: Key Concepts for Business Professionals
By
–
This week, we have the fourth free video from our recent course "AI for Business Professionals" to give you an idea of the concepts we teach in the full course, and share some insights. In this one, we focus on this new type of "intelligence" that LLMs have, that is different
-
LLM Capabilities Without Tools Exceed Expectations in Math
By
–
I thought that was true, but apparently I was wrong – LLMs without tools are a lot more capable than I expected, at least for things like IMO math problems I still think tools are the most important technique in AI engineering generally
