10. A Survey of AIOps This survey analyzes 183 papers to evaluate how LLMs are being used in AIOps, focusing on data sources, task evolution, applied methods, and evaluation practices.
@dair_ai
-
Q-Chunking: Reinforcement Learning for Long-Horizon Sparse-Reward Tasks
By
–
9. Q-Chunking Q-chunking is a reinforcement learning approach that uses action chunking to improve offline-to-online learning in long-horizon, sparse-reward tasks.
-

Machine Bullshit: LLMs and Indifference to Truth
By
–
7. Machine Bullshit This paper introduces the concept of machine bullshit, extending Harry Frankfurt’s definition, discourse made with indifference to truth, to LLMs.
-
Scaling Reinforcement Learning for Enhanced Reasoning in Small Models
By
–
6. Scaling up RL This paper investigates how prolonged RL can enhance reasoning abilities in small language models across diverse domains.
-

REST Benchmark Framework Evaluates Large Reasoning Models Robustness
By
–
5. Stress Testing Large Reasoning Models Proposes a new benchmark framework called REST to evaluate the robustness of Large Reasoning Models (LRMs) under multi-question stress.
-

Chain-of-Thought Monitorability for AI Safety Oversight
By
–
4. Chain-of-Thought Monitorability Proposes that language-based CoT reasoning in LLMs offers an opportunity for AI safety by enabling automated oversight of models’ internal reasoning processes.
-

Agentic-R1: 7B Language Model with Dynamic Tool-Based Reasoning
By
–
3. Agentic-R1 This paper introduces Agentic-R1, a 7B language model trained to dynamically switch between tool-based execution and pure text reasoning using a novel fine-tuning framework called DualDistill.
-

Study Evaluates LLM Performance Across Extended Context Lengths
By
–
2. Context Rot This comprehensive study by Chroma evaluates how state-of-the-art LLMs perform as input context length increases, challenging the common assumption that longer contexts are uniformly handled.
-

xLSTMAD: Anomaly Detection for Multivariate Time Series
By
–
10. xLSTMAD This paper introduces xLSTMAD, the first anomaly detection method using an encoder-decoder xLSTM architecture tailored for multivariate time series.
-

Visual Structures Enhance GPT-4o Feature Binding Reasoning
By
–
9. Visual Structures Help Visual Reasoning This study shows that adding simple spatial structures (like horizontal lines) to images significantly boosts GPT-4o’s visual reasoning by improving feature binding.