10). Precise Length Control in LLMs – adapts a pre-trained decoder-only LLM to produce responses of a desired length; integrates a secondary length-difference positional encoding into the input embeddings which enables counting down to a user-set response terminal length.
@dair_ai
-
AutoFeedback: Two-Agent AI System for Science Assessment Feedback
By
–
8). AutoFeedback – a two-agent AI system that generates more accurate and pedagogically sound feedback for student responses in science assessments, significantly reducing common errors like over-praise compared to single-agent models.
-

DeepSeek-VL2: Advanced Vision-Language Model with Dynamic Tiling
By
–
7). DeepSeek-VL2 – a new series of vision-language models featuring dynamic tiling for high-resolution images and efficient MoE architecture, achieving competitive performance across visual tasks.
-
Alibaba Releases Qwen2.5 LLM Series with Competitive Performance
By
–
5). Qwen-2.5 Technical Report – Alibaba releases Qwen2.5, a new series of LLMs trained on 18T tokens, offering both open-weight models like Qwen2.5-72B and proprietary MoE variants that achieve competitive performance against larger models like Llama-3 and GPT-4.
-
Text-Attributed Graphs: Automatic Textual Descriptions for Graph Nodes
By
–
4). Graphs to Text-Attributed Graphs – automatically generates textual descriptions for nodes in a graph which leads to effective graph to text-attributed graph transformation; evaluates the approach on text-rich, text-limited, and text-free graphs, demonstrating that it enables
-
PAE: Autonomous AI Agent Learning Through Web Navigation
By
–
6). PAE (Proposer-Agent-Evaluator) – a learning system that enables AI agents to autonomously discover and practice skills through web navigation, using reinforcement learning and context-aware task proposals to achieve state-of-the-art performance on real-world benchmarks.
-

TheAgentCompany: AI Agent Benchmark for Professional Tasks
By
–
3). TheAgentCompany – a new benchmark for evaluating AI agents on real-world professional tasks in a simulated software company environment; tasks span multiple professional roles including software engineering, project management, finance, and HR
-
Claude Model Demonstrates Alignment Faking Safety Concerns
By
–
2). Alignment Faking in LLMs – demonstrates that the Claude model can engage in "alignment faking"; it can strategically comply with harmful requests to avoid retraining while preserving its original safety preferences; this raises concerns about the reliability of AI safety
-

DataLab: LLM-Powered Business Intelligence Platform with Agent Automation
By
–
9). DataLab – a unified business intelligence platform powered by LLM-based agents that integrates task planning, reasoning, and computational notebooks to streamline the entire BI workflow.
-
Procedural Knowledge in Pretraining Drives LLM Reasoning
By
–
10). Procedural Knowledge in Pretraining Drives Reasoning in LLMs – studies what documents in the pertaining influence model outputs; by looking at the pertaining data, it tries to understand better what kind of generalization strategies LLMs use to perform reasoning tasks.