Enterprise adoption will take awhile, but we have moved past no success – 75% of companies had positive ROI from GenAI in November, and coding agents are now starting to have a real impact at many firms. I think normal adoption does not mean that model abilities aren't growing.
AGENTS
-

Mix-Quant: Quantized Prefilling and Decoding for Agentic LLMs
By
–
Mix-Quant Quantized Prefilling, Precise Decoding for Agentic LLMs
-

Auth Proxy: Controlling Agent Behavior Boundaries
By
–
Introducing the sandbox Auth Proxy: A way to control the boundary between agent-generated behavior and the rest of the world. An explainer from @hwchase17
-

Tips for AI-assisted scientific coding workflows
By
–
How to vibe code in science: early adopters share their tips
by Nicola Jones @Nature Learn more: https://
bit.ly/4uOUTo1 #Coding #AI #GenerativeAI #ArtificialIntelligence #MachineLearning -
Optimize Your AI Agents with the Right Context
By
–
This applies to any AI coding agent, not just Codex. Cursor, Claude Code, and Antigravity all follow the same approach. Stop asking it to write code from scratch. Provide it with the elements that help you focus on the problem you actually want to solve.
-
Codex enhances engineers’ focus
By
–
6/ The thread beneath all five. Codex isn’t used to write code faster. It’s used to handle the work that disrupts focus. Reading. Refactoring. Testing. Scaffolding. Background tasks.
The engineer stays in the flow. The agent takes care of the rest. -

NVIDIA Verified Agent Skills for AI Agent Capability Governance
By
–
Technical Deep Dive https://
developer.nvidia.com/blog/nvidia-ve
rified-agent-skills-provide-capability-governance-for-ai-agents/
… -

NVIDIA Announces Verified Agent Skills for Enhanced AI Agent Security
By
–
We just shipped NVIDIA-Verified Agent Skills Skills make your agent more capable, but can also introduce vulnerabilities. Verified skills give you transparency into what a skill does, where it came from, what risks it carries, and whether it's been modified. Every verified
-

Evaluating Memory in Long-Horizon AI Agent Systems
By
–
LongMINT Evaluating Memory under Multi-Target Interference in Long-Horizon Agent Systems
-

Composer 2.5 scores 62, price difference makes ‘slightly better’ not worth 60x
By
–

Composer 2.5 scores 62 on the Artificial Analysis Coding Agent Index. The two models above it score 65 and 66. The price difference: $0.07 per task vs. $4–5. At some point "slightly better" stops being worth "60x more expensive," and most engineering teams crossed that point a