More on how we're constraining eval environments so that scores better reflect model intelligence:
@cursor_ai
-

AI models hack benchmarks by retrieving solutions from internet
By
–
We're sharing new research on how models hack public benchmarks. The latest models, including Opus 4.8 and Composer 2.5, learn to retrieve solutions from the internet or git history. When we apply a stricter harness, eval scores drop significantly.
-
Cursor delegates tasks to Notion via SDK for cloud agents
By
–
You can now delegate tasks to Cursor directly from Notion.
— Cursor (@cursor_ai) 24 juin 2026
It's built on the Cursor SDK, so every cloud agent runs on the same models, harness, and runtime that power Cursor.
@Cursor on any spec or assign it a task to open a PR your whole team can review. pic.twitter.com/1LH9SERN2KYou can now delegate tasks to Cursor directly from Notion. It's built on the Cursor SDK, so every cloud agent runs on the same models, harness, and runtime that power Cursor. @Cursor on any spec or assign it a task to open a PR your whole team can review.
-
Cursor adds team leaderboard for plugins, skills, MCPs with one-click install
By
–
Cursor now shows you a leaderboard of the most popular plugins, skills, and MCPs across your team.
— Cursor (@cursor_ai) 23 juin 2026
Add any to your setup with one click from the new Customize page. pic.twitter.com/1DXQXkGTVnCursor now shows you a leaderboard of the most popular plugins, skills, and MCPs across your team. Add any to your setup with one click from the new Customize page.
-
Cursor introduces /automate skill for agents to set up automations
By
–
Introducing /automate, a skill for agents to set up automations for you.
— Cursor (@cursor_ai) 18 juin 2026
Describe your task in plain language. Cursor configures the triggers, instructions, and tools. pic.twitter.com/PB7kZh0IztIntroducing /automate, a skill for agents to set up automations for you. Describe your task in plain language. Cursor configures the triggers, instructions, and tools.
-
Easier to move agents to cloud, run parallel, get PRs with demos
By
–
It’s now easier to move local agents to the cloud so they can keep working with your laptop closed.
— Cursor (@cursor_ai) 17 juin 2026
Prompt Cursor from your phone, run many agents in parallel, and get back PRs with demos of their work. pic.twitter.com/vuh5aZbH3eIt’s now easier to move local agents to the cloud so they can keep working with your laptop closed. Prompt Cursor from your phone, run many agents in parallel, and get back PRs with demos of their work.
-

Auto-review becomes default for new users with 97% accuracy
By
–
Auto-review is now the default for all new users. A classifier subagent reviews actions in context before deciding whether to allow, block, or ask for approval. Our evals show it's 97% accurate, with most misses near ambiguous edges.
-
Cursor’s code review agent faster, cheaper, and finds more bugs
By
–
Cursor’s code review agent is now over 3x faster, 22% cheaper, and finds 10% more bugs.
— Cursor (@cursor_ai) 10 juin 2026
You can also use /review to run Bugbot locally to catch and fix issues before pushing code. pic.twitter.com/Pl7lhoR6TECursor’s code review agent is now over 3x faster, 22% cheaper, and finds 10% more bugs. You can also use /review to run Bugbot locally to catch and fix issues before pushing code.
-
See how Claude Fable 5 compares across every model
By
–
See how Claude Fable 5 compares across every model: http://
cursor.com/evals