The Anthropic Fellows program provides funding and mentorship for a small cohort of AI safety researchers. Here are four exciting papers that our Fellows have recently released.
SAFETY
-

Language Models are Injective and Invertible with SIPIT
By
–
Hottest paper on AlphaXiv Language Models are Injective and Hence Invertible Every prompt maps to a unique hidden state and can be exactly reconstructed with this paper’s algorithm SIPIT. This means the model’s internal activations are the full prompt in disguise!!
-
Codex AI Catches Real Bugs in Code Review Process
By
–
Just had an excellent experience where Codex code review caught two real bugs which would have been easy for human reviewers to miss. It's a comforting, and very novel, feeling to have such a strong safety net watching over every PR.
-
Heed warnings from those declaring plans of extermination
By
–
History says, pay attention to people who declare a plan to exterminate you — even if you're skeptical about their timescales for their Great Deed. (Though they're not *always* asstalking about timing, either.)
-

Stress-Testing LLM Constitutional Specifications and Behavioral Guidelines
By
–
7. Stress-Testing Model Specs This research examines how well large language models adhere to their stated behavioral guidelines by stress-testing AI constitutional specifications through value-tradeoff scenarios.
-

LLMs Show Limited Introspective Awareness Capabilities
By
–
2. Introspective Awareness Anthropic research demonstrates that contemporary LLMs possess limited but functional introspective capabilities, the ability to recognize and accurately report on their own internal states.
-
Open-source AI transparency crisis: Chinese base models audit challenges
By
–
state of open-source AI in 2025:
– almost all new open American models are finetuned Chinese base models
– we don’t know the base models’ training data
– we have no idea how to audit or “decompile” base models who knows what could be hidden in the weights of DeepSeek -
Investigation into Codex Model Degradation Issues
By
–
an excellent investigation into reported Codex degradations, super interesting read imo:
-
Seven AI Opportunities to Transform the World Positively
By
–
7 Great AI Hopes That Could Change The World Beyond caution and risk, there are powerful, positive possibilities for AI—from improving lives to enabling new forms of collaboration and creativity. Read more https://
bernardmarr.com/7-great-ai-hop
es-that-could-change-the-world/
… #AIforGood #FutureTech #PossibilityMindset -
Agreement on terminology for anti-ASI coalition
By
–
Cool, I'm fine with "anti-ASI condition" or "anti-extinction coalition".
