My first experiences with Codex are that it makes the same mistakes other models do when one-shotting something in Cursor. It also failed to respond to feedback and hallucinated a bunch of stuff. Ended up fixing PR in Cursor. OpenAI also re-implemented a bunch of Github stuff
GENERATIVE AI
-
Gemini App Uses Ensemble of Models for Query Responses
By
–
No, Gemini app uses an ensemble of models to answer queries
-
Negative Understanding of LLMs: A Gary Marcus Critique
By
–
to garymarcus harder than the man himself, you had to acquire not zero but negative understanding of LLMs; respect
-
AGI Alpha: Unprecedented $15 Quadrillion Market Opportunity
By
–
"AGI Alpha" : The Greatest Alpha Opportunity Ever Predicting the sectors AGI will revolutionize first is the key to unlocking a historic, unprecedented opportunity—capturing a share of a projected $15 Quadrillion shift. Join the frontier: http://
github.com/MontrealAI/AGI
-Alpha-Agent-v0
… -
Kling v1.6 Pro struggles with multi-character consistency
By
–
Pretty good coherency here, but I used 4 different characters as elements in the generation and only one character persisted, the other 3 aren't represented (kling v1.6 pro).
— fofr (@fofrAI) 18 mai 2025
When I've tried two characters it's worked as expected.
Prompt: "group hug" pic.twitter.com/AeZfCZpkpjPretty good coherency here, but I used 4 different characters as elements in the generation and only one character persisted, the other 3 aren't represented (kling v1.6 pro). When I've tried two characters it's worked as expected. Prompt: "group hug"
-

CellVerse: LLM Benchmark for Single-Cell Biology Tasks
By
–
10. CellVerse Introduces a benchmark to evaluate LLMs on single-cell biology tasks by converting multi-omics data into natural language.
-

RL Framework Teaches LLMs Efficient Search Tool Usage
By
–
7. RL for Search-Efficient LLMs Proposes a new RL-based framework (SEM) that explicitly teaches LLMs when to invoke search and when to rely on internal knowledge, aiming to reduce redundant tool use while maintaining answer accuracy.
-

Nemotron-Research-Tool-N1: LLM Tool-Using with Rule-Based RL
By
–
6. Nemotron-Research-Tool-N1 Introduces Tool-N1, a family of tool-using LLMs trained using a rule-based reinforcement learning (R1-style RL) approach, without reliance on supervised reasoning trajectories.
-

AM-Thinking-v1: 32B Open-Source Model Rivaling Larger MoE Systems
By
–
4. AM-Thinking-v1 Introduces a dense, open-source 32B language model that achieves state-of-the-art performance in reasoning tasks, rivaling significantly larger Mixture-of-Experts (MoE) models.
-

RL Improves LLM Mathematical Reasoning with Single Example
By
–
3. RL for Reasoning in LLMs with One Training Example This paper shows that Reinforcement Learning with Verifiable Rewards (RLVR) can significantly improve mathematical reasoning in LLMs even when trained with just a single example.