Los modelos OSS ya tienen muy buenas capacidades para resolver muchas tareas, peeeeero aún están atrás de los modelos frontera.
LLMS
-

ASCII Emoticons Can Trick LLMs Into Generating Harmful Code
By
–
Can smiley faces break AI? Researchers from Xi’an Jiaotong University, NTU, and UMass Amherst reveal a new LLM vulnerability: emoticon semantic confusion. ASCII emoticons like 🙂 can trick models into misinterpreting intent, leading to harmful code generation. Their study,
-
ChatGPT GPT-5.5 Excels at Searching Gmail Inbox
By
–
ChatGPT is now superhuman at one specific annoying thing – searching Gmail. GPT-5.5 is got a LOT better & quicker. I honestly hate searching on gmail by hand, for a search company their gmail search is surprisingly awful.
-
Domain-Specific Neurosymbolic AI Offers More Hope Than LLMs for Science
By
–
Agreed with the first part, but I think the nuanced part is what you mean by “this technology”. AI may help all of this, but pure, domain-general LLMs per se probably won’t. Domain-specific neurosymbolic hybrids offer more hope for science than chatbots do. See my October NYT
-
Diverse Simulated Personas Boost ReviewerToo Performance
By
–
Oh, and in case that's interesting: one data point supporting the value of diversity in point of view from our work on ReviewerToo is that we get the best results when we pool more simulated reviewing personas in our system
-

Codex App Ships GPT-5.5 Browser Control Sheets Slides and More
By
–
Codex for everything. Over the last 2 weeks, we shipped a pretty big set of updates: GPT-5.5, browser control, Sheets & Slides, Docs & PDFs, OS-wide dictation, auto-review mode, /pets, and a .tex plugin. The whole app also got better: dynamic UI for the task at hand, ~20%
-
Google Gemini Flash and New Omni Model Rumors Circulate
By
–
Rumors so far: – Google Gemini Flash 3.2/3.5 (already being tested)
– New Omni Model, maybe even updated Veo in competition to Seedance
– "spark Robin" – new visual model? -

Open vs Closed AI Models: Beyond Benchmark Gaps
By
–
This is a good explanation of why the gap between open and closed models is larger than it appears in benchmarks. I would add in that current open models are also more fragile than closed: they handle out-of-distribution problems far less well & have lower emergent capabilities.
-
Frontier Agent Benchmarking Struggles to Capture Real Progress
By
–
Its getting hard to benchmark frontier agent performance on longer tasks. Repeated measurement is very expensive and there are differences between using models in harnesses versus via APIs. I suspect benchmarks understate progress, they are built for models, not harnessed agents
-
New AI Research Age Emerges Beyond Scale Alone
By
–
In the last decade, I championed scale and engineering. Scale continues to be essential but it’s now time to innovate beyond scale alone. We are entering a new research age. Thanks to more compute, open-source, code and math assistants, any research team (or LLM agent) can now