Enterprise AI just got a new benchmark. http://
H2O.ai’s h2oGPTe is #1 on the GAIA benchmark—again—with 79.7% accuracy, outperforming Google & Microsoft by 30+ points. GAIA tests what really matters: reasoning, research, docs, decisions.
Not just chat—real work.
LLMS
-

h2oGPTe Tops GAIA Benchmark with 79.7% Accuracy
By
–
-

OpenRouter Raises Funding Amid Explosive Growth
By
–
OpenRouter is likely the fastest-growing company I've ever invested in. Absolutely bonkers growth. I'm also an avid user of the product — if you're not using them, you should be. Congrats on the raise!!
-
LLM Reviewers Raise Questions About Evaluation Authenticity
By
–
unfortunately many of my fellow reviewers are LLMs as well
-

LLMs Beyond Hype: Profit Tools for Software Efficiency
By
–
Here are 10 key points to remember from our free AI 2-hour video training if you missed it : 1. LLMs: Beyond the Hype. More than buzzwords, Large Language Models (LLMs) are concrete tools for generating profit, saving money, and boosting efficiency—especially in software
-
O3 Pro Prompt Compatibility with Non-Reasoning Models
By
–
Most o3 prompts will work great for o3 pro! However, this prompt is for non-reasoning models.
-
NeurIPS Review Experience: LLM-Generated Papers Quality Issues
By
–
i just reviewed five papers for NeurIPS and it was an awful experience: – first paper was clearly LLM-generated. it was too short, the references didn't work, had no experiments or theory at all, and a ton of obvious mistakes. the more i read the less it made sense
– two were -
Strong Regularization Prevents RLHF Model Degradation
By
–
May your regularizer be strong, lest you RLHF to slop.
-
Distinguishing between simple prompts and complex application contexts
By
–
Haha I'm not trying to coin a new word or something. I just think people's use of "prompt" tends to (incorrectly) trivialize a rather complex component. You prompt an LLM to tell you why the sky is blue. But apps build contexts (meticulously) for LLMs to solve their custom tasks.
-
Context Engineering Over Prompt Engineering in Industrial LLM Applications
By
–
+1 for "context engineering" over "prompt engineering". People associate prompts with short task descriptions you'd give an LLM in your day-to-day use. When in every industrial-strength LLM app, context engineering is the delicate art and science of filling the context window
-
AI makes litigation asymmetric via ChatGPT legal threats
By
–
AI will also make litigation asymmetric. One person with ChatGPT and a grudge will file more legal threats than a firm used to in a year. And most targets will settle because AI also knows your cost of fighting.