With open-ended evals you mean free-form answers (and measuring conversational performance)?
LLMS
-
Verifier and LLM-as-Judge for Output Conformance
By
–
Good suggestions. I'd say those fall into the verifier category (perhaps also LLM-as-a-judge for output-conformance); or do you use something different?
-
Recent Progress in Reasoning Methods Research Synthesis
By
–
It's actually thriving. I probably have bookmarked at >100 interesting papers on reasoning-related methods in just the last few months. I will share some thoughts / synthesis of all this in the upcoming weeks.
-

GPT-5 Energy Consumption Rivals Small Nations at 16.4 TWh
By
–
GPT-5 consomme autant qu’un pays entier, il se classe 81ᵉ mondial en consommation d’électricité. Sa consommation annuelle estimée est d’environ 16,4 TWh, soit bien plus que celle de la Slovénie.
-
AI Analogy Problems: Disagreements About LLM Nature and Impact
By
–
Many of the disagreements over AI are analogy problems: is an LLM a brain or parrot? Will its effects on productivity be electrification or internet? Is the build-out of AI capability like the cloud or the 1880s railroad boom-and-bust? No one analogy fits, so we fight over them.
-
Relaxing fundamental challenges in LLM training and metalearning
By
–
But 2 and 4 can be easily relaxed. They are not fundamental challenges. For 2, note that PAQ8 LMs were trained in continual style. We could easily do the same with LLMs, but indeed what we’re doing here is accelerating evolution in the outer metalearning loop. 4 is often
-
Network Latency Negligible for LLMs on Gigabit Networks
By
–
not really, on a local network-1 gigabit, not even a 10 gigabit one-any latency is negligible in the context of llms related:
-
Local LLMs 101: Understanding GPU Processing and System Architecture
By
–
– local llms 101 – tired of guides that just tell you to run a script and call it a day?
– want to actually know what your GPU is doing, not just trust a black box?
– here's what really happens when you run a local LLM
– what gets loaded, why, and how it all fits together
– no -
Local LLMs Tutorial: Runtime Fundamentals Explained
By
–
new tutorial just dropped covering how local llms work and all the runtime fundamentals
-

Apple’s MoE Scaling Breakthrough: RoE Hyper-Parallel Inference
By
–
An intriguing paper from Apple. MoEs Are Stronger than You Think: Hyper-Parallel Inference Scaling with RoE Paper: https://
arxiv.org/abs/2509.17238