Do your language models truly grasp meaning, or are they just pretending? Researchers from Beijing University of Science and Technology and BIGAI present SemanticQA — a new benchmark that evaluates LMs' ability to handle expressions.
LLMS
-

Zai IPOs at HK$120, GLM beats DeepSeek as world’s top open model
By
–
btw Zai IPO'ed in Jan at HK$120 a share. when I first met @louszbd nobody really knew anyone using GLM's. now they have beat deepseek with the world's undisputed top open model and in some respects (see @ml_angelopoulos
) say top model period, and are returning to SF -

MiniAppBench: Benchmark for LLMs Building Interactive HTML Apps
By
–
Can your AI assistant actually build an interactive game or science tool from scratch? Researchers from Ant Group, Shanghai Jiao Tong University, and Carnegie Mellon University introduce MiniAppBench — the first benchmark to test if LLMs can generate dynamic, interactive HTML
-
hf-claude works well with GLM 5.2, install via hf extensions
By
–
hf-claude works well with glm 5.2 hf extensions install hf-claude
-

Why Model Routers Fail for Coding: Information Deficit
By
–
What is actually limiting model routers for coding tasks? Most routers treat picking a model as a static, one-off classification. This paper identifies the real bottleneck as information deficit. Simply augmenting a vanilla LLM router with task-dimension-level performance
-
Ask Claude to update its commit attribution setting
By
–
Yep, just ask Claude to update its commit attribution setting
-
Tip: Use Codex to update the global agents file
By
–
Codex tip: ask Codex to review your old PRs/sessions and update your global agents md file with your development workflow details: branch naming conventions, commit messages, attribution, test plan, and more
-
Not surprised about Grok, other models more reasonable now
By
–
Not surprised about Grok, but glad to see even other models to be much more reasonable now.
-

Stanford HAI leverages LLMs for workplace social skills training
By
–
Can we leverage AI for workplace social skills training? @StanfordHAI faculty affiliate @Diyi_Yang and her team developed a framework to help professionals improve their conflict resolution, peer counseling, and therapy skills using LLMs. https://
hai.stanford.edu/news/using-llm
s-to-improve-workplace-social-skills
…
