this guy literally created a Kebab benchmark for LLMs ♨️♨️♨️ pic.twitter.com/9vY9zLdqBe
— Charly Wargnier (@DataChaz) 25 juin 2026
this guy literally created a Kebab benchmark for LLMs
By
–
this guy literally created a Kebab benchmark for LLMs ♨️♨️♨️ pic.twitter.com/9vY9zLdqBe
— Charly Wargnier (@DataChaz) 25 juin 2026
this guy literally created a Kebab benchmark for LLMs

By
–
Our paper, “Rethinking Psychometric Evaluation of LLMs: When and Why Self-Reports Predict Behavior,” has been selected for Oral Presentation at CTB @icmlconf * Paper: https://
arxiv.org/abs/2606.12730
* Website: https://
psychology-of-ai.github.io/psychometric-e
valuation.html
… * Code: https://
github.com/psychology-of-
AI/Rethinking-Pyschometric-Eval-LLMs
… A central question
By
–
Exactly this. You need to be able to switch between models and have the same connectivity to your data and your agent's memory. We're building this, love to share more.

By
–
This is good news, but the fact that Composer 2.5 comes out on top over GLM 5.2 here should trigger a serious re-evaluation of the entire benchmark. It's nowhere near in practice except basic tasks… and that undermines trust in Cursor's entire model evaluation / publicity.
By
–
If AI is doing the writing, it should do the reading too.

By
–
"Building Business-Ready Generative #AI Systems — Build Human-Centered Generative AI Systems with Context-Aware Agents, Memory, and LLMs for the Enterprise" at https://
amzn.to/3Jdcio5 v/ @PacktDataML Learn:
Implement an AI controller with a conversation AI agent and
By
–
The real signal is not "AI does law." It's the target: in-house teams and smaller firms priced out of Harvey- or Lexis-tier tools. Citation-grounded legal AI is moving down-market. That's the trend worth watching.
By
–
Lawyers' biggest AI fear? Citing a case that doesn't exist.
— Futurepedia – Learn to Leverage AI (@futurepedia_io) 25 juin 2026
Two years of peers getting sanctioned for AI-hallucinated citations made "just trust the model" a non-starter in legal. @perplexity_ai just shipped something built around that exact fear. 🧵 pic.twitter.com/RpEcMQi1yS
Lawyers' biggest AI fear? Citing a case that doesn't exist. Two years of peers getting sanctioned for AI-hallucinated citations made "just trust the model" a non-starter in legal. @perplexity_ai just shipped something built around that exact fear.
By
–
This asynchrony really should exist in normal one-on-one chatbot interactions. When you pose ask a hard question to an LLM it should say, immediately, “This is a hard problem — give me 15m.” If that’s an issue, you should be able to say so and get a quicker guess.

By
–
AI Agents in Action: http://
amzn.to/44kbHsL v/ @ManningBooks Gain hands-on experience in these areas: Understand & implement AI agent behavior patterns Design & deploy production-ready intelligent agents Leverage the OpenAI Assistants API and complementary tools