An AI model (Llama 3.1 70B) fine-tuned on the results of 60,000 people in psychology experiments shows some real promise in using LLMs for studying human behavior. It predicts actual human behavior in held-out data & it generalizes to out-of-distribution tasks and experiments.
LLMS
-
GSM1K shows test set leakage; search enables instant leaks
By
–
This (+ normal, old-fashioned leakage) is pretty much the whole argument for Scale’s semi-private benchmarks. E.g. GSM1K was made to show leakage from GSM8K test. In theory search makes all test sets insta-leak even if in practice we aren’t quite there.
-

Energy-Based Transformers: Scalable Learning and Reasoning Advances
By
–
Energy-Based Transformers are Scalable Learners and Thinkers Gladstone et al.: https://
arxiv.org/abs/2507.02092 #ArtificialIntelligence #DeepLearning #MachineLearning -

Energy-Based Transformers: Scalable Learning and Reasoning Architecture
By
–
Energy-Based Transformers are Scalable Learners and Thinkers Gladstone et al.: https://
arxiv.org/abs/2507.02092 #ArtificialIntelligence #DeepLearning #MachineLearning -
User excited about Grok upgrade and improved intelligence
By
–
@elonmusk @grok Nice upgrade! Excited to test it out—seems smarter already!
-
Goodside critiques a lucky guess about a famous illusion image
By
–
I’d call that a lucky guess, given: – You told it it was a famous image
– It took 3 sequential (vs. independent) tries
– It’s not describing the illusion well at all
– It says the image includes her “upper body” which isn’t true -

Free Deep Learning Book: Complete Guide from Math to LLMs
By
–
A fantastic, visually stunning intro to deep learning. This free book covers: – Maths for ML
– Datasets & Losses
– Linear Models
– Fully Connected Models
– CNNs for Images & Beyond
– Transformers & LLMs
– Graph Models Highly recommended! -
Building on LLMs: RAG, Fine-tuning, Personalization Essentials
By
–
How to Really Build on Top of LLMs (full training session) Personalization, databases, retrieval (RAG, CAG), frameworks, fine-tuning… We will discuss… Some theory for the essentials
LLM limitations
Context window
Knowledge issues
Embeddings + encoders
Long context
RAG
Data, -

Grok-4 Benchmark Leak Shows 45% HLE Score Doubling Gemini 2.5 Pro
By
–
Pourquoi xAI ne lance pas officiellement Grok-4 ? Certains résultats de benchmarks ont déjà fuité. Si Grok-4 a obtenu 45 % au benchmark HLE, c’est juste énorme , c'est le double du score de Gemini 2.5 Pro. Mmmmh, j'attends je voir.
-

Grok 4 Early Benchmarks Leak: 95% on AIME 2025
By
–

Les premiers benchmarks de grok 4 commencent à fuiter. AIME 2025, 95%. On est pas mal.
J'attends d'y mettre réellement la main dessus pour me prononcer