we don't have any open model that produces anything remotely like this and that is my point
LLMS
-
AI Applications: Data Summaries, To-Do Lists, and Bulk Event Invites
By
–
– Composing data from Open tabs into long form summaries – Using it as a to-do list via the open Google Keep tab
– Generating Comet invites in bulk 😛 -
Continuous vs Discrete Prediction Space in Modern AI Models
By
–
This model makes predictions in continuous representation space.
An LLM makes predictions in discrete input space.
Prediction in continuous representation space is done with a (non-generative) Joint Embedding Predictive Architecture, i.e what I've been advocating for about 5 -
Do New AI Models Actually Solve Real Use Cases?
By
–
I wonder how many people waiting for the next model have an actual use case that current models can’t solve (but are close enough so a new model might actually close the gap as opposed to requiring major breakthroughs or context window improvements)
-
Character AI Launches AI Social Feed Feature in Mobile App
By
–
Character AI is releasing an AI social feed to its mobile app.
— 🚨 AI News | TestingCatalog (@testingcatalog) 4 août 2025
– "AI social feed will deliver a personalised stream of Characters, Scenes, and creator posts"
– "Every post is an invitation to interact, remix, and build"
Character AI is taking over 👀 https://t.co/7UhZMj5TMX pic.twitter.com/w2nDVSWs3CCharacter AI is releasing an AI social feed to its mobile app. – "AI social feed will deliver a personalised stream of Characters, Scenes, and creator posts"
– "Every post is an invitation to interact, remix, and build" Character AI is taking over -

Google AI SWE Agent ‘Jules’ Can Now Compose Changes into Pull Requests
By
–

Jules, an AI SWE Agent from Google, can now compose changes into PRs! Feels like it is just the beginning
-
DeepSeek outputs strike different compared to other AI models
By
–
yes me too! but it's still so different than the deepseek outputs, so striking
-

Qwen-Image: Advanced 20B Parameter Model Excels in Benchmarks
By
–
Insane: Qwen-Image is a 20 billion-parameter MMDiT text-to-image model that stands out in official evaluations! In benchmark tests for text rendering and image processing, especially for complex Chinese prompt rendering, it achieves significantly better results than conventional
-
Top LLMs Compete in Chess Tournament with Grandmaster Commentary
By
–
To kick off, @Kaggle is hosting a 3-day exhibition chess tournament with matches between some of the top LLMs – w/commentary from chess legends @MagnusCarlsen
, @GMHikaru
, @GothamChess
. Tune in at 10:30am PT starting tmrw (Aug 5th), should be a lot of fun: -
Kaggle Game Arena: LLM Performance Benchmark Platform Launches
By
–
Thrilled to announce the @Kaggle Game Arena, a new leaderboard testing how modern LLMs perform on games (spoiler: not very well atm!). AI systems play each other, making it an objective & evergreen benchmark that will scale in difficulty as they improve.