First large-scale study of physician-supervised LLM-based medical conversational agent • 926 eligible cases, 298 complete patient interactions over 3 weeks • Integration into existing medical chat service at Alan
LLMS
-

First impressions on updates to Claude and GPT-4o
By
–
Co-authored a new blog post — First impressions on the writing and personality of the new updates to Claude and GPT-4o from Scale Prompt Engineer @JeremyKritz Because no benchmark is complete without vibes!
-
Robot Data Scale Compared to Language Models: A Significant Gap
By
–
Maybe, but I wouldn’t use the word “soon”: The data in Llama 3 is approx 80,000 times greater than the data collected by π in a year. And compared with vision and language date, robot data is continuous and high dimensional.
-
Critique of Article on Foundation Models and Robot Learning
By
–
Interesting article but the author drank the Kool-Aid and never sought out other viewpoints: “Foundation models like GPT-4 have largely subsumed [previous methods] and dexterity will probably soon be subsumed, too.” A Revolution in How Robots Learn.
-

SmolVLM: Fast 2B Parameters Vision-Language Model On-Device
By
–
when the model is so fast and small it falls "outside" of the pareto curve SmolVLM: a fast and high performance 2B params Vision-Language Model for on-device/in-browser use
-
Disponibilité des MCP connectors pour les développeurs
By
–
Here is the current statement regarding availability. Not sure why Pro plans are not mentioned there > Developers can start building and testing MCP connectors today. Existing Claude for Work customers can begin testing MCP servers locally, connecting Claude to internal
-
LLMs Better Than Humans at Content Evaluation Tasks
By
–
Yes, I think it's harder for humans to evaluate, but probably not for LLMs
-
RAG vs Long Context Models: Is Retrieval-Augmented Generation Dead?
By
–
RAG vs. Long Context Models: Is Retrieval-Augmented Generation Dead? 3/10 of the "From Beginner to Advanced LLM Developer" course by Towards AI! Follow the channel to see the others 😉 Learn more in the video: https://
youtu.be/qN3vhWlzd4A #rag #llms #llm #gemini -

Discussion with codebases on Gemini mobile
By
–


Chatting with codebases on Gemini also seems to be supported on mobile apps. There you can observe "thinking" statuses while you continue conversation with existing chats. It tries to open uploaded folders in web view which seems broken currently
-

Chatbot Arena Evaluation Flaw: Messages vs Problem-Solving
By
–
I think the main issue with the chatbot arena is that it evaluates messages instead of conversations. Users interact with LLMs to solve specific problems. Ideally, evaluations should capture how well this problem is solved. → When o1 takes 2 minutes to give me an answer