There has been a push to use OpenEvidence AI for doctors. But this paper suggests general models are much better: “Frontier LLMs outperformed clinical AI tools in all three evaluations. Clinical AI tools performed comparably to auto-enabled Google Search AI Overview on the RCQ.”
GENERATIVE AI
-

65% of US physicians use OpenEvidence with 27 million prompts in April
By
–
>65% of US physicians use OpenEvidence, with 27 million prompts in April https://
nbcnews.com/tech/tech-news
/openevidence-ai-doctor-medical-physician-login-app-what-npi-uptodate-rcna341064
… -

General AI models outperform specialized medical sources in study
By
–
For medical information, general AI frontier models (Google, OpenAI, Anthropic) outperformed specialized @EvidenceOpen and @UpToDate as assessed by 12 US clinicians, randomized and blinded to which model and extensive testing/benchmarks. This was not anticipated.
-

GPT 5.6 likely drops June 23, potentially harming Anthropic
By
–
I’m now 99.99% certain that GPT 5.6 will drop, in Codex and everywhere else, on June 23rd. This could end up being Anthropic’s biggest self own.
-
Fable’s lack of native imagegen limits multimodal output for presentations
By
–
Not having access to native imagegen does hold Fable back somewhat. It is really good at making PNGs, etc, but there are lots of areas (including commercially valuable ones like presentations) where having the ability to have multimodal output would be helpful/token efficient.
-
Toolkits for AIs to build games focusing on gameplay loops
By
–
Are there toolkits (or skillsets) being created specifically for AIs to use for building games? They default to 3js, reinvent how to make sprites from scratch each time, test technical issues but not gameplay loops, etc. It would help to point AIs at some tools to focus them.
-
Techniques for randomness and diversity in language model outputs
By
–
Summary of things: – turn up randomness and ban the most likely words (temperature + min-p + XTC sampling)
– ask for several different options at once, seed each with random constraints (verbalized sampling + entropy injection)
– give it a memory of what it's said and pick the -
Experimentation with Gemma 4 for creative fashion prompts
By
–
I experimented with using modifications of Gemma 4 to create creative prompts repeatedly. There are still a few quirks, but these are all results from the same simple query: "a dynamic fashion photo of a woman".
-
The author admits that ChatGPT and Claude are much smarter than him
By
–
ChatGPT and Claude are MUCH MUCH smarter than me in ALL areas. I am very amazed by the denial. People do not admit that AI is smarter than them. I accept the truth.

