Updated Perplexity Discover will roll out next week. iOS to begin with. Should be pretty good.
GENERATIVE AI
-
Black Forest Labs FLUX model licensing uncertainty clarified
By
–
Black forest labs engineer says its their FLUX model. Maybe they licensed different ones. Not sure.
-
Salesforce Enforces Trusted URL Allowlists for Agentforce
By
–
And details of the Salesforce fix: "Starting September 8, 2025, Salesforce will begin enforcement of Trusted URL allowlists for Agentforce and Einstein Generative AI agents" https://
help.salesforce.com/s/articleView?
id=005135034&type=1
… -
OpenAI ChatGPT Pulse and Luma Ray3: Major AI Launches
By
–
Tomorrow’s AI News video is a big one for multiple reasons (including a fun announcement). Here are all the highlights from the past 2 weeks, what did I miss? – @OpenAI launched ChatGPT Pulse
– OpenAI unveiled age prediction + parental controls
– @LumaLabsAI released Ray3
– -
Parsing and Evaluation: Technical Foundations for AI Systems
By
–
The bottom line: Parsing isn't just a technical detail. It's part of the evaluation story. Don’t forget to follow us @snorkelai and @realjustinbauer – and if you have questions about evals, RL environments, and expert-developed datasets, talk to us!
-

Structured Formats Limit AI Reasoning Capabilities
By
–
Structured formats can actually constrain reasoning Models like GPT-4.1 and Grok-3 performed worse when forced into rigid JSON structures. The formatting requirements limited their ability to think through complex problems.
-
Reasoning Models Show Resilience Across Parsing Methods
By
–
Reasoning-first models stayed resilient Claude Sonnet 4, Gemini 2.5 Pro, and o4-mini showed minimal sensitivity to parsing methods—their strong reasoning held steady across formats.
-
OpenAI Deep Research Tool Transparency and System Prompt Access
By
–
I would use OpenAI Deep Research (and equivalent products from other labs) a whole lot more if I'd seen the full list of tools that are available to them – and ideally their system prompts as well but that's less valuable to me than the tool definitions
-

Medical AI Benchmarks: Beyond Memorization to Real-World Performance
By
–
Another paper pointing out the inadequacy of older public benchmarks for determining whether AI is actually good for tasks like medicine. Models are clearly memorizing or using heuristics for some answers. A new wave of benchmarks based on real-world data will help, more needed.
-
OpenAI Models Solve Programming Challenges With Code Sandbox
By
–
Confirmation from Ahmed El-Kishky – who worked on this project at OpenAI – that their models solved the programming challenges with a code execution sandbox but no internet access