.@josh_bickett (our eng. lead) co-wrote this piece w/ OpenAI on this topic, if you want to learn more about our process: https://
cookbook.openai.com/examples/strip
e_model_eval/selecting_a_model_based_on_stripe_conversion
…
GENERATIVE AI
-
Stripe and OpenAI collaborate on AI model evaluation
By
–
-

Gemini Visualizes Boullée’s Never-Built Newton Centograph
By
–
Never built architecture and AI. Gemini image generator (nano banana) does a pretty good job imaging what Boullée’s Centograph, his fantastical (and never built) tomb for Isaac Newton would have looked like. I gave it the original 1784 black and white drawings to work with.
-

Groq Launches Kimi-K2-Instruct-0905 High-Speed AI Model
By
–
Who needs sleep?
Kimi-K2-Instruct-0905 just landed. 200+ T/s, $1.50/M tokens.
256k context window.
Built for coding. Rivals Sonnet 4. Available now. -
Eval Performance Doesn’t Predict Real-World User Adoption
By
–
Not necessarily. I thought this way for a while, but at this point I've seen so many "wow, this did incredible on an eval, let's put it in prod! oh, shit, our users hate this"… + "this scored terribly, but whatever, let's give it to 1% of users, oh my god, it converts so well!"
-
Multimodal LLMs struggle with figurative language in image generation
By
–
LLM multimodal image generation remains too literal in the face of figurative language.
-
AI-Generated Photorealistic Portrait of Japanese Ceramicist
By
–
Here's the full prompt that was used to generate the first image: A photorealistic close-up portrait of an elderly Japanese ceramicist with deep, sun-etched wrinkles and a warm, knowing smile. He is carefully
inspecting a freshly glazed tea bowl. The setting is his rustic, -
Beyond Evals: A/B Testing Improves Customer Satisfaction
By
–
*Need* is a strong word. It's not a perfect 1:1. We actually no longer use ANY evals, purely A/B testing, and we've only improved our customer satisfaction (measured in retention/LTV)
-
MiniCPM-V 4.5 Outperforms Larger Models Through Efficient Design
By
–
MiniCPM-V 4.5 beats much larger models by focusing on efficiency instead of brute force. Instead, it combines (1) extreme video compression, (2) smarter OCR/document pretraining, (3) controllable fast/deep reasoning, and (4) efficient alignment. So an 8B model can outperform
-
Limited Control and Understanding in Model Fine-tuning Approaches
By
–
Maybe it's a n-of-1 problem, but there's very little control (hyperparams etc.), and very little understanding (both on HOW it's being done (is it LoRA, fft, etc.) and telemetry-wise). May be n-of-1 b/c most people who want this stuff will run their own shit, but I personally
-
Free Step-by-Step AI Agents and RAG Systems Tutorials
By
–
100+ free step-by-step tutorials with code covering: AI Agents RAG Systems Voice AI Agents MCP AI Agents Multi-agent Teams Autonomous Game Playing Agents P.S: Don't forget to subscribe for FREE to access future tutorials.