.@josh_bickett (our eng. lead) co-wrote this piece w/ OpenAI on this topic, if you want to learn more about our process: https://
cookbook.openai.com/examples/strip
e_model_eval/selecting_a_model_based_on_stripe_conversion
…
@mattshumer_
-
Stripe and OpenAI collaborate on AI model evaluation
By
–
-
Eval Performance Doesn’t Predict Real-World User Adoption
By
–
Not necessarily. I thought this way for a while, but at this point I've seen so many "wow, this did incredible on an eval, let's put it in prod! oh, shit, our users hate this"… + "this scored terribly, but whatever, let's give it to 1% of users, oh my god, it converts so well!"
-
Beyond Evals: A/B Testing Improves Customer Satisfaction
By
–
*Need* is a strong word. It's not a perfect 1:1. We actually no longer use ANY evals, purely A/B testing, and we've only improved our customer satisfaction (measured in retention/LTV)
-
Limited Control and Understanding in Model Fine-tuning Approaches
By
–
Maybe it's a n-of-1 problem, but there's very little control (hyperparams etc.), and very little understanding (both on HOW it's being done (is it LoRA, fft, etc.) and telemetry-wise). May be n-of-1 b/c most people who want this stuff will run their own shit, but I personally
-
Evals vs A/B Tests: Measuring What Actually Matters
By
–
Evals measure what DEVELOPERS care about. A/B tests measure what USERS care about. Both have their place, but which do you think is the one that actually matters?
-
User-Centric Metrics: Why A/B Testing Succeeds
By
–
A/B tests work because they measure success on what USERS really care about, rather than what DEVELOPERS care about.
-
Evals Disconnect From Real User Utility In AI
By
–
Often, evals are very disconnected from actual utility. For example, we had an eval for a while that measured 'writing style'. Basically, how well do we prevent AI slop in writing output? We maxed out the eval, put the model in prod, and users hated it.
-
Real-world A/B testing better than evals for AI utility
By
–
This is the correct take. Evals are helpful but not well-correlated with actual utility. At Otherside, we use A/B tests grounded in real-world traffic, measured against subscriptions and retention. We've tried it all. This is the way.
-

OpenPipe Acquisition by CoreWeave: Major AI Infrastructure Exit
By
–
Huge congrats to @corbtt @dvdcrbt and the @OpenPipeAI team on their acquisition by @CoreWeave
! Kyle + David are special founders and are going to do great things at CoreWeave. This is my first exit as an investor! Very exciting, but bittersweet, as I loved working with them. -

OpenPipe AI Team Achieves Remarkable Results
By
–
Another incredible result from the @OpenPipeAI team.