as the sun sets on GPT 4.5, just reflecting that it was really not that great for summarization tasks. today's @smol_ai example here i think the "failure"* of GPT 4.5 relative to o3/o4mini is actually fantastic validation of the bet on the reasoning paradigm. 10-100x larger
@swyx
-

First Podcast Feature with Yi Tay and Machine Learning Experts
By
–
I'm back! and super proud to be the first podcast to feature @YiTayML senpai with special shoutouts to @quocleix
, @_jasonwei
, @hwchung27
, @teortaxesTex
! Special interest callouts:
– why encoder-decoder is not actually that different than decoder-only cc @teortaxesTex @eugeneyan -
Merging on Test Set as Key Machine Learning Approach
By
–
~~training on~~ merging on the test set is all you need
-
Greens, reds, and merge hierarchy decision explanation
By
–
what are the greens and the reds? and how was merge hierarchy decided?
-
Latent Space Podcast: LLM Leaderboard Evaluations Discussion with Clémentine Fourrier
By
–
we -just- did a @latentspacepod with @clefourrier on what evals she is looking for to add to the LLM Leaderboard and she had 3 back to back banger answers!!
-
Token Ratio Parameter and Cost-Performance Tradeoff in LLM Pricing
By
–
fwiw i think its better to have one parameter that blends input:output token ratio (eg, 3:1 is pretty common), and compute blend cost that way. reserve the screen real estate for the obvious issue here with the cost, which is %MMLU or blend benchmark score tradeoff
-
Fireworks AI Recruiting Top DevRel Talent in SF and NYC
By
–
SF/NYC Devrel folks, you should seriously consider joining @lqiao at @FireworksAI_HQ . one of the absolute top ai infra teams right now https://
jobs.ashbyhq.com/fireworks.ai/5
7ec407c-e7e1-48c0-9858-a9a434d9e4fe
…? -
AI Engineering Conference Recommendation and Modal Speculation
By
–
I know an ai engineering conference that would be perfect.. Srsly tho idk what could have changed, this all I imagine is p core to modal
-
Routing Objectives and Modular Expert Systems in MoEs
By
–
i suspect this is a contingent but not necessary truth. if we refocused routing objective to be topic experts, or sparse-upscaled topic experts, then we can also achieve plug and play experts or frankenmerged MoEs
-
10 Essential Facts About Anthropic Claude
By
–
THIS CHANGES EVERYTHING 10 things you need to know about ANTHROPIC CLAUDE
