and it truly is a tortured model. here the model hallucinates a programming problem about dominos and attempts to solve it, spending over 30,000 tokens in the process completely unprompted, the model generated and tried to solve this domino problem over 5,000 separate times
@jxmnop
-

Embedded Generations: AI Model Capabilities in Math and Code
By
–
here's a map of the embedded generations the model loves math and code. i prompt with nothing and yet it always reasons. it just talks about math and code, and mostly in English math – probability, ML, PDEs, topology, diffeq
code – agentic software, competitive programming, -

GPT-OSS Training Data Analysis: Bizarre Results Revealed
By
–
curious about the training data of OpenAI's new gpt-oss models? i was too. so i generated 10M examples from gpt-oss-20b, ran some analysis, and the results were… pretty bizarre time for a deep dive
-
GPT-5 Scaling Laws: Diminishing Returns on General Intelligence
By
–
shortest explanation of GPT-5: this is exactly what the scaling laws predicted! the model is better, the returns are diminishing, and sadly absolute general intelligence improvements will only get smaller the good news is there’s so much still to do. personality, reasoning,
-
Four Years Minimum Wage Career Path Challenges
By
–
no, it requires four years of minimum wage and perpetual headache
-
PhD holders unlikely to produce flawed data visualizations
By
–
if they have a phd then there’s no way they would’ve made this graph after two phds
-

PhD-level rigor in data visualization and research standards
By
–
people arent gonna wanna hear this but i truly do not believe this mistake could’ve been made by someone with a phd. after going through brutal peer review several times you just stop doing stuff like this. whoever made this graph clearly has a bachelors degree. maybe a masters
-
Python and SWE Bench: Comparing Claude and ChatGPT Performance
By
–
python is great also i love swe bench btw just might not be the highest signal point of comparison between Claude and chatGPT these days
-
SWEbench Scores Inflated by Django Training Data Bias
By
–
good time to remind everyone that a high score on SWEbench really just indicates the training data contained a sufficiently large proportion of Django
