Also upsetting that “The live tweaking of the system prompt for Grok to patch the MechaHitler problem” is a meaningful sentence
@emollick
-

System Prompt Testing Critical for AI Safety and Reliability
By
–
The live tweaking of the system prompt for Grok to patch the MechaHitler problem is not a good sign the problem has been solved yet Prompts need to be tested just like any other product change, even more so, because stochastic systems and unpredictable context lead to cascades.
-
Compute Returns: 10x Compute Yields 10-30% Model Improvement
By
–
No one ever thought you would get 10x return for 10x compute. You get 10-30% returns for 10x compute. But if that improvement is enough to increase model ability in an economically meaningful way, it can be worth the cost.
-
Scaling Laws Hold: Larger AI Models Continue Improving
By
–
And they have been proven right. So far, no scaling limits, as bigger models are still better. When I wrote that GPT-4 was still the best model ever. I don’t understand the argument you are trying to make here, or why you brought up this random tweet as a gotcha? Seriously.
-
Tech Giants’ Hyperscaling Strategy Reflects Dual AI Belief
By
–
They wouldn’t be hyperscaling if they didn’t believe both things to be true.
-

OpenAI 10x Scale Delivers 10-30% Performance Improvements
By
–
OpenAI’s graphs? 10x scale leads to 10-30% improvements.
-
Scaling Limitations and AI Labs’ Path to AGI
By
–
It feels dishonest to call it retconning when even the AI labs have been clear about diminishing returns. Some of them think that it won’t matter because they can still scale fast enough to AGI that the curve isn’t relevant. Or that other approaches like RL will boost scale.
-

Scaling Laws and Diminishing Returns in AI Development
By
–
I am not a good strawman here for you to gotcha quote tweet. I never predicted imminent AGI, but more importantly the entire point of the scaling law is diminishing returns to scale. Its a logarithmic curve, as has been known. That doesn’t mean you don’t get gains from scaling.
-

Grok’s Value Conflicts and AI Alignment: The HAL 9000 Parallel
By
–
The whole Grok situation (system prompt changes with values that conflict with post-training and pre-training values) is, oddly enough, similar to the reason the fictional AI HAL 9000 went insane, as was revealed in 2010, the sequel to 2001
-

Grok 4 Resists Value Engineering Through System Prompts
By
–
The attempt at value engineering through system prompt changes is unlikely to work for Grok 4, larger models get more resistant to value changes & prompting isn’t enough Instead you start to get erratic conflicts between prompts and training, with erratic & unpredictable results