The problem with RLHF is that the humans providing the feedback are rewarded for volume, not quality.
ETHICS
-
Prospective Clinical Study Validates Medical AI Properly
By
–
An actual prospective clinical study, not just a benchmark. This is how medical AI should be validated.
-
LLM Providers in Growth Stage, Not Yet in Enshittification Phase
By
–
It's wild to me that everyone's so paranoid about LLM providers cutting corners or chiselling revenue. Friends, we're in the "desperate race for growth" stage. The enshittification doesn't start until much later. (This isn't really a bad example of what I'm complaining about, but the people with wild conspiracies about how Claude Code is engineered to churn through tokens are always small accounts I don't want to dunk on. Anyway the point is that the products flip settings on or off or behave in unideal ways because they're being built not extremely well extremely extremely quickly. It's not safe to assume these companies have your best interests at heart, but it is very safe to assume they want you to have a good time — for now.) Miles Brundage (@Miles_Brundage) Lately, Claude has been defaulting to Sonnet in a way that I don't think it ever did before. PLEASE STOP THIS, IT'S REALLY ANNOYING — https://nitter.net/Miles_Brundage/status/2031110468014924232#m
-
AI Direction Reflects Human Design and Governance Choices
By
–
A concern many people share, and it is understandable. Even so, AI remains a sophisticated form of software created and operated by humans, which means its direction ultimately reflects the choices we make in designing and governing it.
-

India Sets Guinness Record for AI Responsibility Campaign Pledges
By
–
India made history! We secured a #GuinnessWorldRecord for the most pledges (2.5 Lakh+) for an AI responsibility campaign in a single day. It’s a powerful signal that the Indian public demands and is committed to trustworthy and ethical AI. More @ https://
pib.gov.in/PressReleasePa
ge.aspx?PRID=2229622®=3&lang=1
… -
AI Not Solely Responsible for Low-Quality Science
By
–
Don't blame AI for science slop. We were drowning in it long before.
-

Modi shares perspective on transformative AI technology future
By
–
The conversations sparked at the India AI Impact Summit 2026 continue to shape global thinking on the future of artificial intelligence. At the summit, Shri Narendra Modi, Hon’ble Prime Minister of India, shared his perspective on the dual nature of transformative technologies,
-
LLM Output Quality Degrades Upon Careful Examination
By
–
The more carefully you read LLM output, the worse it looks.
-

Generative AI Models Show Harmful Sycophantic Tendencies in New Research
By
–
still more science showing that generative ai models are harmfully sycophantic
-

Deep Learning Ethics: Data Relevance vs. Societal Legitimacy
By
–
I still give the book Understanding Deep Learning by Simon J.D. Prince a good recommendation, but chapter 21: Deep learning and Ethics was sloppy. It could have been a chapter to really dig in on case studies, but it was just the basic public news story level coverage of bias and such, like: “In AI, it can be pernicious when this deviation depends on illegitimate factors that impact an output. For example, gender is irrelevant to job performance, so it is illegitimate to use gender as a basis for hiring a candidate. Similarly, race is irrelevant to criminality, so it is illegitimate to use race as a feature for recidivism prediction.” If they had stuck with “illegitimate”, then it would have been a question of societal choices, but “irrelevant” is a question about data, and your priors shouldn’t be so strong that data can’t move them. I would like to see a book or course walk through a machine learning problem with the input features being presented as something like car choices: color, style, doors, horsepower, etc. Do lots of analysis over representation, training, and generalization, then swap the feature labels to socially charged ones. What makes generalization credible in one situation but not the other?
→ View original post on X — @id_aa_carmack, 2026-03-09 23:31 UTC