Leave a few shrimp tempura for me Jokes aside — humans need to be very careful about setting the right objectives and incentives for narrow and general AI systems!
SAFETY
-
LLMs Cannot Verify Truth: The Persistent Problem
By
–
What LLMs say is sometimes true, sometimes not. They can’t tell the difference: they don’t know how to do validity checks (eg crossreferencing WIki or their own training corpus). That’s what makes it BS. First said it in @techreview 2020; still true: https://
technologyreview.com/2020/08/22/100
7539/gpt3-openai-language-generator-artificial-intelligence-ai-opinion/
… -
AI Safety Requires Better Control and Regulation for All
By
–
“AI safety isn’t a right or a left issue. New, poorly controlled AI that is unreliable…can easily fool people & may set off a wave of cybercrime, & lead to…atmosphere of distrust. It’s in everybody’s interest that we move to a safer, better controlled form of AI.” –
@garymarcus -
AI Intelligence Growth and Existential Risk Concerns
By
–
It's not smarter than us yet. I expect it to get smarter later, the same as it's been getting smarter over the past, and then to kill everyone. Do you remember how you acquired the belief "Eliezer thinks LLMs are smarter now"? I want to understand how to prevent this.
-
PCA’s Relevance to AI Alignment Work Questioned
By
–
*Sigh.* I knew that meant "principal components analysis" without looking it up. I also knew about the controversy with ICA being patented. Now, what specific implication do you think PCA has for alignment work, exactly?
-
Dataset Filtering for Large Language Model Training Runs
By
–
Possibly yes. More broadly, if we're going to be doing more large training runs at all, I'd guess we should start filtering the datasets soon (presumably using a previous-generation LLM finetuned for that?). It's hard to finetune out a cognition once it's learned by the base.
-
LLMs Confabulate Inherently Despite Infinite Data
By
–
Conjecture: even with infinite data, LLMs would still confabulate, because they blur the inputs and don’t reliably create precise representations of individuals and their properties. (see The Algebraic Mind, 2001 for related discussion which has thus far has held true)
-
AGI Alignment Knowledge and LLM Developer Expertise Skepticism
By
–
Maybe I shouldn't, but one last shot at explaining my position on AGI gatekeeping. I'm not saying that I know everything known to the high-status inventor or engineer of the largest LLM. I'm saying that I'm dubious that they have secret knowledge deeply relevant to *alignment*,
-

Machiavelli Benchmark: Evaluating LLM Ethics in Adventure Games
By
–
6/ Machiavelli Benchmark – a new benchmark of 134 text-based Choose-Your-Own-Adventure games to evaluate the capabilities and unethical behaviors of LLMs.
-

Accountability in AI: Rejecting Excuses for Harmful Behavior
By
–
Good morning to everyone except those who make their money off making excuses for this behavior.