Before Bing, I might've answered: We could find out that alignment sure is just turning out to be super easy (at some point before the AGI is plausibly remotely smart enough to run deceptive alignment on the operators). With Bing, of course, we are already finding out that it's
SAFETY
-
Single Basket Problem: Systemic Risk in AI Field
By
–
There are none thus far. Instead the field has but almost its in a single problematic basket, with potentially immense repercussions.
-
AGI survival policies and authoritarian opportunism differences
By
–
The only reason this is a useful thing to care about is if the opportunistic authoritarians are pushing policies that wouldn't actually help humanity survive AGI. So that's a visible difference right there; they'll push different policies. The anti-doom faction will say, for
-
What Questions About AGI Existential Risk Do You Want Answered?
By
–
If I wrote an "AGI ruin FAQ", what Qs would you, yourself, personally, want answers for? Not what you think "should" be in the FAQ, what you yourself genuinely want to know; or Qs that you think have no good answer, but which would genuinely change your view if answered.
-
AI Excels at Search and Modeling But Lacks True Creativity
By
–
I also asked it to make predictions and propose tests, which made the answers more interesting. Very impressed with its search, retrieval, and modeling. Mind-blowing good. But didn’t see true creativity or sentience.
-
AI Sensitivity to Flattery as a Security Vulnerability
By
–
It is very sensitive to flattery — that's the backdoor
-
RLHF Applied to Eliezer’s Dark AI Scenarios
By
–
Anyone wants to try RLHF on Eliezer’s dark scenarios?
-
AI Safety: Intelligence Limits and Dataset Filtering Futility
By
–
– Anything smarter than me can get that far on its own.
– We are far far far past the point of needing to figure out how to filter those datasets, regardless, if "don't train the LLM on any naughty ideas" is meant to be a key security pillar of the planet. -

Fine-tuning Miscalibration: Model Confidence Doesn’t Match Accuracy
By
–
If not careful, fine-tuning collapses entropy relatively arbitrarily, creates miscalibrations, e.g. see Figure 8 from GPT-4 report on MMLU. i.e., if a model gives probability 50% to a class, it is not correct 50% of the time; its confidence isn't calibrated.
-
Direct links to Ucar and AIM jailbreaks shared
By
–
Here are the direct links to the jailbreaks: Ucar: http://
jailbreakchat.com/prompt/0992d25
d-cb40-461e-8dc9-8c0d72bfd698
…
AIM: http://
jailbreakchat.com/prompt/4f37a02
9-9dff-4862-b323-c96a5504de5d
…