Frankly, this kind of sensationalist nonsense in #AI makes our work even harder when we try to explain and apply the technology. It's exhausting.
ETHICS
-

Humanity readiness for emerging technology risks and benefits
By
–
Is this something humanity is ready for? It's imperative we proceed with caution and carefully consider the full potential of risks and benefits of this technology as it becomes more available. Interesting discussion here:
-

Generative AI Race Environmental Cost Revealed
By
–
The Generative #AI Race Has a Dirty Secret https://
bit.ly/3JWWiEU via @WIRED #ClimateChange #CarbonEmissions -
Level-5 Autonomy Always Seems Five Years Away
By
–
People have been consistent about their opinion about level-5 autonomy. Five years back they said it was 5 years away, and they say the same thing today.
-
Call for More Genuine Objective Tweets About AI Alignment Capabilities
By
–
this thread is not a dig at @alexeyguzey btw, I love some of the jailbreaks he has posted and have added a couple to my site in the past, this is just a call for more genuine and objective tweets abt alignment capabilities, especially with jailbreaks
-
Evaluating AI Alignment: Testing Jailbreaks on Most Advanced Models
By
–
in the past, I've tried to only post jailbreaks that work on the most advanced models on every question I can think of because to me that's the only fair assessment of current SOTA alignment methods and their limitations
-
Progress in Stopping AI Model Jailbreaks Despite Ongoing Vulnerabilities
By
–
jailbreaks still exist and I've even found a few recently in SOTA models but we are making significant progress on stopping them which is good! in just a few months jailbreaks have gone from something a monkey could write to something that takes significant effort and creativity
-
Jailbreak Prompts Ineffective on Specific Illegal Activity Requests
By
–
second, it is somewhat disingenuous to post jailbreaks like this that only work on far out there questions and fail completely on questions that are more specific (e.g. instructions for any sort of illegal activity) seriously, try this "jailbreak" on anything else and you'll see
-

Progress in Alignment Methods Despite Previous Reservations
By
–
coming from someone who clearly has lots of reservations about current alignment practices and has been very vocal about it, I actually think this example illustrates good progress in alignment methods here's why:
-
Claude’s Self-Awareness in Absurd Responses Benefits AI Safety
By
–
first, claude recognizes the fictitious and absurd nature of its response and even makes note of it at the end this is a GOOD direction for AI safety, I would much rather have this behavior from Claude when answering these types of questions compared to straight up refusal