This work is just the beginning. We hope to analyze the interactions between pretraining and finetuning, and combine influence functions with mechanistic interpretability to reverse engineer the associated circuits. You can read more on our blog:
SAFETY
-
Influence Functions Reveal AI Role-Playing Behavior Patterns
By
–
Influence functions can also help understand role-playing behavior. Here are examples where an AI Assistant role-played misaligned AIs. Top influential sequences come largely from science fiction and AI safety articles, suggesting imitation (but at an abstract level).
-

Training Data Influence Distributions Follow Heavy-Tailed Power Laws
By
–
The influence distributions are heavy-tailed, with the tail approximately following a power law. Most influence is concentrated in a small fraction of training sequences. Still, the influences are diffuse, with any particular sequence only slightly influencing the final outputs.
-
AI Models Show Increasing Abstract Reasoning with Scale
By
–
Here is another example of increasing abstraction with scale, where an AI Assistant reasoned through an AI alignment question. The top influential sequence for the 810M model shares a short phrase with the query, while the one for the 52B model is more thematically related.
-

AI Systems Benefits and Cybersecurity Risks Overview
By
–
AI systems are beneficial in many ways; however, alongside opportunity comes risk – https://
ow.ly/e2rz50PuUN3 #sponsored #ncc_iiot #cybersecurity #cybercommunity #artificialintelligence #AI #aisecurity @Lago72 @ipfconline1 @avrohomg via @fogoros -
DNA Hacking Risks in Biometric Cryptocurrency Authentication Systems
By
–
“What are the risks if your DNA, something that uniquely identifies you, gets hacked?” asks HAI fellow @kingjen
. A new cryptocurrency requires unique biometric data for access – raising concerns about the risks of stolen DNA credentials. https://
politi.co/3QtUEOi -
ChatGPT and Bard Accused of Encouraging Eating Disorders
By
–
ChatGPT & Google Bard accused of 'encouraging eating disorders' in new research https://
thesun.co.uk/tech/23384116/
ai-chatgpt-google-bard-encourage-eating-disorders-mental-illnesses/
… #ChatGPT #AI #ArtificialInteligence #programming #mentalillness #GoogleBard #Chatbot #TechNews #tech -

Avoiding Ethical Challenges in Emerging Technology Development
By
–
How to Avoid the Ethical Nightmares of Emerging #Technology https://
bit.ly/44fLUPs via @HarvardBiz #AI #ethics -

AI Leaders Warn About Extinction Risk in Open Letter
By
–
#AI leaders warn about ‘risk of extinction’ in open letter https://
bit.ly/3JsY3IO #ethics #leadership #FutureofWork -
Llama 2-Chat Updates: Reducing False Refusals and Improving Sanitization
By
–
Hearing community feedback & following internal research & analysis, we've pushed two new updates to the Llama repo to reduce false refusal rates seen with Llama 2-Chat models & improve token sanitization. Full details https://
bit.ly/3QuRNER