4/ While LMSYS and other efforts in the community are awesome, we still think there's a lot to be desired in 3rd party evaluations. One of our design principles is to produce evals that are impossible to overfit. As we saw with our prior GSM1k research, we think it's critical
SAFETY
-
Moving AI from Prototyping to Production with RAG Testing
By
–
How can your organization move beyond AI prototyping to production solutions with confidence?
— DataRobot (@DataRobot) 29 mai 2024
Our new launch brings:
1️⃣ Complete set of testing and evaluation techniques to ensure your RAG workflows are built correctly
2️⃣ Suite of safeguards and intervention methods to keep your… pic.twitter.com/J9308oOSZCHow can your organization move beyond AI prototyping to production solutions with confidence? Our new launch brings: Complete set of testing and evaluation techniques to ensure your RAG workflows are built correctly Suite of safeguards and intervention methods to keep your
-
Reproducible Inference Lacks Scientific Rigor Without Hypothesis
By
–
Reproducible inference means nothing for models, who cares? "I put in this data and got this data out." Congrats. What was the hypothesis?
-
LLM Interpretability and Refusal Mechanisms Research Acknowledgments
By
–
Special thanks to: – @failspy for his notebook and abliterated models
– Arditi et al. (cc @NeelNanda5
) for the excellent "Refusal in LLMs" blog post (
https://
lesswrong.com/posts/jGuXSZgv
6qfdhMCuJ/refusal-in-llms-is-mediated-by-a-single-direction
…)
– All the LLM mergers and fine-tuners for the source models
– Charles Goddard and @arcee_ai for MergeKit -

Removing Meta Alignment from Daredevil-8B Model
By
–
But Daredevil-8B is still censored. Many people asked me for uncensored models, so I wanted to try something. I used @failspy
's abliteration notebook to remove Meta's alignment. Unfortunately, abliteration also slightly degrades performance: 1-2% on every benchmark. -

OpenAI trains next model; safety lead joins Anthropic; AI tools launched
By
–
Top stories in AI today: -OpenAI begins training the next model
-Former OpenAI safety co-lead joins Anthropic
-Create unique AI music tracks in seconds
-Google Chromebooks get AI infusion
-6 new AI tools & 4 new AI jobs Read more: http://
therundown.ai/p/sam-altmans-
new-safety-squad
… -
AI Model Reproducibility and Training Data Transparency Issues
By
–
The models are trained with unknown data, they are neither theoretically reproducible. Nor are they practically reproducible since there's no training code and details matter.
-

Claude AI Implements EU Election Information Warning Banner
By
–
Claude AI is rolling out a new EU Parlament elections warning This banner indicates that Claude AI cannot provide accurate information on this topic and refers to the official resource.
-
Google AI Overviews failures reveal disruption of information relationships
By
–
We've all been laughing at the obvious fails from Google's AI Overviews feature, but there's a serious lesson in there too about how it disrupts the relational nature of information. More in the latest Mystery AI Hype Theater 3000 newsletter:
-
AI Co-dependency: Why Writers Must Shape the Conversation
By
–
Must-read on our AI co-dependency & proof why we need writers to join the conversation.“It occurred to me that I wasn’t really training Brenda to think like a human, Brenda was training me to think like a bot, and perhaps that had been the point all along.”