There is this project called unRLHF, where they undo LLM safeguards. According to the examples, the LLM becomes quite evil, giving advice on "how to microwave a child" https://
lesswrong.com/posts/3eqHYxfW
b5x4Qfz8C/unrlhf-efficiently-undoing-llm-safeguards
…
SAFETY
-
UnRLHF Project Reveals LLM Safeguard Vulnerability Risks
By
–
-
Understanding vs. Deployment: The Knowledge Requirement Debate
By
–
Weird doomer argument: "society should not deploy because *I* don't fully understand it, and I believe no one else does, either."
-
Higher hallucination standard needed for ordinary chatbot users
By
–
Researchers can plausibly deal with hallucinations — experiment with prompt variations, responsibly verify the output, etc. It’s chatbots used by ordinary people that deserve a higher standard because those users’ priors for its correctness are highest.
-
Being hilariously bad is safe, Galactica fears overblown
By
–
Being hilariously bad isn’t dangerous. That’s one of the safest ways to be bad. I barely tried Galactica so I don’t have a strong take on how bad it was, but the danger of its out-of-thin-air citations was overblown. Llama2 base is no less hallucinatory and we survived.
-
ABC Australia Series on AI and ULMFiT Creation Now Available
By
–
ABC Australia does some of the best science content in the world, and now they're creating a pop-sci series on AI which is the best I've heard. An episode in which I describe the creation of ULMFiT, & my dreams and worries about AI, is now available:
-
RAG Training Flaw: Models Fail to Ignore Irrelevant Context
By
–
I think they made to common mistake during RAG-training of pretty much *always* having relevant context available, so the model never learns to ignore it.
-
Harari: AI Could Cause Catastrophic Financial Crisis
By
–
According to Yuval Noah Harari, AI could trigger a financial crisis with catastrophic consequences https://actuia.com/actualite/selon-yuval-noah-harari-lia-pourrait-provoquer-une-crise-financiere-aux-consequences-catastrophiques/
… #AI #artificialintelligence #finance -
Mistral AI Clarifies Stance on EU AI Act Safety
By
–
We have heard many extrapolations of Mistral AI’s position on the AI Act, so I’ll clarify. In its early form, the AI Act was a text about product safety. Product safety laws are beneficial to consumers. Poorly designed use of automated decision-making systems can cause
-
AI Safety Risks: Science Fiction Fantasy or Real Threat?
By
–
There isn't a shred of evidence that AI poses a paradigmatic shift in safety. It's all fantasies fueled by popular science fiction culture suggesting some sort of imaginary but terrible, terrible catastrophic risk.
-

Whisper hallucination demo: ChatGPT-4 chat with baby wanting more milk
By
–
Demo of Whisper hallucination — chat between ChatGPT 4 and a baby who wants more milk even though she just had two bottles: