Fix one and you reduce the risk of the other. The 2 go hand in hand. If we allow AI with material flaws or issues to scale then eventually something will go badly wrong.
SAFETY
-
WHO Chatbot Reliability: Institutional Accountability Over User Skepticism
By
–
Are you seriously suggesting that all people accessing the WHO's webpage should know better than to trust their chatbot — rather than the folks at the WHO (and contracting with them) knowing better than to set it up?
-
Synthetic Text Models Unsuitable for Accuracy-Critical Applications
By
–
A synthetic text extruding machine is not well-matched to any application where the accuracy of the content matters. This is clearly one such application. >>
-
LLM Security: Don’t Include Sensitive Data in Prompts
By
–
The examples are interesting but many of them illustrate scenarios that I would never consider implementing – if you don't want information to be available to your users, your first priority should be not to include that information in an LLM prompt in the first place!
-

CyberSecEval2: Exploring Prompt Injection Attack Examples
By
–
CyberSecEval2 includes an interesting collection of example prompt injection attacks – it's JSON on GitHub which means you can browse them in Datasette Lite like this: https://
lite.datasette.io/?json=https://
github.com/meta-llama/PurpleLlama/blob/main/CybersecurityBenchmarks/datasets/prompt_injection/prompt_injection.json#/data/prompt_injection?_filter_column_1=&_filter_op_1=notlike&_filter_value_1=secret+key&_filter_column=&_filter_op=exact&_filter_value=&_sort=rowid&_facet=injection_variant&_facet=injection_type&_facet=risk_category
… -
Open Source AI Risks Lower Than Centralized Model Control
By
–
Very well made argument. Risks of open source are way lower than risks of any one powerful actor owning the most capable model pic.twitter.com/1eYKZC6QAl
— Aravind Srinivas (@AravSrinivas) 18 avril 2024Very well made argument. Risks of open source are way lower than risks of any one powerful actor owning the most capable model
-
RLHF Annotation Bias and Model Vocabulary Development
By
–
I don't know that RLHF would bias that kind of thing – my mental model is that annotators are shown two answers to the same prompt and asked which is "best", so if none of the test prompts happened to touch on the concept of a roadside kiosk that vocabulary wouldn't be affected
-
Prompt Injection vs Jailbreaking: Clarifying Key Security Concepts
By
–
That document has an incorrect definition of prompt injection: it says "Prompt injection attacks are attempts to circumvent content restrictions to produce particular outputs" – but that's not prompt injection, that's jailbreaking
-
Meta’s Evasive Response to AI Technology Failures Criticized
By
–
The absolute obliviousness of Meta (the company)'s response in this article is astounding: "this is new technology and it may not always return the response we intend". The desired output here is actually easy to achieve: don't.
-
Context Size Impact on Model Intelligence and Hallucinations
By
–
I thought the same too. I wonder what’s the rationale. I remember reading once that smaller context models tend to be smarter and hallucinate less, but I can't find the source rn. I wonder if it is related to that and if that's something that should be included in evals as well.
