– inverting prompts from logits
– training data detection (
@WeijiaShi2 worked on this)
– distillation
@jxmnop
-
Inverting Prompts, Training Data Detection, and Distillation Techniques
By
–
-
Researchers Jailbreak ChatGPT API by Accessing Hidden Token Probabilities
By
–
fun research story about how we jailbroke the the chatGPT API: so every time you run inference with a language model like GPT-whatever, the model outputs a full probabilities over its entire vocabulary (~50,000 tokens) but when you use their API, OpenAI hides all this info from
-
How I Got My First Google Job Through Self-Teaching
By
–
how I got my first job at google > be me
> go to college
> great state school, smallish CS program > google does not recruit here
> that’s ok
> no worries
> build personal website
> start working on cool blog post for site
> write some python code for blog
> confused about -
Current LLM Myopia: Critical Perspective on Language Models
By
–
probably not tbh, except maybe myopia wrt current LLMs
-
Seam Carving Technique Applied to Textual Data Processing
By
–
this is super neat, it’s kind of like the textual equivalent of “seam carving” in computer vision https://
x.com/maxkreminski/s
/maxkreminski/status/1743065463616397424
… -
Basilisk and Singularity: AI Existential Risk Reflection
By
–
I personally find the basilisk to be totally awesome, and admit that the above tweet may come back to bite me when the singularity arises
-
Echo chambers protecting ideas from external criticism in AI
By
–
it's a brilliant way to protect your ideas from external criticism
-

Why AI Doomers Use Complex Abstract Arguments
By
–
one reason many people don't argue with the AI doomers is because their ladder of abstractions is truly sky-high one second you've refuted the existence of roko's basilisk. suddenly they're calling out your spurious counterfactuals and myopia, accusing you of simulacrum level 4
-
Outstanding Master’s Thesis in Artificial Intelligence Research
By
–
best masters thesis ever (so far)