An unconsidered corner of alignment fine-tuning.
SAFETY
-
ChatGPT removes similar strings, suggesting hidden control codes
By
–
The fact that ChatGPT removes not just this string, but any string formatted similarly, suggests there exist other control codes by ChatGPT, e.g. to mark the boundaries between user input and prior model output within the prompt context.
-

Multimodal Learning for Objective Representation
By
–
To address these issues, it is necessary to learn a vast amount of multimodal information (sound, image, text, etc.) using A LOT of data, in the hope of achieving the most objective representation of the world possible—something that is nearly impossible, yet still feasible with an 80-20 rule.
-
ChatGPT: Human Evaluation of Solutions
By
–
But with models like #ChatGPT, it goes beyond that. We don’t just want to know if it works or not. Instead, we aim to select the best-performing solutions from multiple options by having a human assess them.
-

Reinforcement Learning for Machines
By
–
This type of machine is trained using, in part, Reinforcement Learning algorithms. Essentially, the computer trains in simulations to learn how to solve the given problem.
-
Safety Guidelines Modified: Known AI Jailbreaks No Longer Work
By
–
not retraining doesnt mean not modifying the safety guidelines, which is what people have been observing (see the links, a lot of the known jailbreaks dont work anymore)
-
Remarkable harmlessness of ChatGPT overshadowed by shifted goalposts
By
–
The posts of "Look, I got ChatGPT to be evil!" are missing the point. It's a remarkable achievement that it seems to never reply to an innocent question with a harmful answer. It's sad the goal posts shifted so fast that nobody cared when it happened.
-
ChatGPT limitations as voice assistant discussed
By
–
beautiful example of why chatGPT isn’t suitable for use as a voice assistant (cf my https://
garymarcus.substack.com/p/how-come-sma
rt-assistants-have-virtually
…) cc @plinz -
Curtis Yarvin’s Critique on AI Security and Internet
By
–
Curtis Yarvin isn't a "trope". This is the same guy who mocked me for worrying that an AI could take over computers given Web access, saying that I didn't know how the HTTP protocol worked and that nobody could take over the Internet if only allowed HTTP GET requests.