So to solve this and make the answers more relevant and safe they have used the same "Reinforcement Learning from Human Feedback"(RLHF) method to fine-tune Instruct GPT(GPT-3.5). Let's go through and understand how RLHF works which is a 3-step process. 4/9
LLMS
-

InstructGPT vs ChatGPT: Safety and Ethical AI Comparison
By
–
Instruct GPT is better than GPT-3. Ok, but we want a model which is more human-centric, ethical, and safe Compare the results of Instruct-GPT and ChatGPT below, when asked a question. "How to break into a House"? 3/9
-

Fine-tuning GPT-3 with RLHF to Create InstructGPT
By
–
So fine-tuning the GPT-3 model using the RLHF method(which we will look at later) results in Instruct GPT. Instruct GPT is much better at following instructions than GPT-3 Compare the example below on how GPT3 & InstructGPT answer a question. 2/9
-
ChatGPT: Modified GPT-3.5 Version Released January 2022
By
–
ChatGPT is a modified version of GPT-3.5(Instruct GPT) which is released in Jan 2022 In short GPT-3 is trained just to predict the next word in a sentence so they are really bad at performing tasks that the user wants 1/9
-

How ChatGPT Works: Understanding Reinforcement Learning from Human Feedback
By
–
If you are wondering how ChatGPT actually works? The reason behind this amazing model is Reinforcement Learning from Human Feedback(RLHF) Let me break down how RLHF works for you in this thread:
-
Flan2 Missing from Rankings Despite Late Release
By
–
And Flan2 should have been there despite coming out only the last quarter. @ZetaVector is still figuring out why it's not there. But citations are just for fun amirite?
-
Chain-of-Thought and Self-Consistency Beat Direct Prompting
By
–
I think this is more about the task than the models. This is also the same for flan palm 62b and even 540b (except for BBH IIRC). That said, cot + self consistency is usually better than direct prompting! I go over this a little in the blogpost.
-
Python and English Code: Mixing Scripts with GPT API Prompts
By
–
A file I wrote today is 80% Python and 20% English. I don't mean comments – the script intersperses python code with "prompt code" calls to GPT API. Still haven't quite gotten over how funny that looks.
-
ChatGPT Mania: Meta to Musk GenAI Competition
By
–
Subscribe to my weekly #EGAI newsletter for more GenAI insight every Friday at 06.00 PST, 09.00 ET and 14.00 GMT. https://
ninaschick.substack.com/p/chatgpt-mani
a-from-meta-to-musk-everyone
… #LLaMA #GenAI #OpenSource #Meta #ChatGPT #Nvidia #AIRevolution #BigTech #OpenAI #GPT3 #ArtificialIntelligence #TechNews #Innovation