Finally got a chance to spend some time with the 77-page Llama 2 paper (
https://
arxiv.org/abs/2307.09288). Here are some takeaways at a glance. Appreciate that someone finally did a comprehensive supervised finetuning vs RLHF evaluation! (Lower right)
LLMS
-

Llama 2 Paper Analysis: Supervised Finetuning vs RLHF Evaluation
By
–
-
TGI Inference Project Positioning Debate
By
–
don't really agree here:
1/
TGI is/was super early in its history (very few outside contributors, and we discussed with all of them) 2/
it's smaller than other similar inference projects (vllm, …) so unclear it had even "taken the lead" Happy to discuss further though! -
InstructGPT and ChatGPT represent significant advancement beyond previous models
By
–
I’d say that InstructGPT, ChatGPt, and everything after (which were all released after this paper) were quite a leap forward
-

GPT Engineer: Automatic AI-Powered Codebase Generation
By
–
GitHub – AntonOsika/gpt-engineer: Specify what you want it to build, the AI asks for clarification, and then builds it.
https://bit.ly/3pnaWgS It generates an entire codebase based on a prompt. #AI #MachineLearning #DeepLearning #LLMs #DataScience -
AI-Generated Meeting Minutes Raise Questions About Summarizing Vague Discussions
By
–
Génial de pouvoir faire des comptes rendus de réunions automatisés avec l'IA. Problème : dans les grands groupes, comment l'IA fera-t-elle pour résumer du bullshit sans actions concrètes ? https://
searchenginejournal.com/openai-publish
es-tutorial-for-ai-generated-meeting-minutes/492789/
… -

ChatGPT Struggles Generalizing Math to Uncommon Bases
By
–
ChatGPT can do math, but what if it's in base-9? MIT researchers show that while ChatGPT-like models can perform many tasks, they struggle to generalize to uncommon task variants, indicating a lack of underlying abilities: https://
bit.ly/3NZOz9s @zhaofeng_wu -
LLaMA Training and PEFT for Neural Networks Zero to Hero
By
–
LLaMA training & PEFT would be a very appreciated entry in NN Zero to Hero :-). And plays well with recent entries(nanoGPT, State of GPT).
-
Neural Networks and Gzip versus Large Language Models Deep Dive
By
–
Following up on some discussions of the "NN + Gzip versus LLMs" methods a few weeks back here's my little deep dive with a reimplementation and additional experiments: https://
magazine.sebastianraschka.com/p/large-langua
ge-models-and-nearest/
… This includes fixing the tie-breaking and more! -
LLM Hallucination: Document Analysis vs Factual Accuracy
By
–
it is quite amazing how LLMs like claude v2 can in one turn give you solid responses about a document i uploaded and in the next turn when i ask about a specific person, the response is entirely made up. –> further to @bxchen recent piece, LLMs are useful when constrained to a
-
Google Robots Become Smarter With AI Language Models
By
–
Aided by #AI Large Language Models, Google's Robots Are Getting Smart https://
nytimes.com/2023/07/28/tec
hnology/google-robots-ai.html
…