So two things are happening here: first is that the OSS community (and potentially Meta) seems to be settling around 3B and 7B. Second is it's surprising Google did not go straight for GPT-4—it seems to me really unlikely that they wouldn't be able to hit that if they wanted.
@mattlynley
-

PaLM 2 vs GPT-4 MMLU Performance Comparison Analysis
By
–
PaLM 2 (L) also does not outperform GPT-4 on an MMLU 5-shot, but it does outperform LLaMA-I, a fine-tuned model for LLaMA. Here's what I pulled from the technical reports for LLaMA, PaLM 2, and GPT-4.
-
Small Language Models: 3B and 7B Parameters Emerge in Open Source
By
–
We are seeing a lot of models pop up in the 3B and 7B param range, incl RedPajamas-INCITE 3B from Together that is based on the RedPajamas dataset. That dataset is designed to replicate the LLaMA dataset. I've heard Meta is also exploring 3B as an option for its fully OSS model.
-
Google PaLM 2 Comes in Four Sizes with Mobile Capabilities
By
–
When it comes to PaLM 2, Google said it would come in four sizes. It didn't disclose parameter count, but it did give us some tea leaves to evaluate, the first of which is that Gecko (its smallest) can run on a mobile device offline.
-
LangChain’s Challenge Converting Framework into Business Model
By
–
On the first one, no one disputes that LangChain is a transformative technology—especially as we move to more agent-based workflows. But LangChain is also a Startup That Has Raised Real Capital, and it faces an uphill climb to convert a beloved framework into a business.
-
LangChain Business Model and Google’s PaLM Announcement Analysis
By
–
Today's issue of Supervised is gonna cover two topics: first, LangChain and questions about its business; second, breaking down some of Google's PaLM announcement. https://
supervised.news/p/parrots-geck
os-and-frameworks-as
… -
Gandalf: Password Guessing Game via Prompt Engineering
By
–
Fun little game similar to the SQL murder mystery if you've played it: try to guess the password via prompt engineering with the difficulty going up over time. https://
gandalf.lakera.ai -
Context Window Limitations in AI Models Performance Metrics
By
–
Also obvious but there are reasons the context windows are in the 4k range. Will have to wait and see performance metrics/cost before jumping to any conclusions.
-
8K Technology Experimentation Pushes Extreme Hardware Boundaries Forward
By
–
There's been some experimentation in 8K but this takes it to quite the extreme end so really curious what use cases come out of this.
-
Hidden Prompts Impact on LLM Context Windows Development
By
–
The importance of hidden prompts is often wildly underestimated and this is going to be really interesting to see how this develops. Most context windows are in the 4K range. https://t.co/6MVuvX4kt0
— Matthew Lynley (@mattlynley) 11 mai 2023The importance of hidden prompts is often wildly underestimated and this is going to be really interesting to see how this develops. Most context windows are in the 4K range.