In the race for better conversational AI, bigger isn’t always better. Recent AI research by @chai_research introduces “Blending,” a method where smaller AI models collaborate to match or even exceed the performance of massive models like ChatGPT (175B+ parameters).
LLMS
-

Cheaper Better LLM Alternative: AI Model Blending Approach
By
–
Forget about GPT-4, Claude and Gemini. Introducing a Cheaper, Better Alternative to Trillion-Parameters LLMs AI Model Blending Is All You Need:
-
Discussion on the release of Google’s GEMs AI agents
By
–
I feel like they should be able to release GEMs at least It should be trivial to do
-

Google plans Gemini update to coincide with Microsoft Build
By
–

Google is planning to ship another Gemini update on the day of Microsoft Build keynote
-
Hermes 2 Theta Llama 8B Model Now Available on Replicate
By
–
https://replicate.com/nousresearch/hermes-2-theta-llama-8b
-

Hermes 2 Θ: Open Source Language Model Innovation
By
–
This is a really cool new model, and a hint at the future of open source language models. Synthetic data, model merges and RLHF are becoming crucial tools in the AI hacker's toolkit. Try Hermes 2 Θ on Replicate right now at https://
replicate.com/nousresearch/h
ermes-2-theta-llama-8b
… and let us know what you think! -

IMO-Bench: Gemini Math Achieves 25% on Mathematical Olympiad Problems
By
–
Glad to see us keep pushing the frontier on maths! Also a bit of a teaser to IMO-Bench (after #alphageometry), which my team (with @quocleix
's support) built & the best model only got 25% 🙂 Check out the cool answer of the Gemini Math on an APMO problem in @OriolVinyalsML
's -
Model Evaluation and Reasoning-Focused Haystack Testing Improvements
By
–
I don’t always find the middle to be lost — really depends on the model/use-case. We definitely need better (more reasoning-focused) haystack tests!
-
Context Window Usability Metric for Language Models
By
–
Model providers should start publishing 'Usable Context' as a metric. Just because a model can support 10M tokens doesn't mean it can use all 10M effectively. Often, I find models >32K can only effectively use a smaller portion of their context window.
-
OpenAI’s Strategic Model Release Signals Imminent Advancement
By
–
I highly doubt OpenAI has hit a wall. Two strong acceleration signals: – ChatGPT makes the bulk of the money for OpenAI — they wouldn't put out a GPT-4-ish level model free for everyone if they didn't have a much better model coming very soon – If the exiting team members