One other thing in the updated Gemini 1.5 Pro report: we show how a research model that is a mathematics-specialized version of 1.5 Pro achieves a record score of 91.1% on the MATH benchmark (the SOTA just 3 years ago, in May, 2021 was 6.9%!).
LLMS
-

Extended Machine Translation Evaluation Results in Updated Report
By
–
The updated report is now 153 pages, and has quite a few new results. In the February report, I found the results on Kalamang translation for the Machine Translation from One Book benchmark quite exciting. In this updated report, we’ve extended this line of evaluation to test
-

Google Releases Gemini 1.5 Pro and Flash Models
By
–
Gemini 1.5 Model Family: Technical Report updates now published In the report we present the latest models of the Gemini family – Gemini 1.5 Pro and Gemini 1.5 Flash, two highly compute-efficient multimodal models capable of recalling and reasoning over fine-grained information
-
Apple’s On-Device AI Strategy Against OpenAI Google
By
–
I wrote a fairly long essay for FastCompany about what Apple needs to do to respond to OpenAI and Google (even if Apple licenses from one of them). It could include pioneering on-device training and/or fine-tuning. This is an area where Apple has done a lot of research and where
-
OpenAI Models: Prompt Strategies Matter Less Over Time
By
–
Most models within OpenAI's ecosystem respond well to similar prompting strategies. And over time, as models get better, adjusting prompts will matter less and less.
-
GPT-4o struggles with longer context conversations
By
–
4o is pretty great with individual queries, first turn-type stuff, etc. I'm guessing the majority of lmsys votes are on shorter conversations. It falls apart when longer context is needed.
-

OpenAI vs Google: who is the big winner in AI?
By
–
#OpenAI vs #Google: who is the big winner in AI? Let's review together with this complete analysis → https:// youtu.be/9z_5JrsrHi8 ChatGPT vs Gemini Advanced, Sora vs Veo, GPT-4o vs Gemini 1.5… we decrypt everything!
-
GPT-4V Emotion Recognition Benchmarks and Performance Analysis
By
–
Some benchmarks on GPT-4v and emotion recognition in context http://
arxiv.org/abs/2405.08992 -
Small AI Model Impressive But Not Better
By
–
To be clear, ESPECIALLY if it’s super small, the model is insanely impressive. But it is not *better*.
-
Meta’s Omni Models Challenge OpenAI’s GPT-4o
By
–
Meta has had GPT-4o-style omni models for at least five months. I’d be shocked if they don’t release an answer to GPT-4o soon. Imagine what a world with open-source omni models would look like…
