Actually a pretty good instruction following test. Ideally I’d run 10 samples/model with temperature 1.0 and all other bells and whistles (topp/k) off.
LLMS
-
Analysis of underlying models and search indices for AI tools
By
–
SearchGPT uses Bing search index
Perplexity uses ChatGPT models Who will survive? -

Llama 3.1 Uses Synthetic Data Not Distillation Says Expert
By
–
Many people say that Llama 3.1 is distilled, but this is incorrect: it's just synthetic data. Here's a plot from kindacognizant (
https://
reddit.com/r/LocalLLaMA/c
omments/1ed58iu/llama31_models_are_fake_distillations_this_should/
…) showing the difference in probability distributions. Yeah, it's NOT EVEN distilled. Still more performance to tap into. -
MMLU Evaluation Method: Log Probability vs Multiple Choice
By
–
Nice! Btw it's possible (in principle) to also evaluate MMLU in the same way I evaluate HellaSwag, where you swap out the 4 continuations in turn and predict the one with highest average log prob. Though it hurts the model by a few percent because it can't reason by elimination.
-
User experience with SearchGPT accessibility
By
–
I don’t have access to SearchGPT So it is now entirely bad
-
FLAN Model Limitations in Modern LLM Evaluation
By
–
I don't think we can draw any conclusion about modern LLMs since it was made with FLAN
-
SearchGPT Multimodal Capabilities and Weather Widget Analysis
By
–
How is the weather! 🌦️
— 🚨 AI News | TestingCatalog (@testingcatalog) 27 juillet 2024
More SearchGPT examples with weather widgets and media responses.
Seems like it’s image multimodality is quite similar to how 4o can work with images from the internet 👀 https://t.co/qgbBZIp50YHow is the weather! More SearchGPT examples with weather widgets and media responses. Seems like it’s image multimodality is quite similar to how 4o can work with images from the internet
-

Major AI Model Announcements: Llama3, Mistral, SearchGPT Revealed
By
–
𝐋𝐞𝐬 𝐉𝐎 𝐝𝐞 𝐥'𝐈𝐀 𝐬𝐨𝐧𝐭 𝐨𝐟𝐟𝐢𝐜𝐢𝐞𝐥𝐥𝐞𝐦𝐞𝐧𝐭 𝐨𝐮𝐯𝐞𝐫𝐭𝐬
Mais quelle semaine olympique alors que s’ouvre #Paris2024 ! Mark Zuckerberg dévoile sa vision pour #Llama3 #MistralLarge2 est aussi bon que #Llama3 Open AI balance son #SearchGPT au nez de -
Comparing LLM Context Windows for Coding and Research Tasks
By
–
Mainly Projects and higher rate limits. I feel like with 3.5 Opus it might get much more interesting. For projects it still depends: It might be good for research but for Coding it still has a very tiny context window. Gemini 2M for Workspace is more interesting in that sense

