GPT-4 has a usage cap of 40 messages every 3 hours for Pro users. Once you run out, it'll put a message up that you're running on GPT-4o until the next time period.
LLMS
-

Leading AI Models Evaluated: Coding, Math, Multilingual Performance
By
–
3/ We eval'd many of the leading models: – GPT-4o
– GPT-4 Turbo
– Claude 3 Opus
– Gemini 1.5 Pro
– Gemini 1.5 Flash
– Llama3
– Mistral Large On Coding, Math, Instruction Following, and Multilinguality (Spanish). See leaderboard results below. -
Third-party AI evaluations and overfitting prevention strategies
By
–
4/ While LMSYS and other efforts in the community are awesome, we still think there's a lot to be desired in 3rd party evaluations. One of our design principles is to produce evals that are impossible to overfit. As we saw with our prior GSM1k research, we think it's critical
-

Scale Launches SEAL Leaderboards for Frontier Model Evaluation
By
–
1/ We are launching SEAL Leaderboards—private, expert evaluations of leading frontier models. Our design principles:
Private + Unexploitable. No overfitting on evals!
Domain Expert Evals
Continuously Updated w/new Data and Models Read more in http://
scale.com/leaderboard -
GPT-4o vs Claude Opus: Code Generation Performance Comparison
By
–
Not sure they're doing the exact same eval, but GPT-4o reports 90.2% vs. 84.9% for Claude 3 Opus on HumanEval (
https://
openai.com/index/hello-gp
t-4o/
…). I don't expect Codestral 22B to be on par, but it can do FIM -
Awesome LLM Apps Repository Open for Contributions
By
–
the GitHub repository to show your support and stay updated with the new updates: https://
github.com/Shubhamsaboo/a
wesome-llm-apps
… Contributions are Welcome If you have any ideas, improvements, or new apps to add, please create a new GitHub Issue or submit a pull request. -

Awesome LLM Apps Repository Achieves Top Trending Status GitHub
By
–
Opensource is amazing. Awesome LLM Apps repo made it to the top trending GitHub repositories across the globe in just a few weeks of its launch. This is one of my biggest dream come true! Thank you awesome community for the support & keep it coming. If you don't know about the
-

Mistral Releases Codestral-22B: Outperforms LLaMA 3 70B
By
–
Mistral has released a code-specific model! Codestral-22B seems to outperform LLaMA 3 70B while being more than 3x smaller. Very impressive!
-
LLaMa Reproducibility: Data Access Essential for Scientific Validation
By
–
Reproducing a benchmark is a minor managerial task. It's not even engineering. The scientific hypothesis of LLaMa models is: "If we overtrain the same model architecture on our data, it outperforms previous iteration." Without the data, LLaMa data the experiment is not repro.
-

GPT-4o launched, GPT-5 soon: video analysis
By
–
→ GPT-4o just released… And GPT-5 is coming!
OpenAI just announced it: GPT-5 is (almost) ready and some see it released as early as June 14.
What's the real deal? I analyze it in my latest video:
https://youtu.be/R0rN–Def7I
