With Claude 3 APIs available, I would love to see more academics re-running their tests (like the papers on AIs performance in medicine & law). Advantages:
1) Helps us understand the variance among GPT-4 class models
2) Helps us understand if there are universal LLM shortcomings
Academics Should Re-test LLM Performance Across Domains
By
–