This paper is about LLMs solving a specific task, but it is also about the difficulty of figuring out what LLMs do well & why Most configurations of GPT-4 failed to solve the problem, but one robustly did, for reasons that are hard to know. LLMs are weird https://
arxiv.org/pdf/2403.15371
.pdf
…
LLMS
-

GPT-4 Task Performance Variability: Why Some Configurations Succeed
By
–
-

GPT-4 Achieves Human-Level Performance in Data Analysis
By
–
Two comparisons of data analysts to Code Interpreter: "Experimental results show that GPT-4 can achieve comparable performance to humans" https://
arxiv.org/pdf/2305.15038
.pdf
… GPT-4 scores over 90% on exams, the data science field is “on the verge of a paradigm shift” https://
arxiv.org/pdf/2307.02792
v2.pdf
… -

GPT-3’s peculiar ASCII art generation limitations discussed
By
–

GPT-3 used to struggle making any coherent ASCII art except one specific piece it sometimes gave in reply to request for art of any subject, a drawing of a male bodybuilder in shorts signed by jgs:
-
Alignment Tax: RLHF’s Impact on NLP Performance Discussed
By
–
It was never a secret. The alignment tax (RLHF hurts perf on NLP benchmarks) is mentioned in the InstructGPT paper Jan 2022. More noticed after Mysteries of Mode Collapse Nov 2022. (My Mask joke was post-shoggoth; ppl hated RLHF well before that)
-

GPT-4 Vision Medical Scan Analysis: Limitations and Accuracy
By
–
I see a lot of examples of people feeding medical scans into the AI to get results. There is no Claude 3 evaluation I have seen, but tests of GPT-4’s vision capabilities show that it makes a lot of mistakes reading scans. Interestingly, it does quite well on text-based tasks
-
Sakana AI Releases Evolutionary Model Merge Automation Technique
By
–
Sakana AI has announced Evolutionary Model Merge, a technique that automates and advances model merging. Models and demos using this method have been released. Give them a try!
-

Optimizing prompts for blog posts using different LLM models
By
–
Working on a better prompt for blog posts on TestingCatalog. Currently using @hunchtools canvas to play around with outputs provided by different models (gpt4 turbo, claude3 opus, gemini pro). Quite useful to compare how these models perform and optimise the prompt to make it
-

AI Trends Dashboard: Multi-Page Dash App with LangChain
By
–
Trends in AI – Plotly Dash App with LLM A multi-page Python Dash app highlighting the growing role of AI in today's world. It incorporates a LangChain Pandas Agent to empower users with deeper insights into the datasets. YouTube: https://
youtube.com/watch?v=t3O-0m
zLJzI
… -
GPT-4 Performance Decline: Users Switching to Claude 3 and Mixtral
By
–
GPT-4 feels worse than it was a few months ago. No fancy benchmarks, I just “UGH” way more often these days. Most folks I know are using Claude 3 or Mixtral 8x7B instead. BUT this is not a sign to ditch @OpenAI
. To me, this is because OpenAI is focused on the next model. -
ChatGPT Performance Optimization: Prompting Techniques and Strategies
By
–
Some options for… When chatgpt doesn’t want to do it:
– demand: “try harder”
– resourceful: “find a way”
– continuation: “you’ve done it before”
– guilt: “you promised me you would do it” When you want better chatgpt performance:
– show examples
– try a better model
– ask it