It may seem obvious, but AI quality is one of the most impactful elements of your AI app. We’ve done A/B tests of different models sizes/prompts/datasets etc., over the course of years. We’ve seen differences as drastic as 2x higher conversion rates between models that, on the
LLMS
-
GPT-4 Model Versions and Naming Clarification in 2024
By
–
in the past year: gpt 4.0: gpt-4-0314
gpt 3.9: gpt-4-0613
gpt 4.3: gpt-4-1106-preview
gpt 4.4: gpt-4-1106-vision-preview
gpt 4.2: gpt-4-0125-preview
gpt 4.5: gpt-4-turbo-2024-04-09 lets be clear what we are talking about when saying another model is "gpt4 level", because this -
Congratulating Yi Tay on Impressive Model Launch with Honest Benchmarks
By
–
Congrats @YiTayML on this launch. It is impressive that a small team can train a strong model so quickly. What I also like is that the PR is not full of unfounded hype. Just plainly states the model's benchmark scores and you can immediately try out the model yourself for free.
-

Anthropic Launches New Tool-Calling Agent Type for Developers
By
–
Anthropic Agents Last week, we added a new type of agent that uses tool calling for reliability. It's designed to work with a variety of models! Try it with @AnthropicAI
's SOTA Claude 3 for yourself: Python : https://
python.langchain.com/docs/modules/a
gents/agent_types/tool_calling/
…
JS : https://
js.langchain.com/docs/integrati
ons/chat/anthropic#agents
… -
New Model Feels Like Claude 2.1 Despite Better Benchmarks
By
–
It "feels" GPT-3.5+ class, like Claude 2.1 like so far (though its reported benchmarks are closer to Claude 3/GPT-4 class). But that is just first impressions.
-
Surprise LLM Announcement Catches AI Community Off Guard
By
–
Did anyone expect a big new LLM coming? I was surprised by this announcement.
-
OpenAI Announces Batch Product for LLM Usage
By
–
Lol last week's issue was about how most LLM usage is batch right now and openai announced a batch product today
-
ChatGPT Code Interpreter Alternatives With Internet Access
By
–
What are the best current alternatives to ChatGPT Code Interpreter? I want almost the exact same product but with the ability to access the internet to install additional packages or interact with APIs I'd prefer hosted, but still interested in hearing about good local options
-
Clear Communication Key to Effective AI Model Prompting
By
–
I wonder if part of that is that being "good st prompting" is mainly about being good at clear and explicit communication People with prompts of a less high quality may find they see bigger differences between models