As far as I know, you are not getting the base model through APIs, there is definitely fine-tuning and many guardrail instructions that are likely the result of prompting. Refusals to mess with copyright work, for example.
@emollick
-
Why Humans Prefer Anthropomorphic AI Over Formal Assistants
By
–
It is similar to how Sydney/Bing caused a huge stir. We are very much more freaked out and impressed by an AI that is allowed to act human than one that insists it is just an assistant.
-
Claude 3 Performance: Design Over Actual Model Capability?
By
–
It is really hard to know how much of the Twitter reaction to the "smarts" of Claude 3 is due to the fact that Claude's system prompt/design is pushing the AI to act more human. I am not sure the model is actually better than GPT-4, but it more willing to play along with users.
-

AI Predicts Neuroscience Experiment Outcomes Better Than Experts
By
–
Interesting new result on how AI can help advance scientific research by predicting in advanced which neuroscience experiments would yield positive findings better than human experts could And they only used GPT-3.5 class models & found fine-tuning helped https://
arxiv.org/abs/2403.03230 -

Inflection Pi GPT-4 Class Upgrade: AI Designed as Your Best Friend
By
–
I think we should be talking about Inflection’s Pi more. They released a near GPT-4 class upgrade to their AI this week in service of a very different vision of AI, one designed to be your best friend rather than an assistant. And it seems to be working, for better or worse.
-
OpenAI vs Apple: Different AI Model Deployment Strategies
By
–
Feels like OpenAI is going for the first model. I suspect Apple might do the second (your local Siri on your phone connects to SIRIAC when it can’t help you).
-
Two Competing Visions for AI Agent Organization Architecture
By
–
I see two competing visions of the future of organizing AI agents: In one, you talk to the smartest AI model first, and it decides what to delegate to dumber (but cheaper) models. In the other you start with a dumb local model & it calls for help from bigger models when needed.
-

Claude 3 Excels at Long Context Needle-in-Haystack Test
By
–
Claude 3 does a good job with the needle-in-a-Great-Gatsby test, where I load the entire text of the novel with a couple alterations into the context window. Much better than Claude 2.1 (no hallucinations!), not quite as good as Gemini (not quite as insightful about content).
-

AI Opinions and Biases: Understanding What Artificial Intelligence Does Well
By
–
I think the obsession with asking AI for their opinions is not a particularly useful way to understand what AIs do well or even their underlying biases… But if you do care about their opinions, AIs are VERY excited about my book, out April 2. Pre-order: https://
penguinrandomhouse.com/books/741805/c
o-intelligence-by-ethan-mollick/?ref=PRH410E2C567AF
… -

Testing LLM Expertise: Building Collective Evaluation Framework
By
–
As our research shows, AI has a "jagged frontier" – it is good at some tasks, bad at others in unpredictable ways. Testing AI is hard. We need a collective of experts in various fields agree to test each LLM generation as a way of seeing if AI reaches expert level in that area.