
Claude Sonnet 3.7 evals are here too A huge jump on the SWE bench

By
–

Claude Sonnet 3.7 evals are here too A huge jump on the SWE bench
By
–
yes, also known as "new Claude 3.5 Sonnet", which I always found super confusing.
By
–
oh yes, of course they are limited. They were a bit late to the party. They are also not known for AI/ML work (see Siri being notoriously bad). That being said, I doubt they will have competitive large-scale LLMs. They don't have the server capacities for that at that user

By
–
Yeah. To be fair, every major company does the "safer" marketing. Yet, as a end-user the flag ship LLMs feel little different on the "safety" spectrum. Just an observation.
By
–
Maybe to add a bit more context. What I mean is, based on the marketing, Claude is safer than ChatGPT because it has more content moderation guardrails, and Grok is safer than ChatGPT because it has fewer guardrails? Please make it make sense
By
–
but they do have their own LLMs? https://
machinelearning.apple.com/research/intro
ducing-apple-foundation-models
…
By
–
Maybe this is a good opportunity to ask to ELI5: how are Claude and Grok safer than ChatGPT? Or taking a step back: has "safer than" (not "safety" per so) become just a marketing term?
By
–
It's going to be an eventful week! And I am glad that it's not called "new 'new Claude 3.5 Sonnet' "
By
–
Ouch. But to be fair, this could be similar to DeepSeek mistakenly identifying itself as a ChatGPT model. Likely a result of "poisoned" training data from the internet. In other words, it might just mean it's due to insufficient data filtering & system prompt design.
By
–
Tomorrow, Feb 25th at 8am PT/11am ET — join @codeSTACKr (
@MongoDB
) & @Hacubu (
@langchain
) to explore LangGraph.js + MongoDB for AI agents. Learn to: Integrate LangGraph.js and MongoDB Build a controllable AI agent with LangGraph.js Persist conversation state