No, generally it isn’t (although I’ve had some weird outliers where it outperforms Opus) I just prompt so much that I’ll use Sonnet for less important stuff so that I don’t max out my Opus usage
LLMS
-
GPT-4o for Analysis, Claude Sonnet for Creative Tasks
By
–
My list is the basically the same although I use GPT-4o for my 'analytical' workhorse and Claude Sonnet for stuff that's more creative https://t.co/tBgnIJ4DXP
— Rob Lennon 🗯 | AI Whisperer (@thatroblennon) 26 mai 2024My list is the basically the same although I use GPT-4o for my 'analytical' workhorse and Claude Sonnet for stuff that's more creative
-

Models Show Refusal Behavior on Harmless Instructions
By
–
Yes they also show it in the blog post. In this figure, they've added the refusal direction and you see models refusing harmless instructions
-
DeepSeek-Prover: AI Model Generates Lean 4 Mathematical Proofs
By
–
9/ DeepSeek-Prover – introduces an approach to generate Lean 4 proof data from high-school and undergraduate-level mathematical competition problems; it uses the synthetic data, comprising of 8 million formal statements and proofs, to fine-tune a DeepSeekMath 7B model…
-

Efficient Multimodal LLMs: Survey of Structures and Applications
By
–
10/ Efficient Multimodal LLMs – provides a comprehensive and systematic survey of the current state of efficient multimodal large language models; discusses efficient structures and strategies, applications, limitations, and promising future directions.
-

INDUS: Comprehensive LLM Suite for Scientific Earth and Planetary Research
By
–
8/ Scientific Applications of LLMs – presents INDUS, a comprehensive suite of LLMs for Earth science, biology, physics, planetary sciences, and more; includes an encoder model, embedding model, and small distilled models.
-

Layer-Condensed KV Cache Optimizes LLM Inference Efficiency
By
–
6/ Efficient LLM Inference – proposes a layer-condensed KV cache to achieve efficient inference in LLMs; only computes and caches the key-values (KVs) of a small number of layers which leads to saving memory consumption and improved inference throughput.
-

Guide for Evaluating Large Language Models with Open-Source Library
By
–
7/ Guide for Evaluating LLMs – provides guidance and lessons for evaluating large language models; discusses challenges and best practices, along with the introduction of an open-source library for evaluating LLMs.
-

Hierarchical Reasoning Aggregation Framework for LLM Answer Selection
By
–
4/ Enhancing Answer Selection in LLMs – proposes a hierarchical reasoning aggregation framework for improving the reasoning capabilities of LLMs; the approach selects answers based on the evaluation of reasoning chains.
-
New Abliterated Models Released for Larger AI Systems
By
–
They also released a collection of abliterated models. I recommend giving it a try (especially for larger models).