OpenAI’s o1-preview and o1-mini now top Scale’s SEAL Leaderboards! On the Scale blog, I recap benchmarks for the new o1 models and share some early impressions from my time using them. (Link in comment!)
LLMS
-

Try Gemini-1.5-Flash-8B Experimental Chatbot on Hugging Face
By
–
try out Gemini-1.5-flash-8b-exp-0924 chatbot on @huggingface
: https://
huggingface.co/spaces/akhaliq
/gemini-1.5-flash-8b-exp-0924
… -
LLM Development Enters Innovation Phase: Path to Proto-AGI
By
–
We are entering the 3rd phase of LLM Development. 1st phase was early tinkering, Transformer to GPT-3
2nd phase was scaling
3rd phase is an innovation phase: what breakthroughs beyond o1 get us to a new proto-AGI paradigm https://
x.com/a16z/status/18
38594324956864644/video/1
… -
Can LLM-as-Judge Replace Human Evaluators?
By
–
Human preference studies continue to be the gold standard when it comes to evaluating #LLMs But as models become more sophisticated, we asked ourselves: can LLM-as-a-Judge replace human evaluators? Read more in our blog https://
sambanova.ai/blog/can-llama
-405b-outperform-gpt4
… #GPT4 -
User inquiry regarding AI voice mode generation capabilities
By
–
Was it able to produce sounds in the past? Seems like now it refuses it or my voice mode is not advanced
-

Updates to AI Voice Conversation Interface and Message Controls
By
–

Also chat history now looks different for voice conversations. There is a duration next to user messages, they cannot be modified. Assistant messages now can be “replayed” instead of read aloud
-

ChatGPT rolls out new voice UI capabilities to users
By
–

Looks like a new voice UI rolling out to everyone regardless if they have access to advanced voice mode or not. ChatGPT now has $SOL voice! Which one is your favourite???
-

Cerebras Showcases Top AI Projects from PennApps Hackathon
By
–
Cerebras just finished an incredible weekend at the PennApps Hackathon! Featured below are the top 3 projects using Cerebras. 1st: Electra – Predict election outcomes accurately and quickly with 100's of LLM agents simulating political scenarios 2nd: Optica – For the blind and
-
LLM-as-Judge vs Human Evaluation: Key Differences
By
–
When it comes to evaluating #LLMs, how does a human differ from LLM-as-a-Judge? In our blog, we explore the ways in which LLM-as-a-Judge can offer an alternative to human evaluations. Read more https://
sambanova.ai/blog/judging-l
lm-judgements
… #AI -

OpenAI o1 Dominates SEAL Leaderboard Rankings
By
–
SEAL Leaderboard Update OpenAI’s o1 is dominating SEAL rankings! o1-preview is dominating across key categories:
– #1 in Agentic Tool Use (Enterprise)
– #1 in Instruction Following
– #1 in Spanish o1-mini leads the charge in Coding
