Which LLMs hallucinate the most? The least? A leaderboard for how often these models produce hallucinations when summarizing a document: https://
bit.ly/4gyXsDG
LLMS
-

LLM Hallucination Leaderboard: Which Models Perform Best
By
–
-
AI Engineer 2024: GitHub Integration, Feature Building, Autonomous Bug Fixing
By
–
As we near the end of the year, here is where AI Engineer is today: •It can connect to your GitHub, connect to your codebase, and make PRs
•Build simple features without supervision
•Examine stack traces and fix bugs autonomously 10-15% of the time
•Combines LLMs like Sonnet -
Big Bench Audio Dataset Released by Research Team
By
–
and.. they also release the Big Bench Audio dataset:
-

Speech-to-Speech Models: Cascaded Approaches Still Outperform End-to-End
By
–
We're so early! even the current SoTA speech-to-speech models are so dumb compared to cascaded speech-to-speech (ASR + LLM + TTS) There's loads of low hanging fruits to pick still! Looking forward for open models to climb up this benchmark in 2025
-

AI Builders Virtual Summit: Master LLMs, Agents, RAG January 2025
By
–
Boost your skills at the #AI Builders Virtual Summit (Jan 15–Feb 8, 2025)! Master #LLMs, #AI Agents, #RAG & more with 40+ experts, 30+ hours of hands-on training, & a community of 800+ AI pros. Register now: https://
summit.ai @_odsc #ODSC #DataScience #DataScientist -

Simple Framework for Understanding AI Agents and Their Architecture
By
–
From @AbacusAI at https://
abacus.ai/ai_agents A simple framework for understanding agents: • Connect to any data source or vector store
• Use a code execution engine
• Orchestrate with other ML models
• LLM-agnostic for flexibility
• Designed for chat or task-based operations -
ModernBERT Release: Technical Paper Code Weights Apache Licensed
By
–
it's not everyday that people drop such a deep technical paper full of insights, codebase and model weights (all apache 2.0 licensed) congratulations again on the release! eagerly waiting for ModernBERT-Huge
-
User compares AI model capabilities
By
–
Have you seen it do anything o1-mini can’t yet? Seems weaker on the handful of tasks I’ve tried so far
-

LLMs Work Well Because Humans Are LLMs
By
–
Si les LLM fonctionnent si bien, c'est parce que nous sommes des LLM.