8/ StarCoder 2 – a family of open LLMs for code with three different sizes (3B, 7B, and 15B); the 15B model was trained on 14 trillion tokens and 600+ programming languages with a context window of 16K token and employing a fill-in-the-middle objective.
@dair_ai
-

Societal Impact of Open Foundation Models Position Paper
By
–
7/ On the Societal Impact of Open Foundation Models – a position paper with a focus on open foundation models and their impact, benefits, and risks.
-
EMO: Direct Audio-to-Video Synthesis Framework for Expressive Videos
By
–
6/ EMO – a new framework for generating expressive video by utilizing a direct audio-to-video synthesis approach; by leveraging an Audio2Video diffusion model it bypasses the need for intermediate 3D models or facial landmarks. https://
x.com/_akhaliq/statu
s/1762686465777999932?s=20
… -

LearnAct: Open-Action Learning Strategy for Language Agents
By
–
5/ LearnAct – explores open-action learning for language agents through an iterative learning strategy that creates and improves actions using Python functions.
-

Comprehensive Overview and Analysis of LLM Datasets
By
–
4/ Dataset for LLMs – a comprehensive overview (180+ pages) and analysis of LLM datasets.
-

BitNet b1.58: High-Performance 1-bit LLM Architecture
By
–
3/ The Era of 1-bit LLMs – introduces a high-performing and cost-effective 1-bit LLM variant called BitNet b1.58 where every parameter is a ternary {-1, 0, 1}.
-

Mistral Large: Multilingual LLM with Advanced Reasoning and Code
By
–
2/ Mistral Large – a new LLM with strong multilingual, reasoning, maths, and code generation capabilities.
-

Survey of Major LLM Families: GPT, Llama, PaLM
By
–
9/ Survey of LLMs – reviews three popular families of LLMs (GPT, Llama, PaLM), their characteristics, contributions, and limitations; includes a summary of capabilities and techniques developed to build and augment LLM.
-

ChemLLM: Specialized LLM Outperforms GPT-4 Chemistry Tasks
By
–
8/ ChemLLM – a dedicated LLM trained for chemistry-related tasks; claims to outperform GPT-3.5 on principal tasks such as name conversion, molecular caption, and reaction prediction; it also surpasses GPT-4 on two of these tasks.
-

LLM Agents Demonstrate Autonomous Website Hacking Capabilities
By
–
10/ LLM Agents can Hack – shows that LLM agents can automatically hack websites and perform tasks like SQL injections without human feedback or explicit knowledge about the vulnerability beforehand.
