10/ AI Agents That Matter – analyzes current agent evaluation practices and reveals shortcomings that potentially hinder real-world application; proposes an implementation that jointly optimizes cost and accuracy and a framework to avoid overfitting agents.
@dair_ai
-
Ctrl-G Framework: Logical Control for LLM Generation
By
–
7/ Adaptable Logical Control for LLMs – presents the Ctrl-G framework to facilitate control of LLM generations that reliably follow logical constraints; it combines LLMs and Hidden Markow Models to enable following logical constraints (represented as deterministic finite
-
Synthetic Data Effects on LLM Bias and Internal Attributes
By
–
8/ LLM See, LLM Do – closely investigates the effects and effectiveness of synthetic data and how it shapes a model’s internal biases, calibration, attributes, and preferences; finds that LLMs are sensitive towards certain attributes even when the synthetic data prompts appear
-

SummHay: New Benchmark for Haystack Summarization and Source Citation
By
–
9/ Summary of a Haystack – proposes a new task, SummHay, to test a model’s ability to process a Haystack and generate a summary that identifies the relevant insights and cites the source documents.
-

OpenAutoEncoder-Agentless Solves 27.3% GitHub Issues
By
–
6/ Agentless – introduces OpenAutoEncoder-Agentless which offers an agentless system that solves 27.3% GitHub issues on SWE-bench Lite; claims to outperform all other open-source AI-powered software engineering agents.
-
Self-Evaluation Defends LLMs Against Adversarial Attacks
By
–
5/ Self-Evaluation as a Defense Against Adversarial Attacks on LLMs – proposes the use of self-evaluation to defend against adversarial attacks; uses a pre-trained LLM to build defense which is more effective than fine-tuned models, dedicated safety LLMs, and enterprise
-

RAG Best Practices: Multimodal Retrieval and Performance Optimization
By
–
3/ Searching for Best Practices in RAG – shows the best practices for building effective RAG workflows; proposes strategies that focus on performance and efficiency, including emerging multimodal retrieval techniques.
-
Scaling Synthetic Data Creation with One Billion Diverse Personas
By
–
4/ Scaling Synthetic Data Creation – proposes 1 billion diverse personas to facilitate the creation of diverse synthetic data for different scenarios; uses a novel persona-driven data synthesis methodology to generate diverse and distinct data covering a wide range of
-
CriticGPT: New Model Critiques ChatGPT Responses
By
–
2/ CriticGPT – a new model based on GPT-4 to help write critiques for responses generated by ChatGPT; trained using RLHF using a large number of inputs that contained mistakes for which it had to critique.
-
Open-Sora: Open-Source Video Generation Model Supports Image-to-Video
By
–
9/ Open-Sora – an open-source video generation model that can generate 16-second 720p videos; it’s a 1.1B parameter model trained on more than 30m data and now supports image-to-video.https://t.co/eZO4A3uf2e
— DAIR.AI (@dair_ai) 23 juin 20249/ Open-Sora – an open-source video generation model that can generate 16-second 720p videos; it’s a 1.1B parameter model trained on more than 30m data and now supports image-to-video.