SageAttention: Accurate 8-Bit Attention for Plug-and-play Inference Acceleration https://
arxiv.org/abs/2410.02367 https://
github.com/thu-ml/SageAtt
ention
…… This work from Tsinghua University proposes SageAttention, a highly efficient and accurate quantization method for attention.
@jiqizhixin
-
SageAttention: 8-Bit Quantization for Efficient Attention
By
–
-
SageAttention Outperforms FlashAttention2 and Xformers
By
–
The OPS (operations per second) of our approach outperforms FlashAttention2 and xformers by about 2.1 times and 2.7 times, respectively. SageAttention also achieves superior accuracy performance over FlashAttention3.
-

RDT-1B: Largest Diffusion Foundation Model for Robot Manipulation
By
–
Great work from Tsinghua University! RDT-1B, the largest diffusion-based foundation model for robotic manipulation.
RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation https://
rdt-robotics.github.io/rdt-robotics https://
arxiv.org/pdf/2410.07864 -
Aligning Large Language Models with Human Values and Intentions
By
–
align any modality large models (any-to-any models), including LLMs, VLMs, and others, with human intentions and values. More details about the definition and milestones of alignment for Large Models can be found in AI Alignmen
-
PKU releases Align Anything framework for AI alignment
By
–
PKU opensourced 「Align Anything」framework https://
github.com/PKU-Alignment/
align-anything
… -
SORA-like Models Bridge Academic Research and Industry Video Generation
By
–
This report studies a series of SORA-like models to bridge the gap between academic research and industry practice, providing a more profound analysis of recent video generation advancements.
-

SORA-like Video Generation Models: Preliminary Evaluation Framework
By
–
The Dawn of Generation: Preliminary Explorations with SORA-like Models https://
arxiv.org/pdf/2410.05227 https://
ailab-cvc.github.io/VideoGen-Eval/ -
LLaVA-Critic: Open-Source Multimodal Model Evaluator
By
–
LLaVA-Critic: Learning to Evaluate Multimodal Models https://
arxiv.org/abs/2410.02712 https://
llava-vl.github.io/blog/2024-10-0
3-llava-critic/
…
the first open-source large multimodal model (LMM) designed as a generalist evaluator to assess performance across a wide range of multimodal tasks. -
Open-source Framework Enhances LLM Reasoning Capabilities
By
–
an open-source framework designed to integrate key components for enhancing the reasoning capabilities of large language models (LLMs)
-
OpenR: Open Source Framework for Advanced LLM Reasoning
By
–
OpenR: An Open Source Framework for Advanced Reasoning with Large Language Models https://
github.com/openreasoner/o
penr/blob/main/reports/OpenR-Wang.pdf
… https://
github.com/openreasoner/o
penr
… https://
openreasoner.github.io
#OpenAI o1 #LLM Reasoning #UCL