The OPS (operations per second) of our approach outperforms FlashAttention2 and xformers by about 2.1 times and 2.7 times, respectively. SageAttention also achieves superior accuracy performance over FlashAttention3.
LLMS
-

Active Statistics: Machine Learning and Deep Learning Insights
By
–
Active Statistics https://
bit.ly/3Y7NPpl
#AI #MachineLearning #DeepLearning #LLMs #DataScience -
The Limits of Fine-Tuning and the Need for AI Supervision
By
–
It still feels that with current capabilities, even if every individual will be able to fine tune the model for themselves and just for one specific task, it still will remain to be “high risk” / require supervision. In other words, ppl are expecting “delegation level 4+” while
-
OpenAI o1 Models Planning Abilities Self-Evaluation Constraints
By
–
9). On the Planning Abilities of OpenAI’s o1 Models – reports that o1-preview is particularly strong in self-evaluation and constraint-following; also mentions that these o1 models demonstrate bottlenecks in decision-making and memory management, which are more pronounced in
-

Model Kinship Strategy for Improved LLM Merging
By
–
8). Model Kinship for Merging LLMs – proposes model kinship to measure the degree of similarity between LLMs; model kinship is used to build a model merging strategy (Top-k Greedy Merging with Model Kinship) which yields better performance.
-
Janus: Unified Autoregressive Framework for Multimodal Understanding
By
–
5). Janus – proposes a unified autoregressive framework for multimodal understanding and generation; it decouples visual encoding into independent pathways and leverages a single transformer architecture to improve flexibility and performance on both visual understanding and
-

Inference Scaling Laws for Long-Context RAG Systems
By
–
6). Inference Scaling for Long-Context RAG – uses two strategies to investigate scaling laws for RAG: in-context learning and iterative prompting; RAG performance improves with the expansion of the effective context length under optimal configurations.
-
LLM Introspection Enables Self-Knowledge and System Interpretability
By
–
4). Introspection in LLMs – reports that LLMs can acquire knowledge through introspection that cannot be inferred from their training data; suggests that LLMs contain privileged information about themselves that can potentially lead to more interpretable and controllable systems.
-

Pre-order LLM Engineers Handbook and Support on GitHub
By
–
If you're interested, you can pre-order the book at this address: https://
a.co/d/6O2hhZ4 You can also check our GitHub repo made by Paul and star it to support us 🙂 https://
github.com/PacktPublishin
g/LLM-Engineers-Handbook
… -

LLM Engineer’s Handbook Becomes Number One Neural Networks Release
By
–
Really happy to see that the LLM Engineer's Handbook is #1 New Release in Neural Networks We put an insane amount of work into this book. I hope it'll help a new generation of LLM engineers and tinkerers to build production-level AI systems.