This blog post is really cool: to understand everyday Transformers optimizations like KV cache, FlashAttention or PagedAttention: https://
astralord.github.io/posts/transfor
mer-inference-optimization-toolset/
… Image below is the interactive visualization for KB cache!
LLMS
-

Everyday Transformers Optimizations: KV Cache, FlashAttention, PagedAttention
By
–
-
Google NotebookLM Transforms Books Into Podcasts and Study Guides
By
–
This is the most impressive and instantly useful #AI demo I’ve seen so far!
— Pascal Bornet (@pascal_bornet) 2 octobre 2024
It is called Google's NotebookLM.
I uploaded my entire new book "IRREPLACEABLE" and it transformed it into a podcast, study guide, FAQ, timeline, and an accurate chatbot that references specific parts… pic.twitter.com/2o4DSat7ShThis is the most impressive and instantly useful #AI demo I’ve seen so far! It is called Google's NotebookLM. I uploaded my entire new book "IRREPLACEABLE" and it transformed it into a podcast, study guide, FAQ, timeline, and an accurate chatbot that references specific parts
-

Daily Papers Project Translates and Summarizes Huggingface Research into Korean
By
–
daily_papers_ko This project aims to automatically translate and summarize Huggingface's daily papers into Korean using ChatGPT.
-

GraphRAG-UI: User-Friendly Interface for RAG Text Indexing
By
–
GraphRAG-UI
— AK (@_akhaliq) 2 octobre 2024
GraphRAG-UI is a user-friendly interface for GraphRAG, a powerful tool that uses the Retrieval-Augmented Generation (RAG) approach to index and query large text data. This project supports the latest version graphrag-0.3.3 and aims to provide a convenient management… pic.twitter.com/davBb07LcBGraphRAG-UI GraphRAG-UI is a user-friendly interface for GraphRAG, a powerful tool that uses the Retrieval-Augmented Generation (RAG) approach to index and query large text data. This project supports the latest version graphrag-0.3.3 and aims to provide a convenient management
-

vLLM 0.6.2 now supports SolarPro
By
–
Super happy that @vllm_project 0.6.2 now supports #SolarPro. Check it out! https://
github.com/vllm-project/v
llm/releases
… -

Rethinking ML Generalization and Scaling Paradigms
By
–
Rethinking Conventional Wisdom in Machine Learning: From Generalization to Scaling https://
bit.ly/4dtpKgP
#AI #MachineLearning #DeepLearning #LLMs #DataScience -

Real-time Speech to Speech with Whisper-Turbo
By
–
real-time speech to speech with whisper-turbo: https://t.co/8qDGac419x
— AK (@_akhaliq) 1 octobre 2024real-time speech to speech with whisper-turbo:
-

Real-time speech to speech with Whisper Turbo
By
–
real-time speech to speech with whisper turbo https://t.co/8qDGac419x
— AK (@_akhaliq) 1 octobre 2024real-time speech to speech with whisper turbo
-

Real-time Speech to Speech with Whisper-Turbo Technology
By
–
real-time speech to speech with whisper-turbo: https://t.co/8qDGac419x
— AK (@_akhaliq) 1 octobre 2024real-time speech to speech with whisper-turbo:
-
CGPO: Mixture of Judges Outperforms RLHF Approaches
By
–
New paper from GenAI and Meta FAIR. CGPO uses Mixture of Judges and consistently outperforms SOTA RLHF approaches across various tasks. More details and key results in the full thread