This blog post is really cool: to understand everyday Transformers optimizations like KV cache, FlashAttention or PagedAttention: https://
astralord.github.io/posts/transfor
mer-inference-optimization-toolset/
… Image below is the interactive visualization for KB cache!
SYSTEMS
-

Everyday Transformers Optimizations: KV Cache, FlashAttention, PagedAttention
By
–
-
MPMC Graph Neural Networks Improve Data Point Uniformity
By
–
MPMC uses graph neural networks to allow points to "communicate" for better uniformity.
— MIT CSAIL (@MIT_CSAIL) 1 octobre 2024
Often, the more evenly you can spread out those data points, the more accurately you can simulate complex systems. pic.twitter.com/sfiYLge84yMPMC uses graph neural networks to allow points to "communicate" for better uniformity. Often, the more evenly you can spread out those data points, the more accurately you can simulate complex systems.
-

CIOs Prioritizing DEX Gain Greater Organizational Influence
By
–
@GoIvanti study explores #DEX from various perspectives: IT professionals, executive leadership and office workers. One key finding: 3 in 4 leadership executives say that “CIOs who prioritize DEX earn greater influence with other organizational leaders.” The role of the
-
Cognitive Architecture and Prompting Strategies with LangGraph
By
–
Less about the framework, more about the prompting/cognitive architecture. I think there are a few examples of doing it with Langgraph,
-
Distributed Neural Network Training: From 1990 to Large Scale
By
–
I did an undergrad thesis on parallel training of neural networks in 1990, then detoured through HIV/AIDS forecasting, compiler optimizations, information retrieval, distributed computation and storage systems before returning to large scale distributed training of neural
-
Universe Maximizes Future Options: Strategic Betting Principle
By
–
The universe wants to maximize its future options; always bet on this.
-
Microservices Approach for Query Optimization and Reusable Functions
By
–
I’ll probably do another post on it, but with this system, I think you want a micro service inspired approach where every query gets broken down into reusable functions first and then wrapped in a function that uses those. Overtime, you should have to create less low-level
-
Computer Functionality: Meeting Basic Computational Expectations
By
–
It'll be a while before we can answer "is it good". I mean… it's a computer, and so far it seems to be correctly doing what a computer ought to do. Can't ask for much more than that!
-

Anticipatory AI System with Autonomous Function Generation
By
–
Level 3: Anticipatory A system that generates synthetic queries based on its understanding of the user, where each query is fed into the system 2 need-based system – which generates required functions as necessary. The result is a system that can autonomously generate functions
-
OpenAI’s Audacious Plan to Make AI Flow Like Electricity
By
–
Behind OpenAI’s Audacious Plan to Make #AI Flow Like Electricity https://
nytimes.com/2024/09/25/bus
iness/openai-plan-electricity.html?smid=nytcore-android-share
…