6). Star Attention: Efficient LLM Inference over Long Sequences – introduces Star Attention, a two-phase attention mechanism that processes long sequences by combining blockwise-local attention for context encoding with sequence-global attention for query processing and token
Star Attention: Efficient LLM Inference for Long Sequences
By
–