HeadInfer: Unlocking Long-Context LLM Inference on Consumer GPUs (Million-level Tokens)
*long-context inputs require large GPU memory.
*A standard LLM like Llama-3–8B requires 207GB of GPU memory for 1 million tokens — far beyond the capabilities of consumer GPUs like the RTX
HeadInfer: Long-Context LLM Inference on Consumer GPUs
By
–