March 27, 2026

IndexCache, a new sparse attention optimizer, delivers 1.82x faster inference on long-context AI models

Modern office space with glass walls and light decor
Deliberate Directions / Unsplash

Processing 200,000 tokens through a large language model is expensive and slow: the longer the context, the faster the costs spiral. Researchers at Tsinghua University and Z.ai have built a technique called IndexCache that cuts up to 75% of the redun...