Learn/Core Concept How does sparse attention save compute? Sparse attention limits which tokens interact, cutting memory and compute while preserving useful context. It matters when serving long inputs, as in the CPU model course’s sparse-attention example, where efficient code keeps hardware costs manageable. AttentionTokenisation |