Windowed-MTP Removes Full-Context Draft KV Reads Windowed-MTP消除草稿模型的全上下文KV读取
arXiv:2607.21535 · 2026-07-23T17:21:44Z
Windowed-MTP applies a sliding attention window and sink only to the multi-token-prediction draft head, preventing draft cost from growing with a million-token KV cache. Full-attention verification remains unchanged, so the method can alter proposed tokens but not the target model's accepted output. It is a training-free systems optimization; the preprint still needs independent evaluation across model families and serving stacks. Windowed-MTP仅在多Token预测草稿头中应用滑动注意力窗口和attention sink,避免草稿成本随百万Token KV缓存线性增长。全注意力验证保持不变,因此该方法可以改变候选Token,却不会改变目标模型最终接受的输出。这是一项免训练的系统优化,但仍需在不同模型家族和推理栈上进行独立评估。