Deltoris: Bit-sparse speculative inference for real-time VLA models Deltoris:面向实时VLA模型的位级稀疏推测推理
arXiv:2608.04428 · 2026-08-05T04:17:03Z
Deltoris co-designs bit-level temporal sparsity and speculative inference for diffusion-based vision-language-action models. The authors report up to 34.2× speedup over mobile GPUs and 6.1× over a prior accelerator for 50–200 Hz control loops. These are proposed-accelerator evaluations, so hardware cost and model coverage remain open questions. Deltoris针对基于扩散的视觉-语言-动作模型,将位级时间稀疏性与推测推理进行协同设计。作者报告称,面向50至200Hz控制循环时,其速度最高可达移动GPU的34.2倍、此前加速器的6.1倍。这些结果基于所提出的加速器评估,硬件成本和模型覆盖范围仍是未解问题。