semi·news
Headlines要闻 / Research研究 / /
Research digest · Thursday, August 6, 2026 研究摘要 · 2026年8月6日 星期四

Memory movement reshapes accelerator design 数据搬运重塑加速器设计

This week's papers treat memory bandwidth, KV-cache representation, and quantization as first-order architecture choices. They span PIM-GPU systems, dedicated inference datapaths, wafer processing, and power-aware EDA. 本周论文将内存带宽、KV缓存表示和量化视为一阶架构选择。研究覆盖PIM-GPU系统、专用推理数据通路、晶圆加工以及面向功耗的EDA。

Look-back window: 7 days · 6 paper(s) 回溯窗口: 7天 · 6篇

AI Accelerators & Compute-in-Memory AI加速器与存算一体

Deltoris: Bit-sparse speculative inference for real-time VLA models Deltoris:面向实时VLA模型的位级稀疏推测推理

Zheng Liu, Zeyu Guo, et al.

arXiv:2608.04428 · 2026-08-05T04:17:03Z

Deltoris co-designs bit-level temporal sparsity and speculative inference for diffusion-based vision-language-action models. The authors report up to 34.2× speedup over mobile GPUs and 6.1× over a prior accelerator for 50–200 Hz control loops. These are proposed-accelerator evaluations, so hardware cost and model coverage remain open questions. Deltoris针对基于扩散的视觉-语言-动作模型,将位级时间稀疏性与推测推理进行协同设计。作者报告称,面向50至200Hz控制循环时,其速度最高可达移动GPU的34.2倍、此前加速器的6.1倍。这些结果基于所提出的加速器评估,硬件成本和模型覆盖范围仍是未解问题。

Design principles for heterogeneous DRAM-PIM-GPU systems 异构DRAM-PIM-GPU系统的设计原则

Corey Lammie, Hadjer Benmeziane, et al.

arXiv:2608.04169 · 2026-08-04T19:26:54Z

This study evaluates DRAM-PIM-GPU systems for decode-phase LLM inference across OPT and Mamba2 workloads. It finds static power can make dynamic-only models overstate tokens per watt by up to 3.85×, while channel count and hierarchy choices dominate outcomes. It is a design-space study rather than a new silicon implementation, but makes deployment assumptions explicit. 该研究在OPT和Mamba2工作负载上评估了解码阶段LLM推理的DRAM-PIM-GPU系统。研究发现,静态功耗可使仅考虑动态功耗的模型将每瓦token数高估最多3.85倍,而通道数和层级配置主导结果。这是一项设计空间研究而非新的芯片实现,但明确了部署假设的重要性。

AI Systems & Inference AI系统与推理

KV-cache vector quantization that preserves attention 保持注意力特性的KV缓存向量量化

Samuel Fernández-Menduiña, Amir Ziashahabi, et al.

arXiv:2608.04074 · 2026-08-04T16:10:59Z

This paper frames KV-cache quantization as transform coding whose distortion is attention-product error, not generic reconstruction error. It derives calibration-based transforms for keys and values to make two-bit-per-element compression more attention-aware. The serving benefit will depend on transform overhead and workload stability. 该论文将KV缓存量化表述为变换编码问题,其失真度量是注意力乘积误差,而非通用重构误差。它为键和值推导基于校准数据的变换,使每元素两比特的压缩更能保持注意力特性。实际服务收益将取决于变换开销和工作负载稳定性。

Signed-digit KV caches for ternary LLM inference 面向三值LLM推理的有符号数字KV缓存

Ziang Duan, Jiajun Wu, et al.

arXiv:2608.03229 · 2026-08-04T07:00:20Z

The authors store online K/V states as scaled multi-plane signed digits so attention can use the lookup-table machinery of ternary projections. The design avoids dense K/V materialization and handles incomplete cache blocks during causal decoding. The abstract does not establish end-to-end model quality or silicon efficiency. 作者提出将在线生成的K/V状态存为缩放的多平面有符号数字,使注意力计算能够复用三值投影的查找表机制。该设计避免K/V的稠密实体化,并处理自回归解码中未完成的缓存块。摘要未证明端到端模型质量或芯片效率。

Devices & Process 器件与工艺

Self-focusing control for precise 4H-SiC wafer slicing 用于精确4H-SiC晶圆切片的自聚焦控制

Dong Hee Kang, Jaeseung Lim, et al.

arXiv:2608.03814 · 2026-08-04T15:24:12Z

The work links femtosecond-laser pulse energy and processing depth to Kerr self-focusing during 4H-SiC wafer slicing. Experiments, modelling, and ray-optics simulations connect the effect to surface texture and separation stress. Better control could support thin SiC layers, although production yield is not yet demonstrated. 该工作将飞秒激光脉冲能量和加工深度与4H-SiC晶圆切片过程中的Kerr自聚焦联系起来。实验、建模和光线光学仿真将其与表面纹理及分离应力相关联。更好的工艺控制有望支持薄SiC层,但尚未证明量产良率。

EDA & Design Automation EDA与设计自动化

DiffPower: Differentiable GPU switching-power analysis DiffPower:可微分GPU开关功耗分析

Isaac Jacobson, Zheng Zhao, et al.

arXiv:2608.03778 · 2026-08-04T15:00:49Z

DiffPower converts netlists into PDK-agnostic bytecode and uses reverse-mode automatic differentiation for GPU-accelerated switching-power analysis. It reports up to 1,002× speedup over single-threaded CPU propagation and a median toggle-rate correlation of 0.96 across ten designs. Signoff accuracy will require validation against production flows. DiffPower将网表转换为与PDK无关的字节码,并使用反向模式自动微分实现GPU加速的开关功耗分析。它报告相对单线程CPU传播最高1002倍加速,并在十个设计上取得0.96的开关率相关系数中位数。签核精度仍需相对量产流程验证。

Newsletter 邮件订阅

Daily semiconductor briefing. 每日半导体简报。