semi·news
Headlines要闻 / Research研究 / /
Research digest · Saturday, August 8, 2026 研究摘要 · 2026年8月8日 星期六

Hardware brings sparsity and resilience closer 硬件把稀疏性与韧性推向更近处

This week’s papers work on the data paths around neural computation: sparse die-to-die links, asynchronous control, compute-in-memory precision, and error resilience. Several contributions are measured designs, making their energy and area claims more actionable than architecture-only proposals. 本周论文聚焦神经计算周边的数据路径:稀疏的裸片间互连、异步控制、存内计算精度和错误韧性。多项工作给出了实测设计,因此其能耗和面积主张比纯架构提案更具可操作性。

Look-back window: 7 days · 6 paper(s) 回溯窗口: 7天 · 6篇

Quantum Computing 量子计算

Dynamic quantum circuits on a hybrid superconducting qubit-cavity processor 混合超导量子比特-腔体处理器上的动态量子电路

Hongbo Wu, Ling Hu, et al.

arXiv:2608.04780 · 2026-08-05

This experiment uses mid-circuit measurement, reset, reuse, and feed-forward on a superconducting qubit-cavity processor. It reports 82% average success on a 10-bit Bernstein-Vazirani task, sub-10^-3 error in 8-bit phase estimation, and a dynamic implementation of Shor factoring 15 with over 99.8% squared statistical overlap. The result makes a concrete case for dynamic circuits as a way to trade control complexity for fewer physical resources. 该实验在超导量子比特-腔体处理器上使用了线路中测量、复位、复用和前馈控制。它在10比特Bernstein-Vazirani任务上报告了82%的平均成功率,在8比特相位估计中实现低于10^-3的误差,并以超过99.8%的平方统计重叠度动态实现了对15的Shor分解。结果为动态电路提供了具体证据:可用更复杂的控制来换取更少的物理资源。

Circuits & Reliability 电路与可靠性

Permanent-fault analysis for asynchronous neuromorphic NoCs 异步神经形态片上网络的永久故障分析

Giuseppe Chessa, Yuyi Yang, et al.

Proceedings of the International Conference on Neuromorphic Systems · 2026-08-04

The authors couple a gate-level asynchronous NoC model with an SNN simulator to trace permanent routing faults into application behavior. Localized stuck-at faults caused accuracy losses of up to 10% in navigation and 36% in a keyword-spotting case. The cross-layer result shows that spike-routing reliability can dominate even when neuron and memory blocks are sound. 作者将门级异步NoC模型与SNN模拟器耦合,用于追踪永久路由故障如何传导至应用行为。局部卡死故障在导航任务中最多造成10%的准确率损失,在关键词识别案例中最多造成36%的损失。该跨层结果表明,即使神经元和存储模块正常,脉冲路由可靠性也可能成为主导因素。

Reed-Solomon decoding tailored to HBM3 error correction 面向HBM3纠错的Reed-Solomon解码架构

Jaehoon Kwon, Jeongmin Kim, et al.

IEEE Transactions on Very Large Scale Integration (VLSI) Systems · 2026-08-01

The paper optimizes a decoder for HBM3’s mandatory (19,17) single-symbol-error-correcting Reed-Solomon code. By replacing costly Galois-field inverse LUTs with comparison search and constant multiplications, it reports up to 97% less area and 81% higher speed in 28-nm synthesis. Compact on-memory error correction becomes more consequential as HBM stacks grow in capacity and interface width. 论文针对HBM3强制采用的(19,17)单符号纠错Reed-Solomon码优化了解码器。通过以比较搜索和常数乘法替代昂贵的伽罗瓦域逆元查找表,28nm综合结果显示其面积最多降低97%,速度最多提高81%。随着HBM堆叠的容量和接口宽度提升,紧凑的存储器内纠错将愈发重要。

AI Accelerators & Compute-in-Memory AI加速器与存算一体

Temporal-sparse die-to-die communication for heterogeneous neuromorphic systems 面向异构神经形态系统的时域稀疏裸片间通信

Joshua Nardone, Rui-Jie Zhu, et al.

Proceedings of the International Conference on Neuromorphic Systems · 2026-08-04

The design keeps dense ANN computation within a die but uses learnably sparse spiking layers at bandwidth-limited die boundaries. That partitioning aims to reduce inter-die traffic without asking an entire network to adopt spiking computation. System benefit will depend on the sparsity achieved after training, but it is a direct response to packaging bandwidth limits. 该设计将密集ANN计算保留在裸片内部,但在带宽受限的裸片边界使用可学习稀疏的脉冲层。这样的划分旨在减少裸片间流量,而不要求整个网络采用脉冲计算。系统收益仍取决于训练后实际达到的稀疏度,但它直接回应了封装带宽限制。

SACIM: an asynchronous charge-domain CIM for CNN inference SACIM:用于CNN推理的异步电荷域存内计算

Tingran Chen, Xiaoya Wang, et al.

IEEE Transactions on Circuits and Systems - II - Express Briefs · 2026-08-01

SACIM combines signed-weight bit cells, a clock-free controller, and a pulse-based ADC to improve latency and parallelism in charge-domain analog CIM. Fabricated in 40 nm, the prototype reports 655.78 TOPS/W peak efficiency, 1,180.4 GOPS throughput, and 5.1 μW static power. Mapping accuracy and peripheral overhead at larger array and model sizes remain the important scaling questions. SACIM结合了有符号权重位单元、无时钟控制器和脉冲式ADC,以改善电荷域模拟CIM的延迟和并行度。该40nm原型报告峰值能效655.78 TOPS/W、吞吐量1,180.4 GOPS以及5.1微瓦静态功耗。在更大阵列和模型规模下的映射精度与外围开销仍是关键扩展问题。

A 28-nm floating-point SRAM CIM macro for edge AI 面向边缘AI的28nm浮点SRAM存内计算宏

Yiyang Yuan, Zhongyi Sun, et al.

IEEE Journal of Solid-State Circuits · 2026-08-01

This 192-Kb, 28-nm macro combines digital and analog CIM with an outer-product FP/INT dataflow and a residual ADC intended to cut conversion overhead. It reports 72.12 TFLOPS/W peak energy efficiency and only 0.05% ResNet-18 accuracy loss on CIFAR-100 using BF16 inputs, weights, and outputs. The work addresses the CIM challenge of gaining efficiency without abandoning floating-point-compatible accuracy. 这枚192Kb、28nm宏将数字与模拟CIM结合,采用外积式FP/INT数据流和残差ADC以降低转换开销。它报告峰值能效72.12 TFLOPS/W,在CIFAR-100上使用BF16输入、权重和输出运行ResNet-18时准确率仅损失0.05%。该工作处理了CIM的一项关键矛盾:在不牺牲兼容浮点精度的前提下提升能效。

Newsletter 邮件订阅

Daily semiconductor briefing. 每日半导体简报。