Joshua Nardone, Rui-Jie Zhu, et al.
Proceedings of the International Conference on Neuromorphic Systems · 2026-08-04
The design keeps dense ANN computation within a die but uses learnably sparse spiking layers at bandwidth-limited die boundaries. That partitioning aims to reduce inter-die traffic without asking an entire network to adopt spiking computation. System benefit will depend on the sparsity achieved after training, but it is a direct response to packaging bandwidth limits.
该设计将密集ANN计算保留在裸片内部,但在带宽受限的裸片边界使用可学习稀疏的脉冲层。这样的划分旨在减少裸片间流量,而不要求整个网络采用脉冲计算。系统收益仍取决于训练后实际达到的稀疏度,但它直接回应了封装带宽限制。
Tingran Chen, Xiaoya Wang, et al.
IEEE Transactions on Circuits and Systems - II - Express Briefs · 2026-08-01
SACIM combines signed-weight bit cells, a clock-free controller, and a pulse-based ADC to improve latency and parallelism in charge-domain analog CIM. Fabricated in 40 nm, the prototype reports 655.78 TOPS/W peak efficiency, 1,180.4 GOPS throughput, and 5.1 μW static power. Mapping accuracy and peripheral overhead at larger array and model sizes remain the important scaling questions.
SACIM结合了有符号权重位单元、无时钟控制器和脉冲式ADC,以改善电荷域模拟CIM的延迟和并行度。该40nm原型报告峰值能效655.78 TOPS/W、吞吐量1,180.4 GOPS以及5.1微瓦静态功耗。在更大阵列和模型规模下的映射精度与外围开销仍是关键扩展问题。
Yiyang Yuan, Zhongyi Sun, et al.
IEEE Journal of Solid-State Circuits · 2026-08-01
This 192-Kb, 28-nm macro combines digital and analog CIM with an outer-product FP/INT dataflow and a residual ADC intended to cut conversion overhead. It reports 72.12 TFLOPS/W peak energy efficiency and only 0.05% ResNet-18 accuracy loss on CIFAR-100 using BF16 inputs, weights, and outputs. The work addresses the CIM challenge of gaining efficiency without abandoning floating-point-compatible accuracy.
这枚192Kb、28nm宏将数字与模拟CIM结合,采用外积式FP/INT数据流和残差ADC以降低转换开销。它报告峰值能效72.12 TFLOPS/W,在CIFAR-100上使用BF16输入、权重和输出运行ResNet-18时准确率仅损失0.05%。该工作处理了CIM的一项关键矛盾:在不牺牲兼容浮点精度的前提下提升能效。