AI Accelerators & Compute-in-Memory
AI加速器与存算一体
Lorenzo Leone, Philip Wiese, Gamze Islamoglu, Michael Rogenmoser, Davide Rossi, Francesco Conti, Luca Benini
academic team
arXiv:2606.02358 · 2026-06-01T15:06:09Z
CHIMERA implements a 22nm FDX microcontroller with nine RV32IMA cores, a transformer accelerator and a shared-L2 subsystem delivering 563 Gb/s aggregate bandwidth. The headline number is 3.1 TOPS/W, but the more useful contribution is the QoS-managed memory hierarchy, which targets real-time edge inference rather than a narrow benchmark. It is a good example of transformer acceleration moving into MCU power envelopes instead of staying in phone-class NPUs.
CHIMERA采用22nm FDX工艺,实现了包含9个RV32IMA核心、Transformer加速器和共享L2子系统的微控制器,L2聚合带宽为563 Gb/s。标题指标是3.1 TOPS/W,但更有价值的贡献是带QoS管理的存储层次结构,目标是真时边缘推理而非狭窄基准测试。它说明Transformer加速正在进入MCU功耗包络,而不再只停留在手机级NPU。
Ye Ke, Zhengnan Fu, Pao-Sheng Vincent Sun, An Guo, Shuai Dong, Junyi Yang, Yahan Yang, Abdelrahman B. M. Eldaly, Xin Si, Leanne Chan, Arindam Basu
academic team
arXiv:2606.01776 · 2026-06-01T06:58:57Z
The chip combines a dual-threshold delta-modulation frontend, in-memory spike detection and a bipolar LIF SNN decoder in 65nm CMOS. Its 3.53 uW per-channel power and 26x frontend compression are the practical points: implantable systems need to cut data movement before wireless or digital processing dominates the budget. The caveat is task scope, but the design is unusually complete as a sensing-to-decoding neuromorphic pipeline.
该芯片在65nm CMOS中集成双阈值delta调制前端、存内脉冲检测和双极LIF SNN解码器。每通道3.53微瓦功耗和26倍前端压缩是实际重点:植入式系统必须先削减数据搬移,否则无线传输或数字处理会吞掉功耗预算。局限在于任务范围,但作为从感测到解码的神经形态流水线,它的完整度较高。
Aradhana Mohan Parvathy, Soumendu Kumar Ghosh, Shamik Kundu, Arnab Raha, Souvik Kundu, Deepak A. Mathaikutty, Anand Raghunathan
academic team
arXiv:2606.00365 · 2026-05-29T21:07:27Z
SPARQLe attacks activation precision rather than only weight precision by splitting each activation tensor into dense low bits and sparse high bits. That matters for hardware because activations often force wider datapaths even when weights are 4-bit. The proposed hybrid format and accelerator path are relevant to inference chips trying to preserve accuracy without paying full 8-bit activation traffic.
SPARQLe不只处理权重量化,而是把每个激活张量拆成密集低位和稀疏高位。其硬件意义在于,即便权重做到4 bit,激活往往仍迫使数据通路保持更高精度。该混合格式和加速器路径适合那些希望保留精度、又不想支付完整8 bit激活流量的推理芯片。
Inference Systems & Quantization
推理系统与量化
Ekaterina Alimaskina, Darya Rudas, Denis Shveykin, Gleb Molodtsov, Pavel Vasiliev, Aleksandr Beznosikov
academic team
arXiv:2606.02011 · 2026-06-01T10:04:09Z
The paper shows that 2-bit reasoning-model inference can lose speedups because unstable traces inflate the token count. The useful insight is process-level: low-bit failures include loops, delayed commitment and unclosed reasoning, not only lower answer accuracy. FP16 planning and loop rescue are lightweight controls that make quantization an inference-policy problem as well as a kernel problem.
该论文指出,2 bit推理模型不一定带来端到端加速,因为不稳定的推理轨迹会膨胀token数量。关键洞见在过程层面:低比特失效包括循环、迟迟不提交答案和未闭合推理,而不只是答案准确率下降。FP16规划和循环救援说明,量化不仅是内核问题,也是推理策略问题。
Yung-Chin Chen, Chung Peng Lee, Ze-Wei Liou, Naveen Verma
academic team
arXiv:2606.02288 · 2026-06-01T14:09:35Z
This work reframes activation spikes as structural vector biases tied to attention-sink and value-drain mechanisms. The proposed INSERTQUANT clamps spikes and restores their function with precomputed templates, targeting a common reason low-bit activation quantization breaks. The hardware relevance is direct: taming activation dynamic range is one of the cleanest routes to narrower datapaths.
这项工作把激活尖峰重新解释为与attention sink和value drain机制相关的结构性向量偏置。其INSERTQUANT方法先钳制尖峰,再用预计算模板恢复其功能,瞄准低比特激活量化失效的常见原因。硬件相关性很直接:控制激活动态范围,是收窄数据通路最清晰的路径之一。
Haodong Wang, Junjie Liu, Zicong Hong, Qianli Liu, Jian Lin, Song Guo, Xu Chen
academic team
arXiv:2606.01556 · 2026-06-01T02:02:12Z
TwinQuant learns quantization-friendly decomposed subspaces instead of relying on energy-minimizing decompositions that may not minimize post-quantization error. The paper also describes a fused dual-component kernel, which keeps the work relevant to deployed inference rather than only offline compression. The main caveat is added optimization complexity, but the direction matches the industry's need for 4-bit accuracy without separate slow paths.
TwinQuant学习对量化友好的分解子空间,而不是依赖未必最小化量化后误差的能量最小分解。论文还描述了融合双组件内核,使其更接近部署推理而非单纯离线压缩。主要代价是优化复杂度上升,但方向符合行业对4 bit精度且不引入慢路径的需求。
Devices & Photonics
器件与光子
Yicheng Du, Meng Chao, Xiuyou Han, X. Su, Xuan Li, Changjun Liu, Guanchao Wang, Mingshan Zhao
academic team
Journal of Lightwave Technology · 2026-06-01
The paper demonstrates leakage-interference cancellation for CW radar using a silicon photonics integrated chip and spectrum predistortion. Reported cancellation depths of roughly 40 dB over 300 MHz and 37 dB over 1 GHz show why photonic RF processing keeps appearing in sensing and defense-adjacent systems. The device angle is practical: delay and amplitude control move into the optical domain, reducing pressure on purely electronic cancellation chains.
该论文用硅光集成芯片和频谱预失真实现连续波雷达的泄漏干扰抵消。约40 dB/300 MHz和37 dB/1 GHz的抵消深度,解释了为什么光子RF处理持续出现在传感和防务相邻系统中。其器件意义较实际:延迟与幅度控制转移到光域,降低了纯电子抵消链路的压力。
Lingfeng Wu, Zhonghao Zhou, Weilong Ma, Haohua Wang, Ziliang Ruan, Shiqing Gao, Zhishan Huang, Lu Qi, Jie Liu, Jing Feng, Changjian Guo, Dapeng Liu
academic team
Laser & Photonics Reviews · 2026-05-25
This work reports trench-based die-to-wafer bonding of thin-film lithium niobate onto a completed active silicon photonics platform. The point is process sequencing: TFLN is introduced after the CMOS-compatible silicon photonics flow, avoiding some of the incompatibility that has kept high-performance modulators separate from active Si photonics. That is directly relevant to optical transceivers for AI clusters, where bandwidth density and power are now system bottlenecks.
这项工作报道了通过沟槽式die-to-wafer键合,将薄膜铌酸锂集成到已完成的有源硅光平台上。关键在工艺顺序:TFLN在CMOS兼容硅光流程完成后引入,绕开了部分让高性能调制器与有源硅光分离的工艺不兼容问题。这与AI集群光收发器直接相关,因为带宽密度和功耗已成为系统瓶颈。
Qianhou Qu, Sheng Lu, Sungyong Jung, Qilian Liang, Chenyun Pan
academic team
arXiv:2605.31141 · 2026-05-29T10:49:47Z
The authors implement a 1T1R memristive crossbar and integrate-and-fire neuron in SkyWater SKY130 for bio-inspired interception tasks. The reported 10.67 pJ/spike neuron and 96% interception success rate make it more concrete than many SNN papers that stop at algorithmic simulation. The caveat is narrow workload scope, but it is a useful silicon-adjacent data point for in-memory neuromorphic designs.
作者在SkyWater SKY130中实现1T1R忆阻交叉阵列和积分发放神经元,用于仿生拦截任务。10.67 pJ/spike的神经元和96%拦截成功率,使其比许多停留在算法仿真的SNN论文更具体。局限是工作负载较窄,但它为存内神经形态设计提供了有价值的近硅片数据点。