semi·news
Headlines要闻 / Research研究 /
Research digest · Tuesday, June 2, 2026 研究摘要 · 2026年6月2日 星期二

Low-Bit Inference Meets Real Edge Silicon 低比特推理遇上真实边缘硅片

This week's research queue is unusually hardware-grounded: an AI-MCU with transformer acceleration, multiple low-bit inference papers that discuss kernels or failure modes, and photonic / neuromorphic device work aimed at packaging real systems rather than only simulations. 本周研究队列明显更贴近硬件:带Transformer加速器的AI-MCU,多篇讨论内核或失效模式的低比特推理论文,以及面向真实系统封装而非纯仿真的光子与神经形态器件工作。

Look-back window: 7 days · 9 paper(s) 回溯窗口: 7天 · 9篇

AI Accelerators & Compute-in-Memory AI加速器与存算一体

CHIMERA: A 3.1 TOPS/W AI-MCU with Transformer Accelerator CHIMERA:3.1 TOPS/W、带Transformer加速器的AI-MCU AI-MCUtransformeredge inference

Lorenzo Leone, Philip Wiese, Gamze Islamoglu, Michael Rogenmoser, Davide Rossi, Francesco Conti, Luca Benini

academic team

arXiv:2606.02358 · 2026-06-01T15:06:09Z

CHIMERA implements a 22nm FDX microcontroller with nine RV32IMA cores, a transformer accelerator and a shared-L2 subsystem delivering 563 Gb/s aggregate bandwidth. The headline number is 3.1 TOPS/W, but the more useful contribution is the QoS-managed memory hierarchy, which targets real-time edge inference rather than a narrow benchmark. It is a good example of transformer acceleration moving into MCU power envelopes instead of staying in phone-class NPUs. CHIMERA采用22nm FDX工艺,实现了包含9个RV32IMA核心、Transformer加速器和共享L2子系统的微控制器,L2聚合带宽为563 Gb/s。标题指标是3.1 TOPS/W,但更有价值的贡献是带QoS管理的存储层次结构,目标是真时边缘推理而非狭窄基准测试。它说明Transformer加速正在进入MCU功耗包络,而不再只停留在手机级NPU。

A 32-Channel 3.53-uW/Channel Brain-Machine Interface SoC 32通道、每通道3.53微瓦的脑机接口SoC SoCneuromorphicimplantable

Ye Ke, Zhengnan Fu, Pao-Sheng Vincent Sun, An Guo, Shuai Dong, Junyi Yang, Yahan Yang, Abdelrahman B. M. Eldaly, Xin Si, Leanne Chan, Arindam Basu

academic team

arXiv:2606.01776 · 2026-06-01T06:58:57Z

The chip combines a dual-threshold delta-modulation frontend, in-memory spike detection and a bipolar LIF SNN decoder in 65nm CMOS. Its 3.53 uW per-channel power and 26x frontend compression are the practical points: implantable systems need to cut data movement before wireless or digital processing dominates the budget. The caveat is task scope, but the design is unusually complete as a sensing-to-decoding neuromorphic pipeline. 该芯片在65nm CMOS中集成双阈值delta调制前端、存内脉冲检测和双极LIF SNN解码器。每通道3.53微瓦功耗和26倍前端压缩是实际重点:植入式系统必须先削减数据搬移,否则无线传输或数字处理会吞掉功耗预算。局限在于任务范围,但作为从感测到解码的神经形态流水线,它的完整度较高。

SPARQLe: Sub-Precision Activation Representation for Quantized LLM Inference SPARQLe:面向量化LLM推理的子精度激活表示 LLM inferencequantizationaccelerator

Aradhana Mohan Parvathy, Soumendu Kumar Ghosh, Shamik Kundu, Arnab Raha, Souvik Kundu, Deepak A. Mathaikutty, Anand Raghunathan

academic team

arXiv:2606.00365 · 2026-05-29T21:07:27Z

SPARQLe attacks activation precision rather than only weight precision by splitting each activation tensor into dense low bits and sparse high bits. That matters for hardware because activations often force wider datapaths even when weights are 4-bit. The proposed hybrid format and accelerator path are relevant to inference chips trying to preserve accuracy without paying full 8-bit activation traffic. SPARQLe不只处理权重量化,而是把每个激活张量拆成密集低位和稀疏高位。其硬件意义在于,即便权重做到4 bit,激活往往仍迫使数据通路保持更高精度。该混合格式和加速器路径适合那些希望保留精度、又不想支付完整8 bit激活流量的推理芯片。

Inference Systems & Quantization 推理系统与量化

Extreme Low-Bit Inference in Reasoning Models 推理模型中的极低比特推理 2-bitreasoninginference control

Ekaterina Alimaskina, Darya Rudas, Denis Shveykin, Gleb Molodtsov, Pavel Vasiliev, Aleksandr Beznosikov

academic team

arXiv:2606.02011 · 2026-06-01T10:04:09Z

The paper shows that 2-bit reasoning-model inference can lose speedups because unstable traces inflate the token count. The useful insight is process-level: low-bit failures include loops, delayed commitment and unclosed reasoning, not only lower answer accuracy. FP16 planning and loop rescue are lightweight controls that make quantization an inference-policy problem as well as a kernel problem. 该论文指出,2 bit推理模型不一定带来端到端加速,因为不稳定的推理轨迹会膨胀token数量。关键洞见在过程层面:低比特失效包括循环、迟迟不提交答案和未闭合推理,而不只是答案准确率下降。FP16规划和循环救援说明,量化不仅是内核问题,也是推理策略问题。

Massive Spikes in LLMs are Bias Vectors LLM中的巨大激活尖峰是偏置向量 activationPTQlow-bit

Yung-Chin Chen, Chung Peng Lee, Ze-Wei Liou, Naveen Verma

academic team

arXiv:2606.02288 · 2026-06-01T14:09:35Z

This work reframes activation spikes as structural vector biases tied to attention-sink and value-drain mechanisms. The proposed INSERTQUANT clamps spikes and restores their function with precomputed templates, targeting a common reason low-bit activation quantization breaks. The hardware relevance is direct: taming activation dynamic range is one of the cleanest routes to narrower datapaths. 这项工作把激活尖峰重新解释为与attention sink和value drain机制相关的结构性向量偏置。其INSERTQUANT方法先钳制尖峰,再用预计算模板恢复其功能,瞄准低比特激活量化失效的常见原因。硬件相关性很直接:控制激活动态范围,是收窄数据通路最清晰的路径之一。

TwinQuant: Learnable Subspace Decomposition for 4-Bit LLM Quantization TwinQuant:面向4 bit LLM量化的可学习子空间分解 4-bitkernelLLM

Haodong Wang, Junjie Liu, Zicong Hong, Qianli Liu, Jian Lin, Song Guo, Xu Chen

academic team

arXiv:2606.01556 · 2026-06-01T02:02:12Z

TwinQuant learns quantization-friendly decomposed subspaces instead of relying on energy-minimizing decompositions that may not minimize post-quantization error. The paper also describes a fused dual-component kernel, which keeps the work relevant to deployed inference rather than only offline compression. The main caveat is added optimization complexity, but the direction matches the industry's need for 4-bit accuracy without separate slow paths. TwinQuant学习对量化友好的分解子空间,而不是依赖未必最小化量化后误差的能量最小分解。论文还描述了融合双组件内核,使其更接近部署推理而非单纯离线压缩。主要代价是优化复杂度上升,但方向符合行业对4 bit精度且不引入慢路径的需求。

Devices & Photonics 器件与光子

Silicon Photonics CW Radar with Spectrum Predistortion 带频谱预失真的硅光连续波雷达 silicon photonicsradarRF

Yicheng Du, Meng Chao, Xiuyou Han, X. Su, Xuan Li, Changjun Liu, Guanchao Wang, Mingshan Zhao

academic team

Journal of Lightwave Technology · 2026-06-01

The paper demonstrates leakage-interference cancellation for CW radar using a silicon photonics integrated chip and spectrum predistortion. Reported cancellation depths of roughly 40 dB over 300 MHz and 37 dB over 1 GHz show why photonic RF processing keeps appearing in sensing and defense-adjacent systems. The device angle is practical: delay and amplitude control move into the optical domain, reducing pressure on purely electronic cancellation chains. 该论文用硅光集成芯片和频谱预失真实现连续波雷达的泄漏干扰抵消。约40 dB/300 MHz和37 dB/1 GHz的抵消深度,解释了为什么光子RF处理持续出现在传感和防务相邻系统中。其器件意义较实际:延迟与幅度控制转移到光域,降低了纯电子抵消链路的压力。

BEOL Integration of Thin-Film Lithium Niobate on Active Silicon Photonics 薄膜铌酸锂在有源硅光平台上的后段异质集成 TFLNsilicon photonicstransceiver

Lingfeng Wu, Zhonghao Zhou, Weilong Ma, Haohua Wang, Ziliang Ruan, Shiqing Gao, Zhishan Huang, Lu Qi, Jie Liu, Jing Feng, Changjian Guo, Dapeng Liu

academic team

Laser & Photonics Reviews · 2026-05-25

This work reports trench-based die-to-wafer bonding of thin-film lithium niobate onto a completed active silicon photonics platform. The point is process sequencing: TFLN is introduced after the CMOS-compatible silicon photonics flow, avoiding some of the incompatibility that has kept high-performance modulators separate from active Si photonics. That is directly relevant to optical transceivers for AI clusters, where bandwidth density and power are now system bottlenecks. 这项工作报道了通过沟槽式die-to-wafer键合,将薄膜铌酸锂集成到已完成的有源硅光平台上。关键在工艺顺序:TFLN在CMOS兼容硅光流程完成后引入,绕开了部分让高性能调制器与有源硅光分离的工艺不兼容问题。这与AI集群光收发器直接相关,因为带宽密度和功耗已成为系统瓶颈。

Compact Memristive Spiking Neuromorphic Accelerator 紧凑型忆阻脉冲神经形态加速器 memristorSNNSKY130

Qianhou Qu, Sheng Lu, Sungyong Jung, Qilian Liang, Chenyun Pan

academic team

arXiv:2605.31141 · 2026-05-29T10:49:47Z

The authors implement a 1T1R memristive crossbar and integrate-and-fire neuron in SkyWater SKY130 for bio-inspired interception tasks. The reported 10.67 pJ/spike neuron and 96% interception success rate make it more concrete than many SNN papers that stop at algorithmic simulation. The caveat is narrow workload scope, but it is a useful silicon-adjacent data point for in-memory neuromorphic designs. 作者在SkyWater SKY130中实现1T1R忆阻交叉阵列和积分发放神经元,用于仿生拦截任务。10.67 pJ/spike的神经元和96%拦截成功率,使其比许多停留在算法仿真的SNN论文更具体。局限是工作负载较窄,但它为存内神经形态设计提供了有价值的近硅片数据点。