semi·news
Headlines要闻 / Research研究 / /
Research digest · Monday, August 10, 2026 研究摘要 · 2026年8月10日 星期一

Inference efficiency shifts across the stack 推理效率向全栈迁移

This week’s papers connect low-bit formats, sparsity, scheduling, and memory-centric circuits to the practical cost of serving models. The strongest results range from silicon prototypes to system-level measurements, with several claims still based on simulation or preprint evaluation. 本周论文将低比特格式、稀疏性、调度和以存储为中心的电路,连接到模型服务的实际成本问题。最有力的结果覆盖硅原型和系统级测量,但也有若干结论仍基于仿真或预印本评估。

Look-back window: 7 days · 10 paper(s) 回溯窗口: 7天 · 10篇

Devices & Process 器件与工艺

Ferroelectric Hafnium Oxide for In-Memory Computing 面向存算一体的铁电氧化铪

Chen He, Wei Li, Jianjun Li, et al.

Micromachines · 2026-08-04

This review connects hafnia-based ferroelectric materials, memory devices, circuit primitives, and system-level in-memory computing. It explains how polarization engineering and defect control constrain precision, endurance, scalability, and the circuit functions that can be implemented. The article is a synthesis rather than a new device demonstration, but it is useful for separating CMOS-compatible promise from integration bottlenecks. 这篇综述将基于氧化铪的铁电材料、存储器件、电路原语和系统级存算一体联系起来。文章说明极化工程和缺陷控制如何约束精度、耐久性、可扩展性以及可实现的电路功能。它并非新的器件演示,但有助于区分CMOS兼容的潜力与集成瓶颈。

Circuits & Architectures 电路与体系结构

A 350-pW 65-nm CMOS arrhythmia detector with Bayesian uncertainty 采用贝叶斯不确定性量化的350 pW、65 nm CMOS心律失常检测器

Zephan M. Enciso, Jianbo Liu, Boyang Cheng, et al.

IEEE Journal of Solid-State Circuits · 2026-08-01

The authors report a 0.816-mm² implantable ventricular-arrhythmia detection engine in 65-nm CMOS that consumes 350 pW. It combines a digital one-dimensional convolution front end with a mixed-signal compute-in-memory Bayesian classifier and a 360-fJ/sample random-number generator. Integrating uncertainty estimates is important for selective prediction, although clinical performance and robustness remain separate validation questions. 作者报告了一款采用65 nm CMOS、面积0.816 mm²、功耗350 pW的植入式室性心律失常检测引擎。该芯片将数字一维卷积前端与混合信号存算一体贝叶斯分类器及360 fJ/样本随机数发生器结合。集成不确定性估计对选择性预测很重要,但临床性能与鲁棒性仍需独立验证。

A 94.8-nW battery-free intelligent silicon platform 功耗94.8 nW的无电池智能硅平台

Haochen Zhang, Wei-Han Yu, Zhongyu Zhao, et al.

IEEE Journal of Solid-State Circuits · 2026-08-01

This 65-nm CMOS platform targets distributed multimodal sensing with event-driven computing and digital compute-in-memory. It reports 94.8 nW in inference mode, 340.4 nW during training, and 90.6% accuracy on a condition-monitoring task. The result shows how sparsity-aware scheduling and low-bit learning can stretch energy budgets, though the reported workload is much narrower than general-purpose edge AI. 该65 nm CMOS平台面向分布式多模态感测,采用事件驱动计算和数字存算一体。论文报告其推理模式功耗为94.8 nW、训练模式为340.4 nW,并在状态监测任务上获得90.6%的准确率。结果表明稀疏性感知调度和低比特学习能够延展能耗预算,但所报告负载远窄于通用边缘AI。

VCMA-MRAM in-memory arithmetic for subtractors and dividers 用于减法器和除法器的VCMA-MRAM存内算术架构

Qilong Tang, Wei Duan, Zijun Zhu, et al.

IEEE Transactions on Magnetics · 2026-08-01

The paper proposes a voltage-controlled magnetic-anisotropy MRAM architecture that performs logic directly inside the array to build subtractors and dividers. In 45-nm CMOS simulation, the design reports 512 GOPS peak throughput and 99.853 TOPS/W; its 8-bit subtractor is claimed to reduce area by 43.85% and latency by 53.66%. These are simulation-based figures, so device variation, write reliability, and peripheral cost remain central implementation risks. 论文提出一种基于电压控制磁各向异性MRAM的架构,在阵列内直接执行逻辑以构建减法器和除法器。在45 nm CMOS仿真中,该设计报告峰值吞吐量512 GOPS、能效99.853 TOPS/W;其8位减法器据称面积减少43.85%、延迟降低53.66%。这些数值基于仿真,因此器件波动、写入可靠性和外围电路成本仍是关键实现风险。

AI Accelerators & Compute-in-Memory AI加速器与存算一体

AdaMX: Heterogeneity-aware microscaling for low-bit LLM inference AdaMX:面向低比特LLM推理的异构感知微缩放

Junyi Luo, Xin Jiang, Tai-Hao Wen, et al.

arXiv:2608.03867 · 2026-08-04

AdaMX selects precision-recovery schemes per block and different representations for weights and activations, rather than applying one MXFP4 treatment everywhere. The team implements the decoder, compute unit, and quantization logic in a 22-nm FD-SOI accelerator prototype, reporting roughly 1% system-energy overhead versus an otherwise identical MXFP4 accelerator. The work is useful because it puts hardware cost beside quantization flexibility, but the reported comparison depends on the chosen baseline and models. AdaMX按块选择精度恢复机制,并为权重和激活采用不同表示,而不是在所有位置使用同一种MXFP4处理方式。团队在22 nm FD-SOI加速器原型中实现了解码器、计算单元和量化逻辑,报告相对同等MXFP4加速器约1%的系统能耗开销。该工作将硬件成本与量化灵活性并置,但其比较结果依赖于所选基线和模型。

Celty: GPU kernel and SIMT co-design for dual-sparse LLM inference Celty:面向双稀疏LLM推理的GPU内核与SIMT协同设计

Ruokai Yin, Priyadarshini Panda

arXiv:2608.01536 · 2026-08-02

Celty targets the sparse-matrix–sparse-vector workload created when weight pruning meets runtime activation sparsity in single-user LLM decoding. It co-designs an RLC-CSC sparse format, GPU kernel, and a SIMT core with a pipelined decoder to avoid costly index reconstruction and memory accesses. The idea addresses a real mismatch between sparsity algorithms and commodity kernels, while the microarchitecture portion remains a proposed design rather than deployed GPU silicon. Celty针对单用户LLM解码中权重剪枝与运行时激活稀疏叠加形成的稀疏矩阵—稀疏向量负载。它协同设计RLC-CSC稀疏格式、GPU内核和带流水线解码器的SIMT核心,以避免高代价的索引重建和内存访问。该思路解决了稀疏算法与通用内核间的真实错配,但其中的微架构部分仍是提案,而非已部署的GPU硅实现。

DSLA: dual-sparsity LLM accelerator with HiMix-BFP DSLA:采用HiMix-BFP的双稀疏LLM加速器

Zikang Zhou, Siyao Dai, Yaqi Chen, et al.

IEEE Transactions on Circuits and Systems I: Regular Papers · 2026-08-01

DSLA combines a HiMix block-floating-point representation with exploitation of sparsity in both data and bit slices for edge LLM inference. The paper argues that conventional BFP and bidirectional BFP lose too much accuracy at ultra-low precision because of outliers. Its contribution is an accelerator-oriented way to recover accuracy while turning exponent-alignment behavior into an additional energy-saving opportunity; deployment results will determine the practical gain. DSLA将HiMix块浮点表示与数据和位片两个层面的稀疏性利用结合,用于边缘LLM推理。论文认为,传统BFP和双向BFP在超低精度下会因离群值损失过多准确率。其贡献在于以加速器为导向恢复精度,同时把指数对齐行为转化为额外节能机会;实际部署结果将决定其可获得的收益。

AI Systems & Inference AI系统与推理

OrionInfer: adaptive parallelism and live migration for LLM serving OrionInfer:用于LLM服务的自适应并行与在线迁移

Jingqi Feng, Guang Yang, Yukai Huang, et al.

Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 · 2026-08-08

OrionInfer switches between data parallelism and tensor parallelism at runtime, while using live migration to rebalance memory pressure. The authors report up to 25% lower average time-to-first-token than a data-parallel-priority configuration at low load, and 50%–90% lower P99 tail latency than a tensor-parallel-priority configuration in many high-traffic cases. The work makes serving policy dynamic, though benefits will depend on migration overhead and the workload mix at an operator. OrionInfer可在运行时切换数据并行与张量并行,并通过在线迁移重新平衡内存压力。作者报告,在低负载下,相比优先数据并行的配置,平均首Token时间最多降低25%;在许多高流量场景中,相比优先张量并行的配置,P99尾延迟降低50%至90%。该工作使服务策略动态化,但收益仍取决于迁移开销和运营方的实际负载组合。

EdgeXpert: MoE and speculative decoding for memory-efficient edge inference EdgeXpert:面向高效边缘推理的MoE与推测解码

Sangwoo Ha, Hyunwoo Seo, Y. Jo, et al.

arXiv:2608.05303 · 2026-08-05

EdgeXpert is a hardware–software co-design that combines mixture-of-experts routing with speculative decoding for on-device LLMs. It uses prompt-wise expert reuse in prefill and depth-aware expert coalescing in decoding to reduce external-memory traffic. The paper targets an important edge bottleneck, but its claimed advantage rests on how well expert reuse preserves model quality across prompts and candidate tokens. EdgeXpert是一项软硬件协同设计,将Mixture-of-Experts路由与推测解码结合,用于端侧LLM。它在prefill阶段使用按提示词复用专家,在解码阶段使用深度感知的专家合并,以降低外部存储访问。论文瞄准了重要的边缘瓶颈,但其优势取决于专家复用能否在不同提示词和候选Token间保持模型质量。

LLM Serving in the Wild: framework adoption in open-source systems 真实世界中的LLM服务:开源系统中的框架采用情况

Forough Majidi, Mohammad Mehdi Morovati, F. Khomh, et al.

arXiv:2608.03036 · 2026-08-04

This empirical study examines how open-source projects use vLLM, SGLang, TensorRT-LLM, LMDeploy, and FlashInfer. It finds vLLM is the most visible framework, while parallel computation, memory management, and network pruning are the most common serving-method categories; multi-framework use is limited. That evidence is a useful counterweight to benchmark-driven claims, although public repositories may not represent proprietary production deployments. 这项实证研究考察了开源项目如何使用vLLM、SGLang、TensorRT-LLM、LMDeploy和FlashInfer。研究发现vLLM的可见度和采用度最高,并行计算、内存管理和网络剪枝是最常见的服务方法类别;多框架并用较少。这一证据可为基准测试驱动的主张提供现实参照,但公开仓库未必代表专有生产部署。

Newsletter 邮件订阅

Daily semiconductor briefing. 每日半导体简报。