semi·news
Headlines要闻 / Research研究 / /
Research digest · Tuesday, September 1, 2026 研究摘要 · 2026年9月1日 星期二

Memory-Centric Compute Gets More Programmable 以存储为中心的计算更加可编程

New measured CIM macros pursue programmable arithmetic and higher-precision data paths, while device and system papers seek denser memory and lower data movement. The research queue also follows the design and deployment tooling needed to make those hardware gains useful. 新的实测CIM宏单元追求可编程算术与更高精度的数据路径,器件和系统论文则探索更高密度的存储与更少的数据搬运。研究队列同时关注让这些硬件收益真正落地所需的设计与部署工具。

Look-back window: 7 days · 10 paper(s) 回溯窗口: 7天 · 10篇

AI Accelerators & Compute-in-Memory AI加速器与存算一体

PROTEUS: programmable RRAM/SRAM compute-in-memory for edge AI PROTEUS:面向边缘AI的可编程RRAM/SRAM存算一体加速器

L. Zheng, A. M. Bavani, M. Chen, et al.

IEEE Journal of Solid-State Circuits · 2026-09-01

PROTEUS is a 40 nm programmable digital CIM accelerator combining 4 Mb RRAM, 2.6 Mb tensor SRAM, and a 32-bit hierarchical ISA. It supports INT8, INT16, FP8, and FP16 execution and reports 702 GOPS, 6.4 TOPS/W, and 0.039 TOPS/mm². By storing micro-programs in RRAM, it can switch among preloaded kernels without off-chip instruction traffic or RRAM rewrites; the reported results cover CNNs, Transformers, GNNs, and state-space models. PROTEUS是一款40 nm可编程数字CIM加速器,集成4 Mb RRAM、2.6 Mb tensor SRAM和32位分层ISA。它支持INT8、INT16、FP8和FP16运算,报告的性能为702 GOPS、6.4 TOPS/W和0.039 TOPS/mm²。通过把微程序存入RRAM,该设计可在预载内核之间切换而无需片外指令流或重写RRAM;实验覆盖CNN、Transformer、GNN和state-space model。

STAR-SRAM: BF16 SRAM digital CIM macro in 28 nm STAR-SRAM:28 nm BF16 SRAM数字存算一体宏单元

C.-T. Lin, J. Oh, M. Seok

IEEE Journal of Solid-State Circuits · 2026-09-01

STAR-SRAM implements a 16-bit BF16 SRAM-based digital CIM macro in 28 nm CMOS. Its designers combine model-based parameter selection with approximate multipliers and sparsity-aware wordline, clock, and input gating. The paper reports up to 43.06 TFLOPS/W, 1.89 TFLOPS/mm², and 400 Kb/mm², directly targeting the precision and density gap between low-bit CIM and modern neural workloads. STAR-SRAM在28 nm CMOS中实现了基于SRAM的16位BF16数字CIM宏单元。设计团队结合基于模型的参数选择、近似乘法器,以及稀疏性感知的wordline、时钟和输入门控。论文报告最高43.06 TFLOPS/W、1.89 TFLOPS/mm²和400 Kb/mm²,直接瞄准低比特CIM与现代神经网络工作负载之间的精度和密度缺口。

NOVA: near-memory processing for attention–SSM–MoE LLMs NOVA:面向Attention–SSM–MoE大模型的近存计算

I. Jung, J. Min, J.-Y. Kim

arXiv:2608.22613 · 2026-08-23

NOVA proposes a technology–architecture co-design for hybrid LLMs that combine grouped-query attention, state-space models, and MoE layers. It pairs a proposed 4F² vertical-channel DRAM cell and peri-over-cell structure, claimed to provide roughly 2× density at iso-area, with a near-memory architecture that adapts to disparate arithmetic intensities. The work is a preprint, so its density and system claims should be read as a proposed design rather than production silicon results. NOVA提出了针对混合LLM的技术—架构协同设计,此类模型结合grouped-query attention、state-space model和MoE层。它将所提出的4F² vertical-channel DRAM单元及peri-over-cell结构——声称在等面积下约有2倍密度——与能适配不同算术强度的近存计算架构结合。该工作是预印本,因此其密度和系统主张应视为设计方案,而非量产硅片结果。

Devices & Process 器件与工艺

Electroforming-free selector-only memory for logic-in-memory 用于逻辑存内计算的免电形成selector-only memory

J.-K. Kim, D.-S. Woo, M.-J. Han, et al.

Advancement of science · 2026-08-25

This paper reports an electroforming-free, self-rectifying selector-only memory based on diffusive Cu-ion dynamics. The device combines diode-like unidirectional current with reported endurance near 10⁸ cycles and sub-20 ns self-rupturing conductive filaments. Eliminating high-voltage forming and reducing drift could address practical array concerns for 3D cross-point memory and logic-in-memory, although array-level integration remains the key test. 该论文报告了一种基于扩散Cu离子动力学的免电形成、自整流selector-only memory。该器件兼具类似二极管的单向电流,报告的耐久度接近10⁸次,并具有小于20 ns的导电细丝自断裂特性。免除高压forming并降低漂移,有望缓解3D cross-point memory和logic-in-memory的阵列实现难题,但阵列级集成仍是关键验证。

Wafer-scale 3D ternary neuromorphic logic using Te/IGZO devices 基于Te/IGZO器件的晶圆级3D三值神经形态逻辑

C.-H. Kim, M. S. Kim, J. B. Rhim, et al.

Advancement of science · 2026-08-25

The authors demonstrate wafer-scale monolithic 3D multi-valued logic using vertically stacked Te/IGZO heterojunction FETs and crystallinity-enhanced Te FETs. Interface and channel engineering stabilize an intrinsic ternary state using low-temperature CMOS-compatible processing. The result points to a route for higher logic density, but the architectural benefit depends on system-level support for ternary logic rather than device behavior alone. 作者利用垂直堆叠的Te/IGZO异质结FET和结晶性增强Te FET,展示了晶圆级单片3D多值逻辑。界面和沟道工程在低温CMOS兼容工艺下稳定了内禀三值状态。该结果指向更高逻辑密度的路径,但其架构收益还取决于系统级对三值逻辑的支持,而不仅是器件特性。

Circuits & Systems 电路与系统

Dual-mode BEV Transformer accelerator for image/point-cloud fusion 用于图像/点云融合的双模式BEV Transformer加速器

J. Qian, Y. Lu, Z. Qian, et al.

IEEE Journal of Solid-State Circuits · 2026-09-01

A 28 nm chip accelerates Transformer bird's-eye-view perception by combining saliency-driven image/point-cloud fusion, adaptive precision, and a dataflow designed to reuse on-chip data. It supports a high-performance driving mode and a low-power sentry mode, with 26.6 mW active power and 1.68 mW long-term average power reported for sparse-activity scenarios. The dual-mode design addresses a practical autonomous-system problem: retaining perception capability when full compute throughput is unnecessary. 这款28 nm芯片通过显著性驱动的图像/点云融合、自适应精度和面向片上数据复用的数据流,加速Transformer鸟瞰图感知。它支持高性能驾驶模式和低功耗哨兵模式;在稀疏活动场景中,报告的active power为26.6 mW、长期平均功耗为1.68 mW。该双模式设计解决了自动驾驶系统中的实际问题:在不需要完整计算吞吐时仍保留感知能力。

AI Systems & Inference AI系统与推理

Hydra: phase-aware characterization of edge LLM inference Hydra:面向边缘LLM推理的阶段感知性能刻画

A. Taherin, S. Taghipour Anvari, C. Amante, et al.

arXiv:2608.25053 · 2026-08-25T18:43:43Z

Hydra provides a common-schema measurement framework for prefill and decode phases across three edge-SoC generations, 13 instruction-tuned LLMs, five execution formats, and two inference backends. Its released dataset contains roughly 107,000 per-prompt records combined with hardware telemetry. The analysis argues that latency averages conceal important effects from backend choice, quantization, memory traffic, and phase-specific utilization. Hydra提供统一schema的测量框架,用于比较三代edge SoC、13个指令微调LLM、5种执行格式和两种推理后端上的prefill与decode阶段。其公开数据集包含约10.7万条逐prompt记录,并结合硬件遥测数据。分析指出,平均延迟掩盖了后端选择、量化、内存流量和阶段特定利用率带来的重要影响。

Event-triggered implicit perturbation for fine-tuning spiking Transformers 用于微调脉冲Transformer的事件触发隐式扰动方法

T. Lei, P. Katti, R. Dutt, et al.

arXiv:2608.21223 · 2026-08-21

This preprint proposes implicit perturbation for zeroth-order fine-tuning of spiking Transformers on in-memory-computing hardware. It combines perturbation terms with IMC weighted sums to avoid perturbation-induced read-modify-write operations, and exploits spike sparsity to reduce random-number-generator requirements. The idea addresses a hardware-specific obstacle in forward-only optimization, though its claimed benefits require validation on a complete accelerator implementation. 该预印本提出用于IMC硬件上脉冲Transformer zeroth-order微调的隐式扰动方法。它将扰动项与IMC加权和结合,以避免由扰动引起的read-modify-write操作,并利用脉冲稀疏性降低对随机数发生器的需求。该思路针对仅前向优化中的硬件特定障碍,但其所宣称的收益仍需在完整加速器实现上验证。

EDA & Design Automation EDA与设计自动化

DeepSeq3: hierarchical graph learning for sequential-circuit analysis DeepSeq3:用于时序电路分析的分层图学习

J. Zhou, Z. Shi, J. Zhu, et al.

arXiv:2608.28188 · 2026-08-28T10:56:04Z

DeepSeq3 represents a sequential circuit at two levels: combinational subgraphs divided by flip-flops and a super-node graph for register-transfer structure. A dual GNN is pre-trained to predict flip-flop-state reachability, capturing both local logic and temporal behavior. On the reported benchmarks, the framework reduces bounded-model-checking solve time by 18% while preserving correctness, though results remain limited to the evaluated designs. DeepSeq3以两个层级表示时序电路:由flip-flop划分的组合子图,以及表达寄存器传输结构的super-node图。双GNN通过预测flip-flop状态可达性进行预训练,同时捕捉局部逻辑和时序行为。在报告的基准上,该框架在保持正确性的前提下将bounded model checking求解时间降低18%,但结果仍限于所评估的设计。

HOLMES: in-context failure localization for SRAM yield estimation HOLMES:用于SRAM良率估计的in-context失效中心定位

W. W. Xing, X. Zhou, K. Huang, et al.

arXiv:2608.26758 · 2026-08-27T07:53:18Z

HOLMES recasts failure-center localization for high-sigma yield estimation as few-shot binary classification using a tabular foundation model. It adds an SVD-based anisotropic sampling proposal and adaptive mixing to stabilize importance weights in high-dimensional SRAM problems. Across the authors' 6T SRAM cases from 108 to 1,152 dimensions, HOLMES reports relative error within 5.9%, compared with up to 25.8% for the strongest baseline. HOLMES将high-sigma良率估计中的失效中心定位重构为few-shot二分类,并采用tabular foundation model进行in-context推理。它增加了基于SVD的各向异性采样proposal和自适应混合,以稳定高维SRAM问题中的importance weights。在作者的6T SRAM案例中,维度从108到1,152,HOLMES报告的相对误差不超过5.9%,而最强baseline最高达到25.8%。

Newsletter 邮件订阅

Daily semiconductor briefing. 每日半导体简报。