semi·news
Headlines要闻 / Research研究 / /
Research digest · Saturday, September 5, 2026 研究摘要 · 2026年9月5日 星期六

Systems Work Moves Into the Data Path 系统研究深入数据路径

This week's work attacks data movement, verification, and scheduling across analog links, memory-side compute, photonic fabrics, and design tools. Several papers report concrete hardware evidence, while simulation-only results remain useful but conditional. 本周研究从模拟链路、存内计算、光子互连到设计工具,集中解决数据搬运、验证与调度问题。多项工作给出具体硬件证据,纯仿真结果则仍需带着条件解读。

Look-back window: 7 days · 10 paper(s) 回溯窗口: 7天 · 10篇

Devices & Process 器件与工艺

Mesh-Native Physics-Informed Surrogates for TCAD Design Exploration 面向TCAD设计探索的网格原生物理约束代理模型

L. Popryho, A. Sadeghi, I. Partin-Vaisband

arXiv:2609.02988 · 2026-09-02T15:15:29Z

A physics-informed graph-attention surrogate works directly on tetrahedral TCAD meshes and predicts electrostatic potential plus electron and hole quasi-Fermi levels at every node. By enforcing finite-volume current continuity, the model is designed to transfer from few-fin geometries to larger arrays without changing its representation. The paper is a preprint and reports a learned surrogate rather than fabricated-device validation, so its value hinges on error and runtime behavior outside the training geometry range. 该工作提出一种直接运行在四面体TCAD网格上的物理约束图注意力代理模型,可在每个网格节点预测静电势以及电子、空穴准费米能级。模型通过有限体积电流连续性约束,力求从少鳍结构迁移到更大阵列而无需更换表示方式。论文仍是预印本,验证对象是学习型代理模型而非实际器件,因此其价值取决于超出训练几何范围后的误差与运行时间表现。

FeFET-Aware On-Chip Training With Stable Conductance Regions 利用稳定电导区间的FeFET感知片上训练

P. Dang, Y. Huang, Y. He, et al.

arXiv:2609.01948 · 2026-09-01T23:42:09Z

The NOVA architecture calibrates an on-chip-training model to a fabricated two-dimensional FeFET and steers bipolar weights toward conductance regions that are less sensitive to device asymmetry. This couples measured device behavior to a non-ideality-aware training algorithm instead of treating programming error as uniform noise. The FeFET is fabricated, but the abstract does not establish a complete accelerator tape-out, so system-level efficiency claims should be read as model-backed rather than full-chip measurements. NOVA架构基于已制备的二维FeFET校准片上训练模型,并引导双极性权重收敛到对器件不对称性更不敏感的稳定电导区间。该方法把实测器件行为直接纳入非理想特性感知训练,而不是把编程误差简化为均匀噪声。虽然FeFET器件已经制备,但摘要并未证明完整加速器已流片,因此系统级能效结论应视为器件模型支撑的结果,而非整芯片实测。

Circuits & Architecture 电路与架构

A Deployed-Silicon Post-Quantum Accelerator Exposes a Verification Blind Spot 已部署后量子加速器暴露验证盲区

J. Park, E. Kim, W. Kim, et al.

arXiv:2609.04058 · 2026-09-03T16:34:50Z

A unified ML-KEM-768 and ML-DSA-65 accelerator designed through 232 logged agentic-LLM experiments passed standard known-answer tests despite a block-RAM latency bug that skipped final coefficient checks. A byte-exact reference oracle plus adversarial randomized testing then completed 301,343 data-dependent signatures with zero escapes. The deployed-silicon case study matters less as proof of AI-designed hardware than as evidence that fixed-vector acceptance tests miss variable-depth cryptographic paths. 一个统一支持ML-KEM-768与ML-DSA-65的加速器由agentic LLM完成232次有记录的设计实验,尽管存在因block RAM时延导致最终系数未检查的缺陷,却仍通过了标准已知答案测试。研究随后采用逐字节精确的参考oracle与对抗性随机测试,完成301,343次数据相关签名且未再出现漏检。这个已部署硅案例的核心意义并非证明AI能够设计硬件,而是说明固定测试向量会漏掉执行深度随数据变化的密码路径。

Confidence Gating Matters More Than the Learned Prefetcher 置信度门控比学习型预取器更重要

Y. Majdane, S. J. Casartelli, E. Lopedoto

arXiv:2609.04040 · 2026-09-03T16:15:20Z

Matched ChampSim tests show that a 257-parameter online MLP loses its apparent advantage when the same confidence gate is applied to a classical stride prefetcher. The gate cuts prefetch requests by 35% and raises accuracy from 11% to 15%, yet changes DRAM reads by only 0.07%, exposing the gap between proxy metrics and endpoint behavior. These are simulation results on 20 SPEC CPU2017 workloads, not measured silicon. 在匹配的ChampSim测试中,当同一置信度门控同时用于传统步长预取器后,拥有257个参数的在线MLP不再显示出明显优势。门控减少35%的预取请求,并把准确率从11%提高到15%,但DRAM读取量仅变化0.07%,揭示了代理指标与最终系统行为之间的差距。这些结果来自20个SPEC CPU2017负载的仿真,并非芯片实测。

AI Accelerators & Compute-in-Memory AI加速器与存算一体

Risk-Aware Selective Inference for Analog In-Memory Accelerators 面向模拟存内计算加速器的风险感知选择性推理

O. Yousuf, M. Lueker-Boden

arXiv:2609.03149 · 2026-09-02T20:30:32Z

RACE-AIMC selects one accelerator from a heterogeneous physical pool for a given energy budget and attaches a statistically exact upper bound to the error rate when that device answers. A lightweight online check accepts confident outputs and defers uncertain cases, avoiding the energy cost of running a full ensemble on every input. The framework addresses chip-to-chip variation directly, but the abstract does not provide enough array, precision, or energy detail to compare it with measured AIMC macro results. RACE-AIMC针对给定能耗预算,从异构物理加速器池中选择一个器件,并为该器件作答时的错误率给出统计意义上的精确上界。轻量级在线检查会接受高置信度输出并推迟不确定样本,从而避免对每个输入都运行完整集成所带来的能耗。该框架直接处理芯片间差异,但摘要未给出足够的阵列规模、精度和能耗细节,尚难与实测AIMC宏结果直接比较。

A Time-Encoded Photonic Interposer Carries Analog Chiplet Signals 时间编码光子中介层传输模拟chiplet信号

S. Chakraborty, Z. Yin, X. Chen, et al.

arXiv:2609.03125 · 2026-09-02T20:01:11Z

A photonic interposer converts analog amplitudes into timing intervals, transmits them over wavelength-division multiplexing, and reconstructs them with implicit 6-bit quantization rather than a conventional high-precision ADC/DAC chain. In an analog-vision evaluation it improves energy-delay product by 2.04 times over an 8-bit digital electrical baseline while keeping task accuracy within roughly two percentage points across additional datasets. The abstract describes system evaluation but not fabricated interposer measurements, so link fidelity and packaging overhead still need hardware validation. 该光子中介层把模拟幅度转换为时间间隔,经波分复用链路传输后重建信号,以隐式6-bit量化替代传统高精度ADC/DAC链路。在模拟视觉评估中,其能量时延积相比8-bit数字电互连基线改善2.04倍,并在额外数据集上把任务精度差距控制在约2个百分点内。摘要描述的是系统评估而非已制备中介层的实测,因此链路保真度与封装开销仍需硬件验证。

High-Radix Photonics Scales MoE Inference Prefill 高基数光子互连扩展MoE推理Prefill

A. Madhavan, P. Carson, T. Groves, et al.

arXiv:2609.01821 · 2026-09-01T19:54:04Z

Simulations of short-, medium-, and million-token MoE workloads find that 3D-integrated photonic interconnects improve stressed high-batch prefill latency by 2.1 to 3.2 times and communication-limited cases by 2.8 to 5.8 times. The modeled fabric enables a 1,152-GPU scale-up domain that electrical systems cannot hold within one pod. The result is explicitly simulation-based, so optical power, packaging yield, and network-control overhead remain outside the demonstrated speedups. 针对短上下文、中等上下文和百万token MoE负载的仿真显示,3D集成光子互连可将高批量压力场景下的prefill时延改善2.1至3.2倍,在通信受限场景下改善2.8至5.8倍。模型中的互连可支持1,152颗GPU的scale-up域,突破电互连单pod规模上限。该结果明确基于仿真,光学功耗、封装良率与网络控制开销尚未纳入已展示的加速比。

Hardware-Relevant AI Research 硬件相关AI研究

FlowTT Reuses Irregular Tensor-Train Embedding Computation FlowTT复用不规则Tensor-Train嵌入计算

J. Seok, C. E. Rhee

arXiv:2609.03459 · 2026-09-03T07:17:33Z

FlowTT groups tensor-train embedding lookups by shared prefixes, fuses core contractions to retain intermediates on chip, and uses persistent threads with work stealing to control skew. At batch size 32,768, it cuts latency by as much as 42.2% for inference and 49.2% for training relative to EcoRec. The benchmarks are Meta synthetic recommendation workloads, so gains on production index distributions and smaller batches remain to be established. FlowTT按照共享前缀对Tensor-Train嵌入查找进行分组,融合核心收缩以把中间结果保留在片上,并通过持久线程与任务窃取处理负载偏斜。在batch size为32,768时,它相对EcoRec最多降低42.2%的推理时延和49.2%的训练时延。测试采用Meta合成推荐负载,因此在真实生产索引分布和更小batch下的收益仍有待验证。

Quantum Computing 量子计算

A Neutral-Atom MWIS Kernel Reaches an Industrial Scheduling Loop 中性原子MWIS内核进入工业调度闭环

J. Chen, M. Lin, J. Wen, et al.

arXiv:2609.01248 · 2026-09-01T13:46:26Z

A backend-agnostic interface maps the discrete layer of stochastic unit commitment to maximum-weight independent set instances, then leaves continuous dispatch and feasibility recovery to classical computation. During a 15-day campaign on QuEra Aquila, refined hardware solutions on 50-node instances matched or exceeded dispatch margins obtained from exact MWIS each day. At 144 nodes, full-array atom survival becomes the limiting factor, showing that hardware availability rather than graph encoding is the immediate scaling bottleneck. 该后端无关接口把随机机组组合的离散决策层映射为最大权独立集问题,同时由经典计算负责连续调度与可行性恢复。在QuEra Aquila上持续15天的测试中,50节点实例的量子硬件解经经典修正后,每天都达到或超过精确MWIS所得的调度裕量。当规模扩大到144节点时,全阵列原子存活率成为限制因素,说明眼前的扩展瓶颈是硬件可用性而非图编码。

EDA & Design Automation EDA与设计自动化

Batched Proxy Execution Speeds Timing-Aware Logic Rewriting 批量代理执行加速时序感知逻辑重写

P. Su

Fudan University

arXiv:2609.02470 · 2026-09-02T11:39:48Z

BBYT compiles all candidates for one logic-rewrite decision into a single zero-delay proxy image, accepting a clear proxy winner or falling back to the original timed flow. Across 12 counterbalanced sequences it reduces candidate-selection time by 18.05%, removes 32.52% of timed evaluations, and matches exhaustive timed selection on all 250 tested decisions. The evidence covers only a five-workload corpus and two reported design-level reductions, so broader netlists are needed to establish generality. BBYT把一次逻辑重写决策的全部候选方案编译进单个零延迟代理镜像;若代理结果明显领先则直接采用,否则回退到原有时序流程。在12组平衡序列中,它将候选选择时间降低18.05%,移除32.52%的时序评估,并在全部250次测试决策中与穷举时序选择保持一致。现有证据仅覆盖5个负载和两个设计级降幅,仍需更广泛网表验证其通用性。

Newsletter 邮件订阅

Daily semiconductor briefing. 每日半导体简报。