Devices & Process
器件与工艺
L. Popryho, A. Sadeghi, I. Partin-Vaisband
arXiv:2609.02988 · 2026-09-02T15:15:29Z
A physics-informed graph-attention surrogate works directly on tetrahedral TCAD meshes and predicts electrostatic potential plus electron and hole quasi-Fermi levels at every node. By enforcing finite-volume current continuity, the model is designed to transfer from few-fin geometries to larger arrays without changing its representation. The paper is a preprint and reports a learned surrogate rather than fabricated-device validation, so its value hinges on error and runtime behavior outside the training geometry range.
该工作提出一种直接运行在四面体TCAD网格上的物理约束图注意力代理模型,可在每个网格节点预测静电势以及电子、空穴准费米能级。模型通过有限体积电流连续性约束,力求从少鳍结构迁移到更大阵列而无需更换表示方式。论文仍是预印本,验证对象是学习型代理模型而非实际器件,因此其价值取决于超出训练几何范围后的误差与运行时间表现。
P. Dang, Y. Huang, Y. He, et al.
arXiv:2609.01948 · 2026-09-01T23:42:09Z
The NOVA architecture calibrates an on-chip-training model to a fabricated two-dimensional FeFET and steers bipolar weights toward conductance regions that are less sensitive to device asymmetry. This couples measured device behavior to a non-ideality-aware training algorithm instead of treating programming error as uniform noise. The FeFET is fabricated, but the abstract does not establish a complete accelerator tape-out, so system-level efficiency claims should be read as model-backed rather than full-chip measurements.
NOVA架构基于已制备的二维FeFET校准片上训练模型,并引导双极性权重收敛到对器件不对称性更不敏感的稳定电导区间。该方法把实测器件行为直接纳入非理想特性感知训练,而不是把编程误差简化为均匀噪声。虽然FeFET器件已经制备,但摘要并未证明完整加速器已流片,因此系统级能效结论应视为器件模型支撑的结果,而非整芯片实测。
Circuits & Architecture
电路与架构
J. Park, E. Kim, W. Kim, et al.
arXiv:2609.04058 · 2026-09-03T16:34:50Z
A unified ML-KEM-768 and ML-DSA-65 accelerator designed through 232 logged agentic-LLM experiments passed standard known-answer tests despite a block-RAM latency bug that skipped final coefficient checks. A byte-exact reference oracle plus adversarial randomized testing then completed 301,343 data-dependent signatures with zero escapes. The deployed-silicon case study matters less as proof of AI-designed hardware than as evidence that fixed-vector acceptance tests miss variable-depth cryptographic paths.
一个统一支持ML-KEM-768与ML-DSA-65的加速器由agentic LLM完成232次有记录的设计实验,尽管存在因block RAM时延导致最终系数未检查的缺陷,却仍通过了标准已知答案测试。研究随后采用逐字节精确的参考oracle与对抗性随机测试,完成301,343次数据相关签名且未再出现漏检。这个已部署硅案例的核心意义并非证明AI能够设计硬件,而是说明固定测试向量会漏掉执行深度随数据变化的密码路径。
Y. Majdane, S. J. Casartelli, E. Lopedoto
arXiv:2609.04040 · 2026-09-03T16:15:20Z
Matched ChampSim tests show that a 257-parameter online MLP loses its apparent advantage when the same confidence gate is applied to a classical stride prefetcher. The gate cuts prefetch requests by 35% and raises accuracy from 11% to 15%, yet changes DRAM reads by only 0.07%, exposing the gap between proxy metrics and endpoint behavior. These are simulation results on 20 SPEC CPU2017 workloads, not measured silicon.
在匹配的ChampSim测试中,当同一置信度门控同时用于传统步长预取器后,拥有257个参数的在线MLP不再显示出明显优势。门控减少35%的预取请求,并把准确率从11%提高到15%,但DRAM读取量仅变化0.07%,揭示了代理指标与最终系统行为之间的差距。这些结果来自20个SPEC CPU2017负载的仿真,并非芯片实测。
AI Accelerators & Compute-in-Memory
AI加速器与存算一体
O. Yousuf, M. Lueker-Boden
arXiv:2609.03149 · 2026-09-02T20:30:32Z
RACE-AIMC selects one accelerator from a heterogeneous physical pool for a given energy budget and attaches a statistically exact upper bound to the error rate when that device answers. A lightweight online check accepts confident outputs and defers uncertain cases, avoiding the energy cost of running a full ensemble on every input. The framework addresses chip-to-chip variation directly, but the abstract does not provide enough array, precision, or energy detail to compare it with measured AIMC macro results.
RACE-AIMC针对给定能耗预算,从异构物理加速器池中选择一个器件,并为该器件作答时的错误率给出统计意义上的精确上界。轻量级在线检查会接受高置信度输出并推迟不确定样本,从而避免对每个输入都运行完整集成所带来的能耗。该框架直接处理芯片间差异,但摘要未给出足够的阵列规模、精度和能耗细节,尚难与实测AIMC宏结果直接比较。
S. Chakraborty, Z. Yin, X. Chen, et al.
arXiv:2609.03125 · 2026-09-02T20:01:11Z
A photonic interposer converts analog amplitudes into timing intervals, transmits them over wavelength-division multiplexing, and reconstructs them with implicit 6-bit quantization rather than a conventional high-precision ADC/DAC chain. In an analog-vision evaluation it improves energy-delay product by 2.04 times over an 8-bit digital electrical baseline while keeping task accuracy within roughly two percentage points across additional datasets. The abstract describes system evaluation but not fabricated interposer measurements, so link fidelity and packaging overhead still need hardware validation.
该光子中介层把模拟幅度转换为时间间隔,经波分复用链路传输后重建信号,以隐式6-bit量化替代传统高精度ADC/DAC链路。在模拟视觉评估中,其能量时延积相比8-bit数字电互连基线改善2.04倍,并在额外数据集上把任务精度差距控制在约2个百分点内。摘要描述的是系统评估而非已制备中介层的实测,因此链路保真度与封装开销仍需硬件验证。
A. Madhavan, P. Carson, T. Groves, et al.
arXiv:2609.01821 · 2026-09-01T19:54:04Z
Simulations of short-, medium-, and million-token MoE workloads find that 3D-integrated photonic interconnects improve stressed high-batch prefill latency by 2.1 to 3.2 times and communication-limited cases by 2.8 to 5.8 times. The modeled fabric enables a 1,152-GPU scale-up domain that electrical systems cannot hold within one pod. The result is explicitly simulation-based, so optical power, packaging yield, and network-control overhead remain outside the demonstrated speedups.
针对短上下文、中等上下文和百万token MoE负载的仿真显示,3D集成光子互连可将高批量压力场景下的prefill时延改善2.1至3.2倍,在通信受限场景下改善2.8至5.8倍。模型中的互连可支持1,152颗GPU的scale-up域,突破电互连单pod规模上限。该结果明确基于仿真,光学功耗、封装良率与网络控制开销尚未纳入已展示的加速比。