semi·news
Headlines要闻 / Research研究 / /
Research digest · Saturday, September 12, 2026 研究摘要 · 2026年9月12日 星期六

Chiplets Push Optimization Across Physical Boundaries Chiplet推动跨物理边界优化

This week's papers move co-design below the model graph and into chiplet pools, power grids, thermal maps, bonding interfaces, and nonvolatile devices. The common goal is to recover efficiency without ignoring manufacturability, reliability, or security. 本周论文把协同设计从模型图进一步推进到chiplet资源池、供电网络、热图、键合界面与非易失器件。共同目标是在不忽视可制造性、可靠性与安全性的前提下重新获取效率。

Look-back window: 7 days · 9 paper(s) 回溯窗口: 7天 · 9篇

Devices & Process 器件与工艺

Self-Activated, Plasma-Free Direct Bonding With ALD Al2O3 采用ALD Al2O3的自激活无等离子体直接键合

ACS Applied Materials and Interfaces · 2026-09-03

Atomic-layer-deposited Al2O3 on 300-mm wafers forms strong, void-free direct bonds without plasma activation or CMP pretreatment. Multiple interface probes attribute the effect to dense polar Al-OH networks stabilized by the amorphous ALD film. Eliminating activation steps could simplify 3D CMOS and heterogeneous-integration flows, although production reliability and process-window data remain to be established. 在300mm晶圆上沉积的ALD Al2O3无需等离子体激活或CMP预处理,即可形成高强度、无空洞的直接键合。多种界面表征把这一效应归因于由非晶ALD薄膜稳定的高密度极性Al-OH网络。省去激活步骤有望简化3D CMOS与异构集成流程,但量产可靠性和制程窗口数据仍需验证。

A Hardware-Bound Binary Neural Network Using Concealable MTJs 采用可隐藏MTJ的硬件绑定二值神经网络

Sung-Ho Park, K. Lee, Ryun-Han Koo, et al.

IEEE Electron Device Letters · 2026-09-01

A fabricated 28-nm FDSOI embedded-MRAM macro uses stochastic magnetic-tunnel-junction breakdown cells to bind a binary neural network's weights to one physical chip. The network retains high accuracy in reveal mode, while concealment or moving the protected layer to another chip collapses accuracy to near-random levels. The measured result offers model protection without separate cryptographic circuitry, though it is demonstrated on a binary network rather than a large production model. 研究人员在28nm FDSOI嵌入式MRAM宏单元中利用随机磁隧道结击穿单元,把二值神经网络权重绑定到一颗物理芯片。网络在显现模式下保持较高精度,而隐藏受保护层或把它迁移到另一颗芯片会使精度降至接近随机水平。这一实测结果无需额外密码电路即可提供模型保护,但目前验证对象仍是二值网络,而非大型量产模型。

SOT-MTJ Hardware for Noise-Tolerant Probabilistic Binary Networks 面向抗噪概率二值网络的SOT-MTJ硬件

Applied Physics Reviews · 2026-09-01

In-plane SOT-MTJ cells demonstrate 400-ps electrical switching, 10^11 endurance, and voltage-controlled probabilistic states with 3.8% variation. An on-chip probabilistic binary neural network reaches 88.2% CIFAR-10 accuracy and, under 25% read/write noise, achieves seven times the accuracy of the cited 32-bit CNN while cutting parameter size and inference energy by one to two orders of magnitude. The device-to-system result is compelling, but the evidence is still tied to a modest image-classification workload. 面内SOT-MTJ单元实现400ps全电切换、10^11次耐久性,以及变异率为3.8%的电压可控概率态。片上概率二值神经网络在CIFAR-10上达到88.2%精度;在25%读写噪声下,其精度达到所比较32位CNN的7倍,同时把参数规模与推理能耗降低1至2个数量级。该工作打通了器件到系统的验证,但证据仍局限于规模较小的图像分类负载。

Circuits, Packaging & Reliability 电路、封装与可靠性

Thermal Management and Reliability of Advanced HBM Packages 先进HBM封装的热管理与可靠性

Micromachines · 2026-09-08

This review connects vertical heat buildup in denser HBM stacks with warpage, delamination, copper protrusion, voids, electromigration, and joint degradation. It surveys thermal-interface materials, underfill, molding compounds, heat spreaders, and high-conductivity composites as a coupled design space rather than isolated fixes. The paper is a useful failure-mode map for HBM packaging, but it synthesizes prior work instead of reporting new silicon. 这篇综述把更高层数HBM中的垂直积热,与翘曲、分层、铜凸起、空洞、电迁移和焊点退化联系起来。论文系统梳理热界面材料、底部填充、塑封料、均热结构和高导热复合材料,强调它们构成相互耦合的设计空间,而非彼此孤立的补救手段。它为HBM封装提供了有用的失效模式地图,但属于既有工作的综合,并未报告新硅片结果。

Deep Learning Accelerates Thermal-Aware Signal-Integrity Analysis for 2.5D Chiplets 深度学习加速2.5D chiplet热感知信号完整性分析

M. Wei, Anton Sattler, H. Amrouch

IEEE Transactions on Circuits and Systems Part 1: Regular Papers · 2026-09-01

A physics-guided neural model predicts thermal maps for 2.5D chiplet systems and feeds them into signal-integrity analysis, avoiding repeated full multiphysics simulation. Its cross feature captures heat-flux exchange between chiplets, while a new equivalent model supports pillar-style microbumps when generating training data. The approach targets faster design-space exploration, but its transferability beyond the evaluated package configurations remains the key practical question. 该工作用物理引导神经网络预测2.5D chiplet系统热图,并把结果用于信号完整性分析,以避免反复运行完整多物理场仿真。其交叉特征用于捕捉chiplet之间的热流交换,新等效模型则在生成训练数据时支持柱状微凸点。该方法面向更快的设计空间探索,但能否迁移到论文评估范围之外的封装配置,仍是主要工程问题。

AI Accelerators & Compute-in-Memory AI加速器与存算一体

Fengshui Co-Designs Chiplet Ecosystems and Bespoke AI Accelerators Fengshui协同设计chiplet生态与定制AI加速器

arXiv:2609.10970 · 2026-09-10

Fengshui jointly chooses a reusable chiplet pool and composes application-specific accelerators from it, addressing the circular dependency between ecosystem value and accelerator quality. The framework co-explores operator microarchitecture, memory hierarchy, tensor fusion, and pipeline, tensor, and expert parallelism, with place-and-route checks for physical feasibility; its reported pool uses eight selected chiplets. This is a design-space framework rather than measured silicon, so its value depends on how well the cost and physical models predict real implementations. Fengshui联合选择可复用chiplet资源池,并用其组合面向具体应用的加速器,从而处理生态价值与加速器质量之间的循环依赖。该框架协同探索算子微架构、存储层级、张量融合,以及流水线、张量和专家并行,并通过布局布线检查物理可实现性;论文报告的资源池由8种精选chiplet构成。这仍是设计空间框架而非实测硅片,其价值取决于成本与物理模型对真实实现的预测能力。

A Reconfigurable Neuromorphic Core for Biomedical Edge Inference 面向生物医学边缘推理的可重构神经形态核心

Sarah Johari, Suman Kumar, Abhishek Mishra, et al.

arXiv:2609.03174 · 2026-09-02

A programmable FPGA core combines spiking convolutional and fully connected layers, with a PyTorch deployment flow and quantized 16-bit execution. It reaches up to 98% on MNIST, 86% on Fashion-MNIST, and 88.26% average accuracy across five folds for hypoxia classification while consuming 1.455 W of dynamic power. The design demonstrates a complete hardware-software path, although the small benchmarks and FPGA implementation leave ASIC efficiency and broader generalization open. 该可编程FPGA核心结合脉冲卷积层与全连接层,并提供PyTorch部署流程和16位量化执行。它在MNIST上最高达到98%、在Fashion-MNIST上达到86%,用于缺氧分类时5折平均精度为88.26%,动态功耗为1.455W。该设计展示了完整软硬件路径,但小型基准和FPGA实现仍无法回答ASIC效率及更广泛泛化能力的问题。

AI Research & Inference Systems AI研究与推理系统

I/O Lower Bounds and Reinforcement Learning Guide Inference Optimization 以I/O下界与强化学习指导推理优化

ACM Transactions on Architecture and Code Optimization (TACO) · 2026-09-09

A two-level optimizer uses I/O lower-bound theory to choose intra-operator tiling and memory mappings, then reinforcement learning to select fusion boundaries across the graph. For TileAttn, the refined partition theorem raises DRAM-traffic estimation accuracy from 77.5%-82.0% to 86.3%-94.8%, while the overall method cuts shared-memory traffic by 10.79% on average. The combination is attractive for portable tuning, although the reported cross-platform win rate and search cost need to be read against the exact hardware suite. 该两级优化器先用I/O下界理论选择算子内分块与存储映射,再用强化学习确定计算图中的融合边界。在TileAttn上,改进后的划分定理把DRAM流量估计准确率从77.5%-82.0%提高到86.3%-94.8%,整体方法平均降低10.79%的共享内存流量。这种组合有利于跨平台自动调优,但其胜率与搜索成本仍需结合具体硬件测试集解读。

EDA & Design Automation EDA与设计自动化

Reinforcement Learning Optimizes Multichiplet Power Networks 强化学习优化多chiplet供电网络

Wei-Yang Miao, Shaopeng Wang, Wang-Ling Goh, et al.

IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems · 2026-09-01

An RL framework automates redistribution-layer power-network layout for multichiplet packages, using a four-chiplet, 28-power-domain system and a 13x13 design grid. Analytical rewards encode shape and distance-to-bump rules so training does not require full EDA simulation at every step; PPO converges faster, while DDDQN produces higher-quality layouts for the discrete task. The reported designs beat human layouts on impedance, though validation on a wider range of package topologies is needed before treating the policy as general. 该强化学习框架自动生成多chiplet封装的重布线层供电网络布局,测试对象为包含4个chiplet、28个电源域的系统,设计网格为13x13。分析式奖励函数编码形状与到电源凸点距离规则,因此训练无需每一步都运行完整EDA仿真;PPO收敛更快,而DDDQN在离散任务中生成质量更高的布局。论文报告的设计在阻抗上优于人工布局,但在把策略视为通用方法之前,仍需覆盖更多封装拓扑。

Newsletter 邮件订阅

Daily semiconductor briefing. 每日半导体简报。