semi·news
Headlines要闻 / Research研究 / /
Research digest · Thursday, August 27, 2026 研究摘要 · 2026年8月27日 星期四

Thermal limits meet alternative compute architectures 热管理限制遇上新型计算架构

Recent work tackles heat flow from 2.5D/3D packages down to material interfaces, while photonic and time-domain designs pursue different efficiency paths. The common challenge is translating promising physical mechanisms into reproducible systems. 近期研究从2.5D/3D封装到材料界面处理热流问题,同时光子和时域设计探索不同的能效路径。共同挑战在于将有前景的物理机制转化为可复现的系统。

Look-back window: 7 days · 6 paper(s) 回溯窗口: 7天 · 6篇

Electronic Design Automation 电子设计自动化

IC-ThermBench: a progressive benchmark for 2.5D/3D-IC thermal learning IC-ThermBench:面向2.5D/3D-IC热学习的渐进式基准

David Huang, Wenkai Yang, Kuiye Ding, Haiyang Xin

arXiv:2608.23977 · 2026-08-25

IC-ThermBench introduces an open benchmark for learning-based thermal modeling across 3D ICs, industrial packages, and a 50,000-sample 2.5D chiplet extension. Its baselines show a sharp generalization drop on unseen package systems: the best reported RMSE rises from 1.216 K to 15.99 K under cross-package transfer. The result shows that thermal ML still needs stronger out-of-distribution validation before broad use in package design. IC-ThermBench提出面向3D IC、工业封装及包含5万样本的2.5D chiplet扩展集的开放热建模基准。其基线在未见封装系统上的泛化明显下降:最佳RMSE在跨封装迁移时从1.216 K升至15.99 K。该结果表明,热学机器学习在被广泛用于封装设计之前,仍需更严格的分布外验证。

Devices, Materials & Packaging 器件、材料与封装

DAF-less 3D stacking combines back-grinding and silane chemistry 无DAF三维堆叠结合背磨与硅烷化学

Gyu-Sik Park, Suk Jekal, Woohyeon Kim, et al.

Bulletin of the Korean Chemical Society · 2026-08-18

The authors demonstrate die-attach-film-less multi-die stacking using back-grinding roughness control and silane coupling chemistry. Their APG12-treated fine-ground interface achieved a roughly 0.31 μm gap and 4.2 MPa shear strength, close to a polished reference, while thermal imaging indicated better heat transfer than DAF stacking. It is a process-oriented route to reduce polymer interlayers, but long-term reliability remains unproven. 作者利用背磨粗糙度控制和硅烷偶联化学,演示了无die attach film的多芯片堆叠。经APG12处理的精磨界面实现约0.31 μm间隙和4.2 MPa剪切强度,接近抛光参考值;热成像显示其传热优于DAF堆叠。这为减少聚合物中间层提供了面向工艺的路径,但长期可靠性仍待证明。

Modeling cure evolution and thermal endurance in filled epoxy underfill 高填充环氧底部填充胶的固化演化与热耐久建模

Ryan Giang, Ran Tao, Elena Moukhina, et al.

Journal of Polymer Science · 2026-08-16

This work combines DSC, TGA, and diffusion-aware kinetic modeling to predict curing and degradation in highly filled epoxy underfill. A two-step modified Kamal-Sourour model better captured conversion after vitrification and informed an oven cure that reached nearly complete conversion. The approach could improve packaging process windows, although it still needs correlation with interconnect reliability over service life. 该研究结合DSC、TGA和考虑扩散的动力学建模,预测高填充环氧底部填充胶的固化与降解。两步修正Kamal-Sourour模型能更好地描述玻璃化后的转化率,并用于设计实现近乎完全固化的烘箱曲线。该方法有望改善封装工艺窗口,但仍需与互连在服役期内的可靠性建立对应关系。

AI Accelerators & Compute-in-Memory AI加速器与存算一体

Programmable photonic neural engine with 40,000 connections 具备4万连接的可编程光子神经引擎

Hao Sun, Xinyi Zhu, J. Azaña

Science Advances · 2026-08-19

The paper reports a photonic neuromorphic engine that integrates all-optical nonlinearity into a loop-based, time-multiplexed architecture. It demonstrates four network topologies and up to 40,000 optical connections, with reported latency more than two orders of magnitude below state-of-the-art electronic processors. System precision, control, and packaging overhead remain central deployment questions. 论文报道了一种光子神经形态引擎,将全光非线性集成到基于环路的时分复用架构中。该系统演示了四种网络拓扑,最多支持4万条光连接,并报告其延迟比先进电子处理器低两个数量级以上。系统级精度、控制和封装开销仍是部署的核心问题。

Time-domain CIM accelerator for AdderNet with hardware-error-aware training 面向AdderNet并具硬件误差感知训练的时域CIM加速器

Aoming Zhan, Ye Zhao, Yumei Zhou, Shushan Qiao

Applied Sciences · 2026-08-17

This time-domain compute-in-memory accelerator maps AdderNet’s L1 operations to minimum selection and time-domain accumulation, using dual-mode DTC encoding and shared-clock TDC readout. Post-layout simulations in 55 nm cover PVT variation, jitter, offsets, and quantization rather than assuming ideal conversion. That error-aware treatment is useful, but the evidence remains simulation rather than measured silicon. 该时域存算一体加速器将AdderNet的L1运算映射为最小值选择和时域累加,并采用双模式DTC编码及共享时钟TDC读出。其55 nm后仿真覆盖PVT变化、抖动、偏移和量化,而非假设理想转换。这样的误差感知处理有参考价值,但当前证据仍是仿真而非实测芯片。

AI Systems & Hardware Co-Design AI系统与软硬件协同设计

FPGA model compression and hardware-aware acceleration: a co-design taxonomy FPGA模型压缩与硬件感知加速:协同设计分类法

Peter Forcha, H. Kajekusumadhar, Mbua Peter, et al.

arXiv:2608.21657 · 2026-08-21

This survey organizes 25 FPGA deep-learning co-design studies from 2015–2026 into five categories, from DSP elimination to memory hierarchy and deployment tooling. It finds that only one study reports compression and accuracy against a common baseline, and only two normalize energy efficiency against a common GPU baseline. The taxonomy warns that cross-paper FPGA comparisons remain poorly standardized. 这篇综述将2015至2026年的25项FPGA深度学习协同设计研究分为五类,涵盖DSP消除、存储层级和部署工具等。它发现只有一项研究在统一基准下同时报告压缩率和准确率,只有两项研究相对统一GPU基准归一化能效。该分类法提醒我们,跨论文FPGA比较仍缺乏标准化。

Newsletter 邮件订阅

Daily semiconductor briefing. 每日半导体简报。