semi·news
Headlines要闻 / Research研究 / /
Research digest · Wednesday, August 26, 2026 研究摘要 · 2026年8月26日 星期三

Bandwidth and heat move into the package 带宽与散热走向封装内部

This selection follows the physical constraints behind AI systems: cooling dense packages, moving data across chiplets and fabrics, and making photonic I/O practical. Several papers also test how far memory-aware algorithms and execution systems can stretch limited hardware. 本期论文聚焦AI系统背后的物理约束:高密度封装散热、Chiplet与互连网络间的数据传输,以及光子I/O的工程化。多篇工作也检验了内存感知算法和执行系统能够将有限硬件推向何种程度。

Look-back window: 7 days · 8 paper(s) 回溯窗口: 7天 · 8篇

Devices & Advanced Packaging 器件与先进封装

Generative design of liquid-cooling channels for 2.5D and 3D packages 面向2.5D与3D封装液冷通道的生成式设计

M. Acquah, Zheng Liu

arXiv:2608.22787 · 2026-08-24

This work uses a conditional diffusion model to generate liquid-cooling channel layouts for a 2.7 kW package with two GPUs and one CPU. Of 5,000 designs, 229 passed the final topology screens; the best feasible layout is predicted to lower maximum GPU temperature by 33.6%, temperature spread by 52.5%, and pressure drop by 72.8% versus a reference. The results show how generative methods can explore coupled thermal-hydraulic constraints, though they remain model-based rather than a packaged hardware demonstration. 该工作使用条件扩散模型,为包含两颗GPU和一颗CPU、功耗2.7 kW的封装生成液冷通道布局。在5,000个设计中,229个通过最终拓扑筛选;最佳可行布局相对参考方案预计可将最高GPU温度降低33.6%、温差降低52.5%、压降降低72.8%。结果显示生成式方法能够探索耦合的热—流体约束,但目前仍基于模型而非封装硬件实证。

Graphene-plasmonic photodetectors for high-speed silicon photonics 用于高速硅光的石墨烯等离激元光探测器

D. Rieben, S. Koepfli, D. Bisang, et al.

ACS Photonics · 2026-08-19

The authors demonstrate waveguide-integrated graphene-plasmonic photodetectors with a 155 GHz 3 dB bandwidth and responsivity above 200 mA/W in the L-band. A single lane reaches 192 Gbit/s, while an eight-detector WDM receiver achieves an aggregate 816 Gbit/s. The result addresses a critical receiver-side constraint for dense optical I/O, although manufacturing yield and integration with electronic drivers remain important next questions. 作者展示了波导集成的石墨烯等离激元光探测器,在L波段实现155 GHz的3 dB带宽和超过200 mA/W的响应度。单通道线速率达到192 Gbit/s,八探测器WDM接收机的聚合速率达到816 Gbit/s。该结果针对高密度光I/O的关键接收端约束,但制造良率以及与电子驱动器的集成仍是后续重要问题。

Photonic Circuits & Interconnects 光子电路与互连

Silicon photonics for co-packaged optics 面向共封装光学的硅光子技术

Ying Zhang, Xiangsheng Wang, Z. Kong, et al.

Advanced Materials & Technologies · 2026-08-20

This review connects the component stack of silicon-photonic PICs—lasers, modulators, couplers, Ge-on-Si photodetectors and wavelength multiplexers—to co-packaged optics systems. It frames CPO as a response to bandwidth, power and scalability limits in AI and data-center interconnects. Laser integration, thermal control, reliability, bandwidth density and energy efficiency remain the gaps between device demonstrations and widespread deployment. 该综述将硅光PIC的核心器件栈——激光器、调制器、耦合器、Ge-on-Si光探测器和波分复用器——与共封装光学系统联系起来。文章将CPO定位为应对AI和数据中心互连带宽、功耗与可扩展性限制的路径。激光器集成、热管理、可靠性、带宽密度和能效仍是从器件演示走向广泛部署的关键缺口。

A 2 μm silicon thermo-optic switch using artificial-gauge superlattices 采用人工规范场超晶格的2微米硅热光开关

Xuelin Zhang, Jiangbing Du, Ke Xu, Zuyuan He

ACS Photonics · 2026-08-19

The paper experimentally demonstrates a silicon thermo-optic Mach–Zehnder switch at 2 μm using a waveguide superlattice with an artificial gauge field. It reports 2.17 mW switching power, 1.5 dB insertion loss, crosstalk below −22 dB and a 110 nm operating bandwidth. Those figures make the design relevant to scalable photonic switching, while its system value still depends on integration with sources, detectors and control electronics. 论文实验展示了一种工作在2微米波段的硅热光Mach–Zehnder开关,采用带人工规范场的波导超晶格。其报告的开关功耗为2.17 mW、插入损耗1.5 dB、串扰低于−22 dB,工作带宽达110 nm。这些指标使该设计与可扩展光交换相关,但其系统价值仍取决于与光源、探测器和控制电子学的集成。

AI Accelerators & System Architecture AI加速器与系统架构

An HBM FPGA processor for real-time multilayer holograms 基于HBM FPGA的实时多层全息处理器

Wonok Kwon, Sang-kyu Cheon, Keehoon Hong

ETRI Journal · 2026-08-19

This FPGA-based holographic processor combines HBM with an optimized Fresnel-diffraction formulation to produce eight-layer, 4K holograms at 30 frames per second. It reports 28 ms end-to-end latency and an 86% power reduction against GPU-based implementations. The system is a useful example of HBM enabling a specialized real-time pipeline where memory traffic, not only arithmetic throughput, sets the design limit. 该FPGA全息处理器将HBM与优化的菲涅耳衍射公式结合,可在30帧/秒下生成八层4K全息图。论文报告端到端时延为28 ms,相比GPU实现功耗降低86%。该系统说明,在专用实时流水线中,HBM可缓解由内存流量而不仅是算术吞吐量决定的设计上限。

HYDRA explores heterogeneous chiplets for hybrid LLM serving HYDRA探索用于混合LLM服务的异构Chiplet

Jiahao Lin, Alish Kanani, Sang-Won Lee, et al.

arXiv:2608.19395 · 2026-08-19

HYDRA is a design-space exploration framework for serving Transformer-Mamba hybrid LLMs on heterogeneous chiplet systems. It jointly evaluates chiplet mix, placement, inter-chiplet bandwidth, dynamic batching and scheduling, reporting 1.55x average throughput and 43.7% lower time-to-first-token than the cited baselines. The work argues that runtime policy and physical partitioning must be co-designed, though its gains are based on modeled exploration rather than silicon measurements. HYDRA是一个面向异构Chiplet系统上Transformer-Mamba混合LLM服务的设计空间探索框架。它联合评估Chiplet组合、布局、片间带宽、动态批处理和调度,相比所引用的基线报告平均吞吐量提高1.55倍、首token时间降低43.7%。该工作指出运行时策略与物理划分必须协同设计,但其收益来自建模探索而非实测芯片。

AI Systems & Efficient Inference AI系统与高效推理

MoE router locality meets a memory-bandwidth trade-off MoE路由器局部性遭遇内存带宽权衡

Shriniwas Ramesh Suram

arXiv:2608.18261 · 2026-08-18

This pre-registered study measures a 235B-parameter MoE model on a single 8 GB GPU and finds decode limited by memory bandwidth, at 0.44 tokens per second when active experts must be streamed. Router traces show reuse that could make a small expert cache useful, and locality losses reduce cache misses by as much as 60%. Yet every trained configuration failed the study's predefined perplexity gate, a valuable negative result on the tension between routing locality and model quality. 这项预注册研究在单张8 GB GPU上测量235B参数MoE模型,发现当活跃专家必须流式加载时,解码受内存带宽限制,仅为0.44 token/秒。路由轨迹显示存在可让小型专家缓存受益的复用,局部性损失可将缓存未命中最多降低60%。但所有训练配置均未通过预设的困惑度门槛,这一负结果揭示了路由局部性与模型质量之间的张力。

FreeToken adapts MoE serving to edge hardware FreeToken让MoE服务适配边缘硬件

Shuo Yang, Xiao-yun Fan, Melissa Z. Pan, et al.

arXiv:2608.16157 · 2026-08-17

FreeToken is an edge-native MoE serving system that adapts model placement, expert residency, CPU-GPU execution and memory management to the resources available on a local machine. The authors say it supports more than 20 MoE models, from a 35B model on a laptop to a 284B model on a gaming desktop and GLM-5.2 753B on one workstation GPU. The claim is ambitious and system-dependent, but it focuses attention on bandwidth-adaptive execution as a route to practical local inference. FreeToken是一套面向边缘设备的MoE服务系统,可根据本地机器的资源动态调整模型布局、专家驻留、CPU-GPU执行和内存管理。作者称该系统支持超过20个MoE模型,可在笔记本上运行35B模型、在游戏台式机上运行284B模型,并在单张工作站GPU上运行GLM-5.2 753B。该主张较为激进且依赖具体系统,但它凸显了带宽自适应执行作为实现本地推理的路径。

Newsletter 邮件订阅

Daily semiconductor briefing. 每日半导体简报。