semi·news
Headlines要闻 / Research研究 / /
Research digest · Monday, August 24, 2026 研究摘要 · 2026年8月24日 星期一

Memory Boundaries Become Design Variables 存储边界成为设计变量

This week's work moves state across devices, packages, accelerators, GPUs, and edge nodes instead of treating memory as a fixed hierarchy. The strongest results pair that movement with measured hardware or explicit kernel and thermal constraints. 本周研究不再把存储视为固定层级,而是在器件、封装、加速器、GPU与边缘节点之间重新安排状态。更有分量的工作同时给出了实测硬件,或明确纳入内核与热约束。

Look-back window: 7 days · 10 paper(s) 回溯窗口: 7天 · 10篇

Devices & Process 器件与工艺

Gate-Stack Engineering Improves 40 nm IGZO Vertical Charge-Trap Memory 栅堆叠工程改善40 nm IGZO垂直电荷俘获存储器

Y. Jung, C.-S. Hwang, S.-M. Yoon

ACS Applied Materials & Interfaces · 2026-08-21

An a-IGZO vertical charge-trap memory transistor uses an H2O-based Al2O3 tunneling layer and nanocrystalline ZnO traps to improve charge injection and erase behavior. Devices optimized at 60 nm and validated at a 40 nm channel length produce an approximately 10 V memory window. The result is relevant to tighter 3D NAND vertical scaling, although array-level endurance, uniformity, and manufacturability are not established in the candidate record. 该a-IGZO垂直电荷俘获存储晶体管采用H2O基Al2O3隧穿层与纳米晶ZnO俘获层,以改善电荷注入和擦除行为。器件在60 nm条件下优化,并在40 nm沟道长度下验证,获得约10 V存储窗口。结果与3D NAND更紧凑的垂直缩放相关,但候选资料尚未给出阵列级耐久性、一致性与可制造性。

A Dual-Ferroelectric Stack Adds Reconfigurable Optoelectronic Memory 双铁电堆叠实现可重构光电存储

P. Wang, B. Jiang, W. Xue, et al.

Nature Communications · 2026-08-19

A stacked dual-ferroelectric transistor switches polarization through electrical and optical inputs, combining electronic storage, multilevel optical memory, and visual-response emulation in one device. It reports more than 8-bit conductance states, low dark current, and a wide analog dynamic range. The artificial-vision demonstration is simulated, so array scalability and system energy remain open questions despite the device-level functionality. 该双铁电堆叠晶体管通过电输入与光输入翻转极化,在单个器件中结合电子存储、多级光存储和视觉响应模拟。器件报告了超过8 bit的电导状态、低暗电流和较宽的模拟动态范围。人工视觉演示基于仿真,因此尽管器件功能丰富,阵列扩展性与系统能耗仍待验证。

Scanless Confocal Metrology Speeds 3D Package Profiling 308x 无扫描共焦测量将3D封装轮廓检测提速308倍

C. Huang, D. Kong, J. Liang, et al.

Measurement Science and Technology · 2026-08-20

A chromatic differential arrayed-confocal microscope reconstructs package-substrate topography without lateral or axial mechanical scanning. It cuts full-resolution acquisition from 20 seconds to 65 milliseconds, a 308x throughput gain, while step-height and surface-roughness deviations stay below 3% and 1% versus a commercial laser confocal tool. The comparison is on selected substrates, so robustness across reflective materials and production defect modes still needs broader validation. 该色差阵列共焦显微系统无需横向或轴向机械扫描即可重建封装基板三维形貌。其全分辨率采集时间从20秒降至65毫秒,吞吐量提升308倍;相较商用激光共焦设备,台阶高度与表面粗糙度偏差分别低于3%和1%。目前比较基于选定基板,面对不同反射材料与量产缺陷模式的稳健性仍需更广泛验证。

Circuits & Architecture 电路与架构

A Sub-Milliwatt ASIC Decodes EEG Auditory Attention 亚毫瓦ASIC解码EEG听觉注意力

Q. Ma, R. George, S. Scholze, et al.

arXiv:2608.20198 · 2026-08-20

A GF22FDX 22 nm ASIC combines a quantized CNN engine with a Pearson-correlation classifier for real-time EEG auditory-attention decoding. The implementation reports 0.4941 mW at 0.55 V, 7.34 ms inference latency, and 2.09 mm² total area; the inference and classifier engines occupy 0.076 mm². The preprint describes a full implementation, but the candidate record does not establish measured post-silicon accuracy or power. 这款GF22FDX 22 nm ASIC将量化CNN引擎与Pearson相关分类器结合,用于实时EEG听觉注意力解码。实现结果为0.55 V下功耗0.4941 mW、推理延迟7.34 ms、总面积2.09 mm²,其中推理与分类引擎占0.076 mm²。预印本给出了完整实现,但候选资料未能确认流片后的实测精度与功耗。

AI Accelerators & Compute-in-Memory AI加速器与存算一体

APEX Exploits Dual Sparsity for Precise SNN Inference APEX利用双重稀疏性实现高精度SNN推理

D. B. Venkatesh, S. Radhakrishnan, R. Rakshit, et al.

arXiv:2608.19046 · 2026-08-19

APEX integrates the PASC-IF neuron into the LoAS accelerator framework so converted SNNs can preserve source-ANN accuracy at fewer timesteps. Its three-stage neuron datapath is fully combinational, adding no pipeline latency, while the architecture exploits sparsity in both activations and weights. This is an accelerator-design preprint rather than a measured-silicon report, so area and energy claims should be read as implementation results. APEX把PASC-IF神经元集成到LoAS加速器框架中,使转换后的SNN能以更少时间步保持源ANN精度。其三级神经元数据通路采用全组合逻辑,不增加流水线延迟,同时利用激活与权重的双重稀疏性。该工作是加速器设计预印本而非实测芯片报告,因此面积与能效结论应视为实现结果。

A Micro-LED Array Programs IGZO RRAM in Parallel 微型LED阵列并行编程IGZO RRAM

A. Adair, J. Robertson, A. Tsiamis, et al.

arXiv:2608.16807 · 2026-08-17

A compact micro-LED array optically programs multiple form-free IGZO RRAM devices in parallel, using 450 nm light for SET and electrical pulses for RESET. Persistent photocurrent provides fading-memory behavior, and the prototype writes spatial patterns across several devices. The laboratory demonstration connects optical fan-out with neuromorphic state storage, but larger-array uniformity, endurance, and programming energy are not yet shown. 紧凑型微型LED阵列可并行光编程多个无需forming的IGZO RRAM器件,以450 nm蓝光执行SET、电脉冲执行RESET。持续光电流提供了衰减记忆行为,原型还在多个器件上写入空间图案。该实验把光学扇出与神经形态状态存储连接起来,但尚未展示更大阵列的一致性、耐久性与编程能耗。

Hardware-Relevant AI Research 面向硬件的AI研究

Pallas Migrates KV Cache Before Cellular Handover Pallas在蜂窝切换前迁移KV cache

T. Ding, J. Liu, H. Xu

arXiv:2608.16477 · 2026-08-17

Pallas prepares LLM inference state at a predicted target base station before cellular handover, avoiding a cold transfer after the user moves. It reconstructs the stable prefix locally while streaming newly generated suffix KV blocks from the source, then assembles a current cache at handoff. The approach trades interruption time for prediction and duplicate-compute overhead, so its benefit depends on handover accuracy and available edge capacity. Pallas在蜂窝切换发生前,把LLM推理状态预先准备到预测的目标基站,避免用户移动后再进行冷迁移。系统在目标端本地重建稳定前缀,同时从源端流式传输新生成后缀的KV块,并在切换时组装为最新cache。该方案以预测开销和重复计算换取更短中断时间,因此收益取决于切换预测准确率与边缘算力余量。

DynamoServe Pools Stranded GPU Memory for Multi-Tenant LLMs DynamoServe汇聚闲置GPU显存服务多租户LLM

D. Z. Tootaghaj, K. Diab, B. Lantz, et al.

ACM Digital Library · 2026-08-17

DynamoServe pools stranded GPU memory to hold model weights and KV caches for multi-tenant LLM serving. Coordinated placement and demand-driven weight migration reduce fragmentation while trying to preserve locality and latency. The reported evaluation improves memory efficiency without a latency penalty, though the candidate record provides no workload-level figures for judging gains across cluster topologies. DynamoServe汇聚分散在GPU上的闲置显存,用于存放多租户LLM服务的模型权重与KV cache。协同放置与按需权重迁移减少资源碎片,同时尽量保持局部性和延迟。论文报告显存效率提升且没有延迟惩罚,但候选资料未提供按工作负载拆分的数据,难以判断不同集群拓扑下的收益。

FluxBin Co-Designs Binary Quantization With a CUDA LUT Kernel FluxBin协同设计二值量化与CUDA查找表内核

Q. Yang, R. Yang, H. Xiao, et al.

arXiv:2608.15602 · 2026-08-16

FluxBin pairs post-training binary decomposition with a CUDA lookup-table kernel to avoid floating-point arithmetic and runtime dequantization in ultra-low-bit LLM inference. Hessian-guided salient bases preserve critical weights, while virtual column mapping turns irregular work into denser execution; the paper reports up to 5.92x speedup. Results are tied to the tested CUDA kernels and models, so portability and accuracy at broader model scales remain the main caveats. FluxBin把训练后二值分解与CUDA查找表内核结合,在超低比特LLM推理中避开浮点运算和运行时反量化。Hessian引导的显著性基底保留关键权重,虚拟列映射则把不规则计算转化为更稠密的执行;论文报告最高5.92倍加速。结果与测试所用CUDA内核和模型绑定,跨平台可移植性及更大模型上的精度仍是主要限制。

EDA & Design Automation EDA与设计自动化

COOL Accelerates Thermal Prediction for 3D and 3.5D Packages COOL加速3D与3.5D封装热预测

Y. Lu, Z. Guo, Q. Zhang, et al.

ACM Digital Library · 2026-08-16

COOL represents dies, interposers, thermal-interface materials, heat spreaders, and cooling structures as annotated 3D point clouds for package thermal prediction. A physics-informed boundary loss enforces material-interface and cooling constraints; on the authors' benchmark it reaches 2.4% normalized mean absolute error and runs more than 15.7x faster than a commercial FEM solver. Because the benchmark is constructed by the authors, generalization to unseen package geometries and boundary conditions needs independent testing. COOL把裸片、中介层、热界面材料、散热盖与冷却结构表示为带属性的3D点云,用于封装热预测。物理信息边界损失约束材料界面与冷却边界;在作者构建的基准上,其归一化平均绝对误差为2.4%,速度比商用FEM求解器快15.7倍以上。由于基准由作者自行构建,面对未见过的封装几何与边界条件时的泛化能力仍需独立测试。

Newsletter 邮件订阅

Daily semiconductor briefing. 每日半导体简报。