semi·news
Headlines要闻 / Research研究 / /
Research digest · Sunday, June 28, 2026 研究摘要 · 2026年6月28日 星期日

Data Movement Sets the Research Agenda 数据搬移主导研究议程

This week's stronger papers treat memory placement, interconnect overhead, photonic switching, and safety verification as first-order design problems. The common thread is that compute blocks matter less when data movement, topology, and packaging constraints dominate system behavior. 本周较强的论文把存储位置、互连开销、光子交换和安全验证视为一阶设计问题。共同线索是,当数据搬移、拓扑和封装约束主导系统行为时,单个计算模块本身的重要性会下降。

Look-back window: 7 days · 10 paper(s) 回溯窗口: 7天 · 10篇

Devices & Process 器件与工艺

Non-Volatile Silicon Mach-Zehnder Switches Using Graphene Heaters and Sb2Se3 Phase Change Material 基于Graphene加热器和Sb2Se3相变材料的非易失硅Mach-Zehnder开关

J. Samland, F. Hoff, T. Veslin, et al.

Advanced Optical Materials · 2026-06-23

Samland et al. experimentally demonstrate non-volatile silicon Mach-Zehnder switches using Sb2Se3 phase-change material switched by graphene heaters. The reported device reaches a 0.7 pi phase shift, VpiL of about 0.56 Vcm, and a 28 dB extinction ratio. The result is useful for neuromorphic and reconfigurable PICs because the optical state can persist without continuous heater power, though the path to dense, repeatable arrays still depends on process control and cycling data. Samland等人实验展示了使用Sb2Se3相变材料、由graphene加热器切换的非易失硅Mach-Zehnder开关。器件实现了0.7π相移、约0.56 Vcm的VπL和28 dB消光比。这个结果对神经形态和可重构PIC有价值,因为光学状态无需持续加热功耗即可保持,但走向高密度、可重复阵列仍取决于工艺控制和循环可靠性数据。

Scalable Silicon Photonic MEMS Switches with Quasi-Buckling-Free Directional Couplers 采用准无翘曲方向耦合器的可扩展硅光子MEMS开关

Y. Hu, J. Zhu, Y. Liu, et al.

Microsystems & Nanoengineering · 2026-06-15

Hu et al. co-design the mechanics and optics of a 2x2 horizontal adiabatic-directional-coupler MEMS switch to suppress buckling. The fabricated switches show broad 180 nm bandwidth, roughly 2 microsecond switching, more than 7.2 billion cycles, and a demonstrated 64x64 Benes array. That combination makes the work more relevant to AI optical switching fabrics than a single-device photonics result, although packaging and control overhead remain the scaling tests. Hu等人对2x2水平绝热方向耦合器MEMS开关进行机光协同设计,以抑制翘曲。实测开关具备180 nm宽带、约2微秒切换速度、超过72亿次循环,并展示了64x64 Benes阵列。这个组合使其相比单器件光子结果更接近AI光交换网络需求,不过封装和控制开销仍是规模化考验。

Circuits & Architecture 电路与架构

Scratchpad Memory and Heterogeneity Trade-offs in CGRAs CGRA中Scratchpad Memory与异构性的权衡

M. J. Belda, L. Orlandic, F. Castro, et al.

arXiv:2606.27240 · 2026-06-25T16:24:35Z

Belda et al. evaluate how scratchpad memory and heterogeneous processing elements change CGRA behavior on FFT, GEMM, and a seizure-detection transformer workload. The scratchpad reduces memory traffic by 8x versus a memory-less design, while the homogeneous architecture cuts area overhead by 4.4x to 8.2x compared with prior CGRAs. The paper is useful because it separates data-movement wins from PE-specialization wins instead of treating CGRA efficiency as one knob. Belda等人在FFT、GEMM和癫痫检测Transformer工作负载上评估scratchpad memory与异构处理单元如何改变CGRA表现。相较无本地存储设计,scratchpad将存储流量降低8倍;而同构架构相对既有CGRA把面积开销降低4.4倍到8.2倍。论文的价值在于把数据搬移收益和PE专用化收益拆开分析,而不是把CGRA效率视为单一旋钮。

Location-Aware Scheduling for Multitasking MCM-GPUs 面向多任务MCM-GPU的位置感知调度

T. Zhang, G. Zhang, Y. He, et al.

ACM TODAES · 2026-06-24

Zhang et al. show that placement decisions become first-order when multiple applications share a multi-chip-module GPU. Some workload pairs perform better when co-located on the same GPU chip to maximize memory bandwidth use, while others benefit from being split across chips to reduce contention. That observation maps directly to cloud GPU scheduling as chiplet GPUs become the norm rather than an exotic package choice. Zhang等人指出,当多个应用共享多芯片模块GPU时,任务放置会成为一阶性能因素。有些工作负载组合放在同一GPU芯片上更利于利用存储带宽,另一些则适合跨芯片分布以降低争用。随着chiplet GPU成为常态而不再是特殊封装选择,这个观察直接映射到云端GPU调度。

AI Accelerators & Compute-in-Memory AI加速器与存算一体

Protocol and Software Co-Optimization for dGPU In-Kernel ML 面向dGPU内核内机器学习的协议与软件协同优化

W. Li, S. Zhu, P. Cuan, et al.

ACM TACO · 2026-06-23

Li et al. target the PCIe and software-path bottleneck that makes small-batch in-kernel ML offload unattractive on dGPUs. Their co-optimization combines feature selection, mixed-precision quantization, lightweight compression, and adaptive transfer scheduling, cutting transferred bytes by 50% to 87.5%. The key result is that the dGPU break-even batch size drops from 256 to 64 and latency falls 3.2x to 5.8x, showing that protocol overhead can dominate raw accelerator throughput. Li等人针对小批量内核内机器学习卸载到dGPU时的PCIe和软件路径瓶颈。其协同优化结合特征选择、混合精度量化、轻量压缩和自适应传输调度,把传输字节数减少50%到87.5%。关键结果是dGPU盈亏平衡批量从256降到64,端到端延迟降低3.2倍到5.8倍,说明协议开销可能压过原始加速器吞吐。

A Neuromorphic Silicon Suite in SkyWater 130 nm 基于SkyWater 130 nm的神经形态硅IP套件

P. Kumaresan, S. Sivasubramani

arXiv:2606.22635 · 2026-06-21T18:46:33Z

Kumaresan and Sivasubramani present four interface-compatible neuromorphic IP blocks in SkyWater 130 nm: a PVT sensor, stochastic LIF neuron, STDP controller, and memristive-crossbar controller. The blocks share an SPI register file and were verified with 99 cocotb tests at RTL and gate level. It is not a production neuromorphic chip yet, but open, standard-cell blocks lower the barrier for reproducible edge-neuromorphic experiments. Kumaresan和Sivasubramani在SkyWater 130 nm中提出4个接口兼容的神经形态IP模块:PVT传感器、随机LIF神经元、STDP控制器和memristive crossbar控制器。这些模块共享SPI寄存器文件,并通过99个RTL和门级cocotb测试验证。它还不是量产神经形态芯片,但开放的标准单元模块降低了可复现实验门槛。

Low-Power Analogue Neural Networks with Trainable Nonlinear Connections 具备可训练非线性连接的低功耗模拟神经网络

I. T. Vidamour, F. Aguirre, T. J. Hayward, et al.

arXiv:2606.23742 · 2026-06-21T15:55:23Z

Vidamour et al. put trainable nonlinear functions on analog neural-network connections rather than treating device nonlinearities as scalar weights. The approach maps well to smooth continuous-control tasks, transfers across about 35,000 hardware connections, and is projected at roughly 30 microwatts in a dedicated CMOS implementation. The caveat is important: the advantage is task-dependent and does not carry over cleanly to classification-like decision boundaries. Vidamour等人把可训练非线性函数放在模拟神经网络连接上,而不是把器件非线性简单当作标量权重。该方法适合平滑连续控制任务,可迁移到约3.5万个硬件连接,并预计在专用CMOS实现中功耗约30微瓦。需要注意的是,这种优势依赖任务类型,并不能自然迁移到类似分类边界的问题。

AI Research for Hardware 面向硬件的AI研究

Physically Constrained Agents for Hardware-Aware Foundation Model Compression 面向硬件约束基础模型压缩的物理约束智能体

J. Zhang, W. Sun, C. Wang, et al.

arXiv:2606.25532 · 2026-06-24T08:07:59Z

Zhang et al. build a physically grounded multi-agent discovery system for hardware-compliant foundation-model deployment. The generated methods include Q-Enhance for long-context dense models and MoE-Salient-AQ for sparse MoE quantization, with the paper claiming a 235B-parameter model deployed on dual A100s after a 75% memory reduction and 0.64% accuracy loss. The result is relevant because compression is framed around hardware constraints, but the claims need independent reproduction before they should guide accelerator roadmaps. Zhang等人构建了一个具备物理约束的多智能体发现系统,用于符合硬件约束的基础模型部署。生成的方法包括面向长上下文稠密模型的Q-Enhance,以及面向稀疏MoE量化的MoE-Salient-AQ;论文声称在内存需求降低75%、精度损失0.64%的情况下,将235B参数模型部署到双A100服务器。该结果重要在于把压缩问题放入硬件约束中讨论,但相关主张仍需独立复现后才适合作为加速器路线图依据。

EDA & Verification EDA与验证

SafeGen Uses LLMs and Formal Verification for Functional-Safety Assertions SafeGen用LLM和形式验证生成函数安全断言

X. Tan, A. Chaudhuri, R. Parekhji, et al.

arXiv:2606.25296 · 2026-06-24T01:58:38Z

Tan et al. propose SafeGen, an LLM-driven framework that generates functional-safety assertions and evaluates fault criticality with formal verification support. The system links safety documents, FMEDA guidance, RTL information, and generated assertions through a document-level Hyper Knowledge Graph. The value is traceability rather than text generation alone, but automotive-grade adoption will depend on how well the generated assertions survive adversarial and coverage audits. Tan等人提出SafeGen,这是一个由LLM驱动、并结合形式验证来生成函数安全断言和评估故障关键性的框架。系统通过文档级Hyper Knowledge Graph把安全文档、FMEDA指南、RTL信息和生成断言连接起来。其价值不只是文本生成,而是可追溯性;不过要进入汽车级流程,还要看生成断言能否经受对抗性检查和覆盖率审计。

Quantum & Photonics 量子与光子

A Photonic Integrated Long-Distance Quantum Communication Network 光子集成长距离量子通信网络

L. Zhang, J. Pan, T.-Y. Chen, et al.

Nature Photonics · 2026-06-19

Zhang et al. report a photonic integrated long-distance quantum communication network in Nature Photonics. The candidate metadata does not expose detailed performance numbers, but the venue and topic make it relevant to integrated quantum links rather than lab-bench optics alone. The paper belongs on the reading queue because photonic integration is one of the few credible routes to scaling quantum networking hardware. Zhang等人在Nature Photonics报道了一个光子集成长距离量子通信网络。候选元数据没有给出详细性能数字,但从期刊和主题看,它更接近集成量子链路,而不只是实验台光学演示。该论文值得列入阅读队列,因为光子集成是扩展量子网络硬件的少数可信路径之一。