Devices & Process
器件与工艺
J. Samland, F. Hoff, T. Veslin, et al.
Advanced Optical Materials · 2026-06-23
Samland et al. experimentally demonstrate non-volatile silicon Mach-Zehnder switches using Sb2Se3 phase-change material switched by graphene heaters. The reported device reaches a 0.7 pi phase shift, VpiL of about 0.56 Vcm, and a 28 dB extinction ratio. The result is useful for neuromorphic and reconfigurable PICs because the optical state can persist without continuous heater power, though the path to dense, repeatable arrays still depends on process control and cycling data.
Samland等人实验展示了使用Sb2Se3相变材料、由graphene加热器切换的非易失硅Mach-Zehnder开关。器件实现了0.7π相移、约0.56 Vcm的VπL和28 dB消光比。这个结果对神经形态和可重构PIC有价值,因为光学状态无需持续加热功耗即可保持,但走向高密度、可重复阵列仍取决于工艺控制和循环可靠性数据。
Y. Hu, J. Zhu, Y. Liu, et al.
Microsystems & Nanoengineering · 2026-06-15
Hu et al. co-design the mechanics and optics of a 2x2 horizontal adiabatic-directional-coupler MEMS switch to suppress buckling. The fabricated switches show broad 180 nm bandwidth, roughly 2 microsecond switching, more than 7.2 billion cycles, and a demonstrated 64x64 Benes array. That combination makes the work more relevant to AI optical switching fabrics than a single-device photonics result, although packaging and control overhead remain the scaling tests.
Hu等人对2x2水平绝热方向耦合器MEMS开关进行机光协同设计,以抑制翘曲。实测开关具备180 nm宽带、约2微秒切换速度、超过72亿次循环,并展示了64x64 Benes阵列。这个组合使其相比单器件光子结果更接近AI光交换网络需求,不过封装和控制开销仍是规模化考验。
Circuits & Architecture
电路与架构
M. J. Belda, L. Orlandic, F. Castro, et al.
arXiv:2606.27240 · 2026-06-25T16:24:35Z
Belda et al. evaluate how scratchpad memory and heterogeneous processing elements change CGRA behavior on FFT, GEMM, and a seizure-detection transformer workload. The scratchpad reduces memory traffic by 8x versus a memory-less design, while the homogeneous architecture cuts area overhead by 4.4x to 8.2x compared with prior CGRAs. The paper is useful because it separates data-movement wins from PE-specialization wins instead of treating CGRA efficiency as one knob.
Belda等人在FFT、GEMM和癫痫检测Transformer工作负载上评估scratchpad memory与异构处理单元如何改变CGRA表现。相较无本地存储设计,scratchpad将存储流量降低8倍;而同构架构相对既有CGRA把面积开销降低4.4倍到8.2倍。论文的价值在于把数据搬移收益和PE专用化收益拆开分析,而不是把CGRA效率视为单一旋钮。
T. Zhang, G. Zhang, Y. He, et al.
ACM TODAES · 2026-06-24
Zhang et al. show that placement decisions become first-order when multiple applications share a multi-chip-module GPU. Some workload pairs perform better when co-located on the same GPU chip to maximize memory bandwidth use, while others benefit from being split across chips to reduce contention. That observation maps directly to cloud GPU scheduling as chiplet GPUs become the norm rather than an exotic package choice.
Zhang等人指出,当多个应用共享多芯片模块GPU时,任务放置会成为一阶性能因素。有些工作负载组合放在同一GPU芯片上更利于利用存储带宽,另一些则适合跨芯片分布以降低争用。随着chiplet GPU成为常态而不再是特殊封装选择,这个观察直接映射到云端GPU调度。
AI Accelerators & Compute-in-Memory
AI加速器与存算一体
W. Li, S. Zhu, P. Cuan, et al.
ACM TACO · 2026-06-23
Li et al. target the PCIe and software-path bottleneck that makes small-batch in-kernel ML offload unattractive on dGPUs. Their co-optimization combines feature selection, mixed-precision quantization, lightweight compression, and adaptive transfer scheduling, cutting transferred bytes by 50% to 87.5%. The key result is that the dGPU break-even batch size drops from 256 to 64 and latency falls 3.2x to 5.8x, showing that protocol overhead can dominate raw accelerator throughput.
Li等人针对小批量内核内机器学习卸载到dGPU时的PCIe和软件路径瓶颈。其协同优化结合特征选择、混合精度量化、轻量压缩和自适应传输调度,把传输字节数减少50%到87.5%。关键结果是dGPU盈亏平衡批量从256降到64,端到端延迟降低3.2倍到5.8倍,说明协议开销可能压过原始加速器吞吐。
P. Kumaresan, S. Sivasubramani
arXiv:2606.22635 · 2026-06-21T18:46:33Z
Kumaresan and Sivasubramani present four interface-compatible neuromorphic IP blocks in SkyWater 130 nm: a PVT sensor, stochastic LIF neuron, STDP controller, and memristive-crossbar controller. The blocks share an SPI register file and were verified with 99 cocotb tests at RTL and gate level. It is not a production neuromorphic chip yet, but open, standard-cell blocks lower the barrier for reproducible edge-neuromorphic experiments.
Kumaresan和Sivasubramani在SkyWater 130 nm中提出4个接口兼容的神经形态IP模块:PVT传感器、随机LIF神经元、STDP控制器和memristive crossbar控制器。这些模块共享SPI寄存器文件,并通过99个RTL和门级cocotb测试验证。它还不是量产神经形态芯片,但开放的标准单元模块降低了可复现实验门槛。
I. T. Vidamour, F. Aguirre, T. J. Hayward, et al.
arXiv:2606.23742 · 2026-06-21T15:55:23Z
Vidamour et al. put trainable nonlinear functions on analog neural-network connections rather than treating device nonlinearities as scalar weights. The approach maps well to smooth continuous-control tasks, transfers across about 35,000 hardware connections, and is projected at roughly 30 microwatts in a dedicated CMOS implementation. The caveat is important: the advantage is task-dependent and does not carry over cleanly to classification-like decision boundaries.
Vidamour等人把可训练非线性函数放在模拟神经网络连接上,而不是把器件非线性简单当作标量权重。该方法适合平滑连续控制任务,可迁移到约3.5万个硬件连接,并预计在专用CMOS实现中功耗约30微瓦。需要注意的是,这种优势依赖任务类型,并不能自然迁移到类似分类边界的问题。