semi·news
Headlines要闻 / Research研究 /
Research digest · Tuesday, May 26, 2026 研究摘要 · 2026年5月26日 星期二

The Memory Hierarchy Goes Analog, Ferroelectric, and Photonic 存储层级转向模拟、铁电与光子化

This week's literature converges on a single theme: as silicon scaling slows, the cheapest path to AI throughput is changing where computation happens — moving it into memory, into ferroelectric devices, into photonic stacks, and even into probabilistic Ising hardware. 本周文献汇聚于一个主题:硅工艺微缩放缓后,AI吞吐量最便宜的增长路径正在改变 — 计算正在被搬入存储、铁电器件、光子堆叠,乃至概率Ising硬件之中。

Look-back window: 7 days · 12 paper(s) 回溯窗口: 7天 · 12篇

Devices & Process 器件与工艺

Hybrid ferroelectric-ionic memristive hardware for high-scalability in-memory computing 用于高可扩展存内计算的铁电-离子复合忆阻器硬件 journalin-memory computingferroelectric

Tian et al.

multi-institution (CN/EU)

Nature Communications · 2026-05-21

The team integrates a ferroelectric switching layer with ionic-modulation conductance tuning to break the precision/retention trade-off that has dogged single-mechanism memristors. They report 8-bit analog programmability with retention measured in months at array-relevant temperatures and demonstrate a 64×64 crossbar running matrix-vector products at energy/op figures competitive with leading PCM and ReRAM benchmarks. Caveat: the work is still cell-array scale; tape-out-grade variability statistics (cycle-to-cycle, device-to-device) at million-cell density are not in this paper. 团队将铁电开关层与离子调制电导调节相结合,打破了单机制忆阻器中精度与保持的折中。他们报告了8位模拟可编程性,阵列相关温度下保持时间以月计,并在64×64交叉阵列上演示矩阵-向量乘法,每次操作能耗指标可与领先PCM和ReRAM对标。需注意:该工作仍是单元阵列层级,关于百万单元密度下的量产级波动统计(cycle-to-cycle,device-to-device)尚未在本文给出。

Stable analog weight programming in single-crystalline van der Waals ferroelectric transistors 单晶范德华铁电晶体管中的稳定模拟权重编程 journalferroelectric2D materials

Park, Lee, Kim et al.

KAIST / partners

Advanced Materials · 2026-05-20

Most analog-weight ferroelectric FETs degrade rapidly under repeated programming because polycrystalline domains drift. This work uses an exfoliated single-crystalline 2D ferroelectric channel to lock in 6-bit weight states with <2% drift over 10⁶ write cycles. The result is one of the cleanest demonstrations to date that 2D ferroelectrics can hit the endurance numbers compute-in-memory accelerators actually need — but practical foundry integration (uniformity across 300 mm wafers) remains an open problem. 多数模拟权重铁电FET在反复编程下因多晶畴漂移而迅速退化。本文采用剥离单晶2D铁电沟道,将6位权重在10⁶次写入循环后漂移控制在<2%。这是迄今为止最干净的演示之一:2D铁电材料能达到存内计算加速器所需的耐久性指标 — 但量产代工集成(300 mm晶圆均匀性)仍是悬而未决的问题。

Memory & Storage 存储

Emerging memory technologies at room and cryogenic temperatures 室温与低温下的新兴存储技术综述 preprintsurveyMRAM

review chapter, multi-author

academic consortium

arXiv:2605.21912 · 2026-05-21

A book-chapter-style survey covering volatile and non-volatile candidates across both regimes — STT/SOT-MRAM, FeRAM, PCM, RRAM at room temperature, and superconducting/quantum-compatible memories at cryogenic. The value here isn't novel results but the side-by-side framing: which technologies actually move into HBM-class roles vs. which stay tied to neuromorphic or quantum-control niches. Useful as a reference for arguing why cryo-CMOS work has bifurcated from mainstream memory roadmaps despite a decade of cross-pollination claims. 一篇章节体例的综述,覆盖室温与低温两种环境下的易失/非易失候选 — 室温侧的STT/SOT-MRAM、FeRAM、PCM、RRAM,低温侧的超导/量子兼容存储。本文价值不在原创结果,而在并排对比的框架:哪些技术真正能进入HBM类角色、哪些注定停留在神经形态或量子控制利基。对于解释为何低温CMOS研究在十年互相借鉴后仍与主流存储路线图分叉,是一份有用的引用文献。

CIM-AD: hardware-efficient SRAM-based computing-in-memory accelerator with sparse matrix in diagonal shift CIM-AD:基于SRAM的硬件高效存内计算加速器(对角移位稀疏矩阵) journalin-memory computingSRAM

co-authors not extracted

CN academic team

Journal of Circuits, Systems and Computers · 2026-05-22

An SRAM CIM macro that handles sparse-matrix MAC by storing weights in a diagonal-shifted layout, eliminating the index-table overhead that has historically eaten the energy savings sparse CIM was meant to deliver. The authors report ~2× energy/op improvement on Transformer MHA workloads vs. dense-baseline CIM, with no accuracy loss at typical 70-90% block sparsity. Caveat: SRAM-CIM still loses to DRAM-anchored solutions on capacity per mm², so this remains a near-cache play, not a main-memory replacement. 一种SRAM存内计算宏,通过对角移位布局存储权重处理稀疏矩阵MAC,消除了原本侵蚀稀疏CIM能耗优势的索引表开销。作者报告在Transformer MHA工作负载上每次操作能耗较密集基线CIM改善约2倍,在70-90%典型块稀疏度下精度无损。需注意:SRAM-CIM在每平方毫米容量上仍输给以DRAM为锚的方案,因此仍只是靠近缓存的方案,并非主存替代。

Architecture & Accelerators 架构与加速器

Provisioning to runtime optimization of a +100 MW AI cluster 百兆瓦级AI集群从规划到运行时的功率优化 preprintdatacenterGB200

industry infra team

hyperscaler (anonymized)

arXiv:2605.24461 · 2026-05-23

First end-to-end public description of power management for a hyperscale AI datacenter — covering 6-12 month pre-GA capacity planning for next-gen accelerators, deployment-time tuning, and dynamic runtime power capping for evolving workloads. The paper publishes detailed power telemetry for a 150 MW datacenter running 83K GB200 GPUs, including the gap between nameplate TDP and actual sustained draw under real training mixes. For anyone modeling datacenter capex or grid impact: this is the first time the inside view has been made citable. 首次端到端公开描述超大规模AI数据中心的电源管理过程 — 涵盖下一代加速器GA前6-12个月的容量规划、部署期调优,以及面向不断变化负载的运行时动态限功。论文公布了运行83K GB200 GPU的150 MW数据中心的详细功率遥测,包括铭牌TDP与真实训练负载下持续功耗之间的差距。对于建模数据中心资本开支或电网影响的人:内部视角首次以可引用形式公开。

MX-SAFE: versatile inference- and training-proof microscaling format with on-the-fly exponent/mantissa allocation MX-SAFE:兼顾推理与训练的可变指数/尾数微缩放格式 preprintnumerical formatsOCP MX

Korean academic team

Yonsei University and partners

arXiv:2605.24391 · 2026-05-23

Builds on OCP's MX format by letting each block dynamically choose between MXINT-style (precision-favored) and MXFP-style (range-favored) bit allocation, with the routing decided per block from the actual distribution. Claims training stability comparable to FP8 and inference accuracy matching MXINT8, in a single hardware unit. If this lands in OCP's next revision, it removes the format split that has been a source of friction between training-side and inference-side accelerator design. 在OCP的MX格式基础上,让每个块依据实际分布动态在MXINT(重精度)与MXFP(重动态范围)之间分配位数,在单一硬件单元中同时实现。论文称其训练稳定性可比FP8,推理精度匹配MXINT8。若该方案进入OCP下一版规范,将消除训练侧与推理侧加速器设计之间因格式分裂带来的摩擦。

EVA: accelerating LLM decoding via an efficient vector-quantization architecture EVA:面向LLM解码的高效向量量化架构 preprintLLM inferencequantization

co-authors not extracted

academic team

arXiv:2605.24144 · 2026-05-22

Targets the GEMV-bound decode stage of LLM inference with a vector-quantization (VQ) datapath that combines on-chip codebook reuse with a memory-conflict-avoiding access scheduler. The reported 2-bit weight compression matches accuracy of 4-bit RTN baselines on Llama-3 70B while delivering ~1.7× tokens/sec on the same memory bandwidth. The architecture is silicon-implementable as an inference-pod accelerator front-end; how it integrates with KV-cache pressure at long-context settings (>32k tokens) is the next open question. 针对LLM推理中受GEMV带宽约束的解码阶段,提出向量量化(VQ)数据通路,结合片上码本复用与避免访存冲突的调度器。论文报告在Llama-3 70B上,2位权重压缩可媲美4位RTN基线的精度,同等存储带宽下token吞吐量提升约1.7倍。架构可作为推理pod前端硅化实现;其在长上下文(>32k token)下与KV缓存压力的协同,是下一个待解问题。

DiSC: resolution-scalable acceleration of diffusion models via sparsity and cached-token reuse DiSC:通过稀疏与缓存token复用实现分辨率可扩展的扩散模型加速 preprintdiffusionASIC

co-authors not extracted

academic team

arXiv:2605.25798 · 2026-05-25

Two intertwined contributions: Cached Token Reuse uses inter-step latent similarity to skip token recomputation, and Softmax Thresholding propagates sparsity masks across iterations to bypass redundant attention work. Together they cut diffusion model inference cost ~3× at high resolutions with negligible FID loss. Notable because it attacks both the iterative-step bottleneck and the quadratic-attention bottleneck at the architecture level rather than as a software optimization, suggesting a clean ASIC target for text-to-image inference at scale. 两项相互嵌入的贡献:缓存token复用(CTR)利用步间隐空间相似性跳过token重计算,Softmax阈值化(ST)将稀疏掩码跨迭代传播以绕过冗余注意力计算。两者结合可在高分辨率下将扩散模型推理成本下降约3倍,FID几乎无损。值得关注之处在于:该工作在架构层而非软件层同时打击迭代步骤瓶颈与二次注意力瓶颈,提示了规模化文生图推理ASIC的清晰目标。

Photonics & Interconnects 光子学与互连

All-band photonic-integrated optical parametric amplification 全波段光子集成参量放大 preprintsilicon photonicsamplifier

EPFL Photonic Systems Lab and partners

EPFL

arXiv:2605.22704 · 2026-05-21

Demonstrates a chip-integrated optical parametric amplifier with usable gain spanning the O-, S-, C-, L- and U-bands — covering essentially all telecom and datacom wavelength windows on one device. Pump-power and chip-length requirements are within commercial silicon-photonics process windows, which is the interesting part: prior OPAs at this breadth needed bulk crystals or kilometers of nonlinear fiber. If this productizes, it would let co-packaged optics scale wavelength count without stacking multiple gain materials. 演示了一种芯片集成光参量放大器,可用增益带宽覆盖O、S、C、L、U波段 — 在单一器件上几乎覆盖所有电信与数通波长窗口。泵浦功率与芯片长度需求处于商业硅光子工艺窗口之内,这正是关键之处:以往达到同等带宽的OPA需要体材料晶体或公里级非线性光纤。若实现产品化,将使共封装光学能够在不堆叠多种增益材料的前提下扩展波长数。

Wideband balanced photodetectors for classical and quantum light detection from optical to EUV/X-rays 覆盖光学至EUV/X射线波段的宽带平衡光电探测器 preprintEUVphotodetector

co-authors not extracted

academic team

arXiv:2605.24199 · 2026-05-22

Solves the junction-capacitance / noise trade-off that has kept balanced photodetection out of the EUV and soft X-ray regimes — exactly the wavelengths that matter for next-gen lithography metrology and EUV-source diagnostics. A bootstrapped transimpedance amplifier feeding a low-noise JFET front-end pushes the noise floor down enough to enable quantum-limited measurements at sub-100 nm wavelengths. Direct relevance to ASML/Zeiss metrology stacks for High-NA EUV nodes. 解决了将平衡光电探测扩展至EUV与软X射线波段时一直存在的结电容/噪声折中 — 这正是下一代光刻计量与EUV光源诊断关注的波长范围。自举式跨阻放大器配合低噪声JFET前端,将噪声底降低到足以在亚100 nm波长实现量子极限测量。对ASML/蔡司在High-NA EUV节点的计量栈具有直接相关性。

Emerging Computing 新兴计算

P-dit probabilistic Ising machine for solving the quadratic assignment problem 用p-dit概率Ising机求解二次分配问题 preprintprobabilistic computingIsing

co-authors not extracted

academic team

arXiv:2605.24408 · 2026-05-23

Extends probabilistic-bit (p-bit) Ising machines to multi-state probabilistic d-dimensional variables, mapping the quadratic assignment problem onto silicon-implementable p-dit hardware. Reports time-to-solution improvements over GPU-based simulated annealing on benchmark instances. Notable as part of a broader trend: probabilistic-computing hardware is starting to clear the bar where the energy/answer ratio actually beats von-Neumann + SA for NP-hard combinatorial workloads — the next 18 months will tell whether this translates into supply-chain or routing problems anyone in industry actually runs. 将概率比特(p-bit)Ising机扩展为多状态、多维概率变量(p-dit),把二次分配问题映射到可硅化实现的p-dit硬件上。基准实例上较GPU模拟退火(SA)在求解时间上有明显改进。意义在于一个更大趋势:概率计算硬件开始在NP难组合问题上实现'每个答案能耗'低于冯诺依曼+SA的门槛 — 未来18个月将验证这是否能落地为工业界真正在跑的供应链或路由问题。

Energy-aware computing in the year 2026 2026年的能耗感知计算 preprintenergyHPC

European HPC consortium

multi-institution EU

arXiv:2605.24569 · 2026-05-23

A position-and-survey paper that argues the carbon and grid-side constraints on HPC and generative AI have crossed a threshold where energy-aware scheduling, not raw FLOPS, is now the binding optimization target. Walks through job-level, system-level and continuum-level (edge-cloud-HPC) levers and where each has measurable impact. Most useful as ammunition for budget conversations with procurement and energy teams; the technical detail is survey-grade rather than novel. 一篇立场+综述论文,主张HPC与生成式AI面临的碳排与电网约束已越过门槛,能耗感知调度(而非裸FLOPS)已成为绑定型的优化目标。文章梳理了任务级、系统级以及边-云-HPC连续体层级的杠杆,及各自可量化的影响。其最大价值是为采购与能源团队的预算对话提供弹药;技术细节属综述层级而非原创。