semi·news
Headlines要闻 / Research研究 /
Research digest · Wednesday, May 27, 2026 研究摘要 · 2026年5月27日 星期三

After Scaling: 2D Channels, Orbital Currents, and Photonic Tensorcores 工艺微缩之后:2D沟道、轨道电流与光子张量核

This week's literature is unusually device-physics-heavy — Ta/W spin-orbit torque MRAM, 2D semiconductor circuit integration, domain-free negative capacitance — alongside system-level work that ties AI supercomputers to grid response and pushes interconnect networks past Fat-Tree limits. 本周文献异常偏向器件物理——Ta/W自旋轨道力矩MRAM、2D半导体电路集成、无畴负电容——同时也有系统级工作:将AI超算与电网响应耦合,并把互连网络推过Fat-Tree拓扑的极限。

Look-back window: 7 days · 10 paper(s) 回溯窗口: 7天 · 10篇

Devices & Process 器件与工艺

Orbital and spin-orbit torque interplay in Ta/W magnetic tunnel junctions with vertical non-local switching Ta/W磁隧道结中轨道与自旋轨道力矩的相互作用与垂直非局域开关 preprintMRAMdevice physics

co-authors not extracted

MRAM research consortium

arXiv:2605.27215 · 2026-05-26

SOT-MRAM has been stalled at ~45% charge-to-spin conversion efficiency (ξ_DL), well below the ~80% needed to compete with advanced transistor nodes for cache memory. This work reports a four-fold increase in spin-orbit torque from a Ta/W bilayer system by exploiting orbital current contributions from Ta — moving the technology meaningfully closer to the write-energy figures cache-level MRAM actually needs. The result is a device-physics step, not a tape-out, but it changes which materials systems show up in mainstream SOT-MRAM roadmaps over the next 24 months. SOT-MRAM长期卡在约45%的电荷-自旋转换效率(ξ_DL),远低于与先进晶体管节点竞争所需的约80%(用于缓存级存储)。本文报告利用Ta层的轨道电流贡献,Ta/W双层体系将自旋轨道力矩提升4倍——使该技术显著靠近缓存级MRAM真正需要的写入能耗指标。这是器件物理层级的进展、并非量产流片,但会改变未来24个月主流SOT-MRAM路线图上出现的材料体系。

Chips in the Flatland: 2D semiconductors for future computing electronics Flatland中的芯片:面向未来计算电子学的2D半导体 preprint2D materialsreview

review, multi-author

academic consortium

arXiv:2605.26555 · 2026-05-26

Honest review of the 2D-semiconductor circuit-integration gap — the 'valley of death' between elegant single-device demos (MoS2, WSe2, hBN gate stacks) and actually-fab-able functional integrated circuits at sub-Angstrom-era nodes. The paper tracks where each material family has stalled (contact resistance, dielectric integration, wafer-scale uniformity) and where it's moved (5-stage logic in MoS2, single-flake DRAM cells). Useful as a calibration against vendor 2D-semi marketing — the bottleneck is rarely the channel material itself. 对2D半导体电路集成「鸿沟」的一次诚实综述——介于单器件优雅演示(MoS2、WSe2、hBN栅堆)与亚埃米代际可实际流片的功能集成电路之间的「死亡之谷」。文章追踪了各材料体系卡在哪里(接触电阻、介质集成、晶圆级均匀性)以及取得了哪些进展(MoS2中5级逻辑、单片DRAM单元)。可作为校准2D半导体厂商营销话术的参照——瓶颈很少是沟道材料本身。

Conditions for domain-free negative capacitance 无畴负电容的条件 preprintnegative capacitanceferroelectric

co-authors not extracted

academic team

arXiv:2605.26536 · 2026-05-26

Negative-capacitance FETs have been a perpetually-promising route to sub-60-mV/dec subthreshold slope, but every NC-FET demonstration to date shows ferroelectric domain formation — which the original theory said should be absent. This paper derives the critical domain-wall energy parameter above which a ferroelectric-dielectric heterostructure stabilizes in a true domain-free NC state. The practical contribution: a clear materials selection criterion (which HfZrO compositions and which dielectric pairings will work) rather than just another fitting attempt. Closer to a foundry-relevant integration recipe than most prior NC-FET work. 负电容FET长期被视为通往亚60 mV/dec亚阈值斜率的可能路径,但迄今为止所有NC-FET演示都显示存在铁电畴——而原始理论认为应不存在。本文推导出临界畴壁能参数:高于该值时铁电-介质异质结将稳定在真正的无畴NC态。实质贡献:给出明确的材料选择准则(哪种HfZrO组分与哪种介质配对能用),而不只是又一次拟合。比此前多数NC-FET工作更接近代工厂相关的集成配方。

Architecture & Accelerators 架构与加速器

Cassandra: enabling reasoning LLMs at the edge via self-speculative decoding Cassandra:通过自我推测解码在边缘运行推理型LLM preprintLLM inferenceedge

co-authors not extracted

academic team

arXiv:2605.26558 · 2026-05-26

Reasoning-style LLMs (the o1 / deepseek-R1 family) suffer disproportionately at decode time because their long chain-of-thought traces multiply per-token overhead. Cassandra proposes an algorithm-hardware co-design where the draft model is constructed by selectively skipping layers of the target — eliminating the need for a separate draft model entirely — and validates that this preserves accuracy while delivering meaningful speedups on consumer-class accelerator hardware. The implication: reasoning models may become deployable on phone-class NPUs sooner than the parameter count suggests, if the decode-side bottleneck is attacked directly. 推理型LLM(o1 / deepseek-R1家族)由于长思维链导致每token开销倍增,在解码阶段尤其受拖累。Cassandra提出算法-硬件协同设计:通过选择性跳过目标模型的层来构建草稿模型,从而完全省去独立草稿模型;论文验证该方案在保持精度的同时,在消费级加速器硬件上带来可观加速。含义:如果解码侧瓶颈被直接攻击,推理型模型或许会比参数规模所暗示的更早可在手机级NPU上部署。

Architectural limits of cloud TPUs in finite-field cryptography 云TPU在有限域密码学中的架构极限 preprintTPUcryptography

co-authors not extracted

academic team

arXiv:2605.25367 · 2026-05-25

Empirically characterizes a 5,558×-6,908× cost-efficiency deficit for TPUs vs A100 GPUs on finite-field cryptography (the workload behind zero-knowledge proofs and modern PKI). Roots the gap analytically in the absence of wide-integer ALUs on TPU silicon, plus a spatial penalty in multi-tenant Montgomery reduction. The paper matters because it's the first principled statement of a hardware category where Google's most expensive in-house silicon is structurally uncompetitive — useful evidence for anyone modelling whether custom-silicon design teams should generalize their datapaths or stay narrowly AI-focused. 实证刻画了TPU相对A100 GPU在有限域密码学(零知识证明与现代PKI背后的负载)上5,558-6,908倍的成本效率差距。分析层面将该差距归因为TPU硅片缺乏宽整数ALU,以及多租户Montgomery规约带来的空间惩罚。论文意义在于:这是首次有原则地指明一类硬件场景——谷歌最贵的自有硅片在结构上不具备竞争力。对于思考定制硅片设计团队应通用化数据通路还是专注AI的人,这是有用证据。

Datacenter & Systems 数据中心与系统

GridPilot: real-time grid-responsive control for AI supercomputers GridPilot:面向AI超算的实时电网响应控制 preprintgrid integrationHPC

co-authors not extracted

academic team

arXiv:2605.26384 · 2026-05-25

Implements a three-tier predictive controller (millisecond / second / hour scales) that translates grid operator requests into actual GPU power changes at the facility meter — with a deterministic safety-island bypass for fast-response events. Validated on real hardware on a three-stage testbed. The contribution is that this is the first published end-to-end measurement of how fast an AI/HPC datacenter can actually deliver grid-flexibility services, which is increasingly a precondition for interconnection at sites above ~100 MW. Direct relevance to anyone building hyperscale capacity in PJM, ERCOT or the UK grid. 实现一个三层预测控制器(毫秒/秒/小时三种时间尺度),将电网调度请求转化为机房计量表处实际GPU功率变化——并配有用于快速响应事件的确定性安全岛旁路。论文在真实硬件三阶段测试床上验证。意义在于:这是首次端到端测量AI/HPC数据中心实际可提供电网灵活性服务的速度——而在装机容量超过约100 MW的站点,这正在成为电网接入的前置条件。与在PJM、ERCOT或英国电网建设超大规模产能的团队直接相关。

Extreme-scale interconnection networks: Multipass Random Leaf-Spine beats Fat-Tree 极大规模互连网络:多通路随机叶脊网络优于Fat-Tree preprintinterconnectnetwork topology

co-authors not extracted

academic team

arXiv:2605.26960 · 2026-05-26

Argues — with simulation data — that Multipass Random Leaf-Spine (MRLS) topologies, which combine Orthogonal Fat-Tree and Random Folded Clos ideas, outperform Fat-Tree at the endpoint counts (>1M GPUs) that hyperscalers now actually plan for. Reports better throughput at lower wiring cost, plus higher flexibility under partial-failure conditions. Practical relevance: this is the kind of paper that quietly shows up two years later as the design rationale in the next Broadcom or Nvidia switch-fabric reference design. Worth tracking, even if the specific topology doesn't win. 通过仿真数据论证:多通路随机叶脊(MRLS)拓扑结合了正交Fat-Tree与随机折叠Clos的设计,在超大规模客户实际规划的端点数(>1M GPU)下优于Fat-Tree。论文报告更高吞吐量、更低布线成本,以及在部分故障下更强的灵活性。实际意义:这类论文通常会在两年后悄然出现在博通或英伟达下一代交换织物参考设计的设计依据中。即便该特定拓扑最终未被采用,本文也值得跟踪。

Photonics & Interconnects 光子学与互连

DUET: a general-purpose photonic computing primitive for contemporary AI DUET:面向当代AI的通用光子计算原语 preprintphotonic computeONN

co-authors not extracted

academic team

arXiv:2605.23051 · 2026-05-21

Proposes Dynamic Universal Encoding Tensorcore (DUET), a photonic computing paradigm built on vectorized operand differential interferometric cells (VODICs). The interesting move is supporting dynamic arbitrary matrix operations — the thing optical neural networks have historically struggled with — by exploiting structural symmetry to cut the hardware overhead that has kept ONNs from scaling. If the silicon-photonics process windows hold, this is a candidate for the kind of photonic primitive that finally puts optical compute into a co-packaged AI accelerator stack rather than a research demo. 提出动态通用编码张量核(DUET),一种基于向量化操作数差分干涉单元(VODIC)的光子计算范式。亮点在于支持动态任意矩阵运算——这正是光神经网络(ONN)此前最难突破的能力——利用结构对称性削减让ONN长期无法规模化的硬件开销。若硅光子工艺窗口能够支撑,这有望成为真正把光计算放入共封装AI加速器堆栈、而非只是研究演示的光子原语候选。

Alignment-free ultra-broadband parametric frequency conversion in lead-halide perovskites 无对准、超宽带的卤化物钙钛矿参量频率转换 preprintperovskitefrequency conversion

co-authors not extracted

academic team

arXiv:2605.25718 · 2026-05-25

Demonstrates four-wave mixing across an exceptionally wide near-IR / mid-IR tuning range in thick single-crystal lead-halide perovskite, with no phase-matching engineering, no angular alignment, and no dispersion optimization required. The practical implication for chip-scale photonics: if perovskite blocks can be integrated as wavelength-conversion stages without the alignment / dispersion plumbing that today's silicon photonics demands, the wavelength count usable per CPO link could go up substantially. Still single-crystal lab work, but unusually friendly to scale-up. 在厚单晶卤化物钙钛矿中演示近红外/中红外极宽带四波混频——无需相位匹配工程、无需角度对准、无需色散优化。对芯片级光子学的实际意义:若钙钛矿模块可作为波长转换级被集成,而不需要当下硅光子所要求的对准/色散水管,则共封装光学(CPO)链路上可用波长数可显著上升。目前仍是单晶实验室级别的工作,但少见地对规模化友好。

Emerging Computing 新兴计算

Co-designing graph-based approximate nearest-neighbor search at billion scale for processing-in-memory 面向存内计算的十亿级图近似最近邻搜索协同设计 preprintPIMANN search

co-authors not extracted

academic team

arXiv:2605.25522 · 2026-05-25

Graph-based ANNS is fundamentally bandwidth-bound — exactly the workload PIM was designed for — but every prior attempt to port it hit the same architectural mismatches: tiny per-PU memory, costly inter-PU communication, host coordination overhead. This work treats PIM and ANNS as a joint design target rather than software-on-hardware, and reports a usable billion-scale PIM-ANNS pipeline. Most useful as a concrete answer to 'when does PIM actually beat host-DRAM by enough margin to justify the integration headache?' for AI-retrieval workloads. 基于图的近似最近邻搜索(ANNS)本质上受带宽约束——正是PIM被设计来处理的负载——但此前每次将其移植到PIM都遇到相同的架构错配:单处理单元内存太小、跨单元通信开销高、主机协调开销大。本文将PIM与ANNS视为联合设计目标,而非「硬件上的软件」,并报告了可用的十亿级PIM-ANNS流水线。对「在AI检索负载下,PIM到底何时能以足够幅度跑赢主机DRAM以抵消集成代价」这个问题,是难得的具体答案。