semi·news
Headlines要闻 / Research研究 / /
Research digest · Sunday, August 23, 2026 研究摘要 · 2026年8月23日 星期日

Verification and Scheduling Move Up the Stack 验证与调度持续向栈上层延伸

This week's systems papers push verification, placement, tracing, and service-level scheduling into more integrated tool flows. Device work on ScAlN and GaN provides a measured counterpoint at the physical layer. 本周系统论文将验证、布局、追踪和服务等级调度纳入更集成的工具流程。ScAlN与GaN器件研究则以实测结果补足物理层视角。

Look-back window: 7 days · 9 paper(s) 回溯窗口: 7天 · 9篇

Devices & Process 器件与工艺

Sc-Composition Sequencing Controls the Structural Stability of Sputter-Epitaxial ScAlN/GaN Multilayers Sc组分排序控制溅射外延ScAlN/GaN多层结构稳定性

N. Mimasaka, K. Ikeda, M. Naito, et al.

Japanese Journal of Applied Physics · 2026-08-20

ScAlN multilayers grown on GaN/6H-SiC remained coherently strained in-plane whether Sc content stepped down toward about 6% or up toward about 22%. The increasing-Sc sequence produced more roughness, needle-like features, and c-axis deviation, while decreasing Sc was structurally more stable. The result is measured by high-resolution X-ray diffraction and reciprocal-space mapping, but it covers short four-step growth sequences under one sputtering condition rather than a manufacturing window. 在GaN/6H-SiC上生长的ScAlN多层结构,无论Sc含量逐步降至约6%还是升至约22%,均保持面内相干应变。Sc递增序列出现更明显的粗糙度、针状表面特征和c轴偏离,而Sc递减序列结构更稳定。结果来自高分辨率X射线衍射和倒易空间映射实测,但仅覆盖单一溅射条件下的四段短时生长,尚不能代表完整制造窗口。

GaN HEMTs Remove Reverse-Recovery Spikes in a 130 kHz WPT Inverter GaN HEMT消除130 kHz无线供电逆变器的反向恢复尖峰

M. Bogdanović, Ž. Despotović, D. Marčetić, et al.

Electronics · 2026-08-18

A measured 130 kHz L-S-tuned GaN full-bridge inverter commutates through third-quadrant conduction during dead time and eliminates the reverse-recovery current spikes seen with Si IGBTs. The comparison isolates GaN's zero-reverse-recovery behavior from system-level variables, making dead-time tuning the central trade-off between shoot-through risk and conduction loss. The evidence is hardware-based, but it comes from one inverter topology and operating regime. 一套实测的130 kHz L-S调谐GaN全桥逆变器在死区期间通过第三象限导通完成换流,消除了Si IGBT方案中的反向恢复电流尖峰。该比较将GaN零反向恢复特性与系统级变量分离,使死区调节成为直通风险与导通损耗之间的核心权衡。证据来自真实硬件,但目前仅覆盖一种逆变器拓扑和工作区间。

Circuits, Architecture & Reliability 电路、架构与可靠性

Performance Verification Across Four Generations of AmpereOne 跨四代AmpereOne核心的性能验证方法

D. Al-Otoom, N. Kelly, A. Lindsay, et al.

arXiv:2608.19300 · 2026-08-19T16:50:02Z

Ampere describes a pre-silicon verification flow used across four generations of AmpereOne, correlating cycle-accurate RTL against a trace-driven performance model. Daily regressions, curated workloads, and a unified event stream let teams isolate units such as branch prediction and L2 prefetch before full-core correlation. The paper is valuable as an industrial methodology report, though the supplied abstract gives no quantitative defect-escape or schedule comparison against alternative flows. Ampere介绍了一套用于四代AmpereOne的流片前性能验证流程,通过周期精确的RTL与轨迹驱动性能模型进行关联。每日回归、精选工作负载和统一事件流,使团队可先隔离分支预测与L2预取等单元,再进行整核关联。该文的价值主要在工业方法论,但现有摘要未给出相对其他流程的缺陷漏检率或进度改善数据。

An Open RISC-V Trace Encoder With N-Trace and E-Trace Back Ends 支持N-Trace与E-Trace后端的开放RISC-V追踪编码器

A. Weiss, A. Schulz

arXiv:2608.18170 · 2026-08-17T16:13:02Z

CTTE provides a synthesizable SystemVerilog trace encoder with one protocol-agnostic front end and selectable RISC-V N-Trace/Nexus or E-Trace back ends. It has been integrated with six cores from five suppliers and tested on 64-bit Linux systems, including two-hart SMP; source-side process filtering doubled observation depth in a fixed buffer. The open implementation fills a practical tooling gap, although the result remains a preprint rather than a standardized reference implementation. CTTE提供一套可综合的SystemVerilog追踪编码器,以统一的协议无关前端连接可选的RISC-V N-Trace/Nexus或E-Trace后端。该实现已集成到5家供应商的6款核心,并在包括双hart SMP在内的64位Linux系统上测试;源端进程过滤使固定缓冲区中的观察深度翻倍。它填补了开放工具链中的实际空白,但目前仍是预印本,并非标准参考实现。

AI Accelerators & Compute-in-Memory AI加速器与存算一体

ODEONN: A Digital ODE Solver for Oscillatory Neural Networks ODEONN:面向振荡神经网络的数字ODE求解架构

B. F. Haverkort, A. Todri-Sanial

arXiv:2608.20110 · 2026-08-20T14:39:55Z

ODEONN is a modular fully digital architecture for oscillatory neural networks that supports complex-valued coupling rather than a single fixed application. Its sine approximation uses half the hardware resources of standard methods, with less than 2% degradation versus full-precision software, and the reported energy-delay product is 45× lower than conventional software execution. The comparison is against software rather than taped-out silicon, so implementation efficiency still needs physical validation. ODEONN是一种模块化的全数字振荡神经网络架构,支持复数耦合,而非仅面向单一固定应用。其正弦近似相较标准方法节省一半硬件资源,相对全精度软件的性能下降低于2%,报告的能量延迟积则比传统软件执行低45倍。当前对比对象是软件而非流片芯片,因此实现效率仍需物理验证。

APEX Exploits Dual Sparsity for Precise SNN Inference APEX利用双重稀疏性实现精确SNN推理

D. B. Venkatesh, S. Radhakrishnan, R. Rakshit, et al.

arXiv:2608.19046 · 2026-08-19T15:38:16Z

APEX maps the PASC-IF neuron into the LoAS accelerator framework to preserve ANN-equivalent accuracy at fewer spiking-inference timesteps. A three-stage PASC-IF datapath is implemented as combinational logic without added latency, while the architecture exploits sparsity in two dimensions. The supplied abstract is truncated before quantitative efficiency results, and the preprint does not establish measured-silicon performance. APEX将PASC-IF神经元映射到LoAS加速器框架中,以更少的脉冲推理时间步保持与ANN等效的精度。三级PASC-IF数据通路采用组合逻辑实现,不增加额外延迟,同时架构利用两个维度的稀疏性。候选摘要在定量能效结果前被截断,且预印本尚未证明实测芯片性能。

AI Research for Hardware 面向硬件的AI研究

Multi-Tier SLA Scheduling for LLM Serving 面向LLM服务的多层SLA调度

A. Vestrum, A. Raeesi, H. Roed

arXiv:2608.16336 · 2026-08-17T09:43:51Z

This work extends Llumnix from two priority levels to an arbitrary number of service tiers using per-tier headroom, tier-aware dispatch, and migration. Evaluation covers uniform, Gaussian, and enterprise-style priority mixes in the Vidur simulator against INFaaS, vLLM, Orca, and Sarathi-Serve baselines. The design better matches production SLA hierarchies, but all evidence in the supplied record is simulation-based rather than from a live inference cluster. 该研究将Llumnix从两个优先级扩展为任意数量的服务层级,引入分层余量、层级感知调度与迁移机制。评估在Vidur模拟器中覆盖均匀、高斯和企业型优先级分布,并与INFaaS、vLLM、Orca及Sarathi-Serve比较。该设计更贴近生产环境的SLA层级,但候选记录中的证据全部来自模拟,而非真实推理集群。

EDA & Verification EDA与验证

Congestion-Aware NoC Placement and Routing for FPGAs 面向FPGA的拥塞感知NoC布局与路由

S. G. Shahrouz, V. Betz

arXiv:2608.17266 · 2026-08-18T01:51:14Z

The authors add NoC link-congestion cost and turn-model packet routing directly to VPR's placement flow for NoC-equipped FPGAs. Jointly considering programmable routing and packet-path diversity lets the placer reject oversubscribed network choices earlier, with evaluation across 29 benchmarks. The candidate abstract is truncated before the full numerical result, and the work remains an open-flow preprint rather than a production-tool study. 作者将NoC链路拥塞成本和转向模型分组路由直接加入面向NoC FPGA的VPR布局流程。通过联合考虑可编程布线与数据包路径多样性,布局器可更早排除链路过度订阅的选择,并在29个基准上进行评估。候选摘要在完整数值结果前被截断,且该工作仍是开放流程预印本,并非生产级工具研究。

GoalEvolve Optimizes Physical-Design Algorithms Against Final QoR GoalEvolve以最终QoR驱动物理设计算法优化

H. Liu, L. Zhou, Y. Ren, et al.

arXiv:2608.16733 · 2026-08-17T15:45:33Z

GoalEvolve evolves physical-design algorithms against end-to-end quality-of-results targets instead of optimizing each stage in isolation. It identifies the dominant target gap, traces it to a flow stage, and uses full-flow evaluation to retain only changes that survive downstream effects; across eight ASAP7 designs, post-route total negative slack improved by 30.67% on average. The gains come from benchmark designs in a research flow, so transfer to commercial nodes and production constraints remains unproven. GoalEvolve以端到端QoR目标演化物理设计算法,而不是孤立优化各个阶段。它识别最主要的目标缺口,将其定位到具体流程阶段,并通过全流程评估只保留下游仍有效的改动;在8个ASAP7设计上,布线后总负裕量平均改善30.67%。这些增益来自研究流程中的基准设计,能否迁移到商用制程和生产约束仍待验证。

Newsletter 邮件订阅

Daily semiconductor briefing. 每日半导体简报。