semi·news
Headlines要闻 / Research研究 / /
Research digest · Sunday, June 21, 2026 研究摘要 · 2026年6月21日 星期日

Hardware Specialization Spreads Beyond Matrix Multiplication 硬件专用化走出矩阵乘法

This week's work stretches specialized hardware across hybrid FeFET memory, privacy-enforced training, polynomial arithmetic, photonics, and KV-cache movement. Most results are still preprints, so implementation detail and silicon evidence matter more than headline novelty. 本周研究把专用硬件扩展到混合模式FeFET存储、隐私约束训练、多项式运算、光子系统和KV cache迁移。多数结果仍是预印本,因此实现细节与硅验证比标题中的新颖性更重要。

Look-back window: 7 days · 10 paper(s) 回溯窗口: 7天 · 10篇

Devices & Process 器件与工艺

A 4T Differential FeFET Cell with Volatile and Non-Volatile Modes 兼具易失与非易失模式的4T差分FeFET单元

J. Wang, W. Zhang, X. Fong

arXiv:2606.19918 · 2026-06-18T08:11:19Z

The paper designs a four-transistor differential bit-cell from cross-coupled FeFETs plus two access transistors, with volatile and non-volatile operation selected through write conditions. It reports 0.13 microwatts of store power and a 2 ns store time without explicit backup-and-restore. The result is a preprint-level cell proposal; the abstract does not establish a fabricated array, endurance, retention, or manufacturability. 论文以交叉耦合FeFET和两个访问晶体管构成4T差分位单元,并通过写入条件选择易失或非易失工作模式。结果给出0.13微瓦存储功耗和2 ns存储时间,且无需显式备份与恢复。该成果仍是预印本级单元方案,摘要没有证明已制成阵列,也未交代耐久性、保持时间或可制造性。

Circuits & Architectures 电路与体系结构

Phase-Modulation Tradeoffs for Testable Photonics and Co-Packaged Optics 面向可测试光子系统与共封装光学的相位调制权衡

P. Agnihotri, P. Kalla, S. Blair

arXiv:2606.19674 · 2026-06-18T00:53:05Z

This study compares thermal phase tuning with carrier-based electrical modulation in Mach-Zehnder and microring devices for integrated test and calibration. It evaluates extinction ratio, tuning efficiency, power, bandwidth, and controllability, clarifying when a slower thermal path is preferable to a faster electrical one. The abstract provides design guidance but no headline measurements, so the strength of the comparison depends on the full evaluation methodology. 研究比较了Mach-Zehnder与微环器件中的热相位调谐和载流子电调制,用于片上测试与校准。论文从消光比、调谐效率、功耗、带宽和可控性展开评估,说明何时较慢的热调谐优于更快的电调制。摘要给出了设计取舍,但没有列出关键测量值,因此比较的可信度仍取决于完整评估方法。

AIA: A 16 nm Multi-Core RISC-V SoC for Discrete Sampling AIA:面向离散采样的16 nm多核RISC-V SoC

S. Zhao, N. Shah, W. Meert, et al.

arXiv:2606.16143 · 2026-06-15T02:59:45Z

AIA is a fabricated 16 nm SoC with a RISC-V host and a 2D mesh of 16 custom RISC-V cores for discrete sampling and approximate inference. Custom instructions and datapaths target non-normalized Knuth-Yao sampling and nonlinear-function interpolation, attacking workloads that map poorly to conventional CPUs and GPUs. The abstract confirms fabrication but does not expose throughput or energy results, which are necessary to judge the specialization cost. AIA是一款采用16 nm工艺制造的SoC,包含RISC-V主机和由16个定制RISC-V内核组成的二维网格,用于离散采样与近似推理。其定制指令和数据通路面向非归一化Knuth-Yao采样及非线性函数插值,处理传统CPU和GPU映射效率不高的工作负载。摘要确认了芯片制造,但未给出吞吐率或能效结果,因此尚难判断专用化代价。

AI Accelerators & Compute-in-Memory AI加速器与存算一体

An Integer-Only Transformer Framework for Versal AI Engines 面向Versal AI Engine的纯整数Transformer框架

G. Koski, S. Lipps, Z. Ma, et al.

arXiv:2606.17500 · 2026-06-16T04:22:06Z

The work maps a quantized, integer-only transformer for CERN jet tagging onto AMD Versal AI Engine tiles, including dense and multi-head-attention layers. Its main contribution is an open-source framework that turns high-level model descriptions into composable AIE blocks and generated Vitis graph code. This is an initial implementation without latency, accuracy, or resource figures in the abstract, so it is more reusable infrastructure than a finished accelerator result. 该工作把用于CERN喷注标记的量化纯整数Transformer映射到AMD Versal AI Engine tile,覆盖全连接层和多头注意力层。主要贡献是一个开源框架,可将高层模型描述转换为可组合AIE模块并生成Vitis图代码。它仍是初步实现,摘要未给出时延、精度或资源占用数据,因此更接近可复用基础设施,而非成熟的加速器结果。

DataGuard Enforces Differential-Privacy Budgets in Systolic Accelerators DataGuard在脉动阵列加速器中强制执行差分隐私预算

P. K. Sanjaya, C. Giannoula, N. Shreekumar, et al.

arXiv:2606.16809 · 2026-06-15T14:53:13Z

DataGuard adds a hardware control path to systolic-array training accelerators so that only outputs satisfying an owner's differential-privacy budget may leave the device. That removes the assumption that a third-party federated-learning application correctly clips gradients and adds noise. The reported evaluation is simulation-based rather than measured silicon, making area, performance, and integration overhead the key open questions. DataGuard为脉动阵列训练加速器增加硬件控制路径,只允许符合数据所有者差分隐私预算的输出离开设备。这消除了对第三方联邦学习应用会正确裁剪梯度并添加噪声的信任假设。其评估基于仿真而非实测芯片,因此面积、性能和集成开销仍是关键问题。

MPX Unifies Matrix and Polynomial Multiplication in One Systolic Array MPX在同一脉动阵列中统一矩阵与多项式乘法

G. Alexakis, D. Schoinianakis, G. Dimitrakopoulos

arXiv:2606.16394 · 2026-06-15T08:29:37Z

MPX extends a conventional systolic array to execute both matrix multiplication and direct polynomial multiplication for FHE and post-quantum cryptography. The dual-mode additions cost 20% more area with negligible power overhead during matrix work, while polynomial mode reports more than 1.2 times lower latency than NTT-based mapping on matrix engines. These are experimental design results in a preprint, not a measured production chip. MPX扩展传统脉动阵列,使其同时执行矩阵乘法以及面向FHE和后量子密码的直接多项式乘法。双模式设计增加20%面积,在矩阵运算时功耗开销可忽略;多项式模式相对在矩阵引擎上映射NTT的方法,时延降低超过1.2倍。这些是预印本中的实验性设计结果,并非量产芯片实测。

AI Research & Inference Systems AI研究与推理系统

SwiftCache Shares KV Cache Across Heterogeneous Models over NVLink SwiftCache通过NVLink在异构模型间共享KV cache

J. Hu, M. Xu, S. Wang, et al.

arXiv:2606.16135 · 2026-06-15

SwiftCache lets models with low KV-cache demand donate idle GPU memory to high-demand models, moving shared prefixes over NVLink instead of reloading them over PCIe from CPU memory or SSD. It also keeps only the active layer's cache locally, creating room for longer multi-turn contexts. The candidate abstract excerpt does not include measured speedups or capacity gains, so workload balance and cross-model interference remain the practical questions. SwiftCache让KV cache需求较低的模型把空闲GPU显存借给高需求模型,并通过NVLink传输共享前缀,避免经PCIe从CPU内存或SSD重新加载。它还只在本地保留当前活动层的cache,为更长的多轮上下文腾出空间。候选摘要未包含实测加速比或容量增益,因此负载均衡与跨模型干扰仍是实际问题。

SMEPilot Schedules LLM Operators Across Arm SME and CPU Cores SMEPilot在Arm SME与CPU内核间调度LLM算子

F. Chen, H. Chen

arXiv:2606.16332 · 2026-06-15

SMEPilot uses a roofline characterization to choose CPU-only, Arm SME-only, or cooperative execution for each LLM operator shape. It partitions matrix tiles, overlaps matrix and vector stages in attention, and preserves packed layouts to avoid repeated conversion. Tests span Llama-3.2-3B, Qwen3-4B, and Qwen3-30BA3B on phone, PC, and server platforms, but benefits will depend on shared-memory bandwidth and the availability of SME-capable CPUs. SMEPilot利用roofline分析,按LLM算子形状选择纯CPU、纯Arm SME或协同执行。系统对矩阵tile进行划分,在注意力中重叠矩阵与向量阶段,并保留打包布局以避免重复转换。测试覆盖手机、PC和服务器上的Llama-3.2-3B、Qwen3-4B及Qwen3-30BA3B,但收益取决于共享内存带宽与支持SME的CPU可用性。

Quantum Hardware & Systems 量子硬件与系统

Quantum Graph States with O(1) Local Feedforward 采用O(1)局部前馈的量子图态生成

X. Zheng, C.-T. Lam, L. Chen, et al.

arXiv:2606.16375 · 2026-06-15T08:10:59Z

The protocol replaces global quantum-network routing decisions with local measurements and amortized O(1) classical feedforward to keep control delay below qubit coherence times. Hybrid simulations map the design to a dual-species trapped-ion platform and identify readout fidelity as the dominant bottleneck, with erasure conversion, spatial multiplexing, and branch independence proposed as mitigations. The evidence is simulated, so experimental latency and accumulated error remain unverified. 该协议以局部测量和摊销O(1)的经典前馈取代全局量子网络路由决策,使控制延迟保持在量子比特相干时间以内。混合仿真将其映射到双物种离子阱平台,并指出读出保真度是主要瓶颈,提出用擦除转换、空间复用和分支独立性缓解。现有证据来自仿真,实验时延与累积误差仍未验证。

EDA & Design Methods EDA与设计方法

ComPart Uses Community Structure After Hypergraph Coarsening ComPart在超图粗化后利用社区结构

Y. Zhu, Z. Guo, Y. Wu, et al.

arXiv:2606.18131 · 2026-06-16T16:27:52Z

ComPart moves community detection beyond the coarsening phase of hypergraph partitioning and uses structures found during initial partitioning and uncoarsening to guide refinement. The framework is designed to escape local optima and accept multiple present or future community-detection methods, making it relevant to heterogeneous MPSoC mapping and multi-FPGA prototyping. The candidate abstract excerpt gives no cut-quality or runtime numbers, so the practical gain over established partitioners still needs inspection. ComPart把社区发现从超图划分的粗化阶段延伸到初始划分和反粗化过程,并利用这些结构指导细化。该框架旨在跳出局部最优,同时兼容现有及未来的多种社区发现方法,适用于异构MPSoC映射和多FPGA原型验证。候选摘要没有给出割质量或运行时间数据,因此相对成熟划分器的实际收益仍需进一步检查。