semi·news
Headlines要闻 / Research研究 / /
Research digest · Tuesday, June 23, 2026 研究摘要 · 2026年6月23日 星期二

Hardware Co-Design Moves Toward Measured Constraints 软硬件协同设计转向实测约束

This week's work couples algorithm choices to accelerator structure, grounds device and process claims in measurements, and closes verification loops around AI-generated designs. The strongest results expose where projections still depend on memory size, process control, or unreported implementation details. 本周研究把算法选择与加速器结构更紧密地结合,以实测结果支撑器件和工艺主张,并围绕AI生成设计补上验证闭环。最有价值的结果也明确暴露了其对存储容量、工艺控制或尚未披露实现细节的依赖。

Look-back window: 7 days · 10 paper(s) 回溯窗口: 7天 · 10篇

Devices & Process 器件与工艺

Bendable Self-Powered Silicon Photodetector Reaches 178.8 dB Dynamic Range 可弯曲自供电硅光探测器实现178.8 dB动态范围

L. Qiu, Z. Tian, Z. Zhao, et al.

Laser & Photonics Reviews · 2026-06-19

A CuO film conformally deposited on a freestanding pyramid-textured Si layer produces a bendable, self-powered photodetector with a measured 178.8 dB linear dynamic range and 0.69/0.92 microsecond rise/decay times at zero bias. A visible-light positioning demonstration reached 1.4 cm RMS error, while the communication link supported 620 kbps. These are device-level measurements; long-cycle bending durability and manufacturing-scale uniformity remain open questions. 研究在自支撑金字塔纹理Si薄层上共形沉积CuO,制成可弯曲、自供电光探测器;器件在零偏压下实测线性动态范围178.8 dB,上升/衰减时间为0.69/0.92微秒。可见光定位演示实现1.4 cm均方根误差,通信链路支持620 kbps。结果来自器件级实测,但长期弯折耐久性与规模制造均匀性仍待验证。

Pulsed RF Bias Tunes Cryogenic Silicon-Grating Etch Profiles 脉冲RF偏压调控低温硅光栅刻蚀轮廓

Z. Shi, M. Lu, N. Tiwale, et al.

Journal of Vacuum Science & Technology A · 2026-06-16

Experiments on 500 nm-wide, 1 micrometer-pitch silicon gratings etched 0.8-1.1 micrometers deep show that 0.5-2 kHz pulsed RF bias and 25%-75% duty cycles strongly affect etch rate, linewidth, and sidewall angle. Moderate 50%-75% duty cycles near 1 kHz best balance anisotropy and throughput by alternating ion-driven etching with passivation recovery. The result adds a practical control knob for cryogenic SF6/O2 etching, but it is demonstrated on one grating geometry rather than a full process window. 在宽500 nm、节距1微米、刻蚀深度0.8-1.1微米的硅光栅上,实验表明0.5-2 kHz脉冲RF偏压和25%-75%占空比会显著影响刻蚀速率、线宽与侧壁角。约1 kHz、50%-75%占空比通过交替进行离子驱动刻蚀与钝化恢复,在各向异性和产能之间取得较好平衡。该结果为低温SF6/O2刻蚀增加了实用调节手段,但目前只在一种光栅几何结构上验证,尚非完整工艺窗口。

Circuits & Architecture 电路与架构

COMPOSE Fuses CGRA Operations Around Static Timing Slack COMPOSE利用静态时序余量融合CGRA运算

R. Juneja, V. Ranjan, R. Harish, et al.

arXiv:2606.21454 · 2026-06-19T14:12:19Z

COMPOSE uses static timing information to form composable CGRA processing elements at compile time, spatially fusing operations across loop iterations to attack recurrence-bound loops. The approach aims to reduce serialized dependency handling, intermediate-result registration, and local data movement that fixed processing elements impose. The preprint's candidate abstract is truncated before quantitative results, so throughput and energy claims need confirmation from the full evaluation. COMPOSE利用静态时序信息在编译期组合CGRA处理单元,并跨循环迭代进行空间运算融合,以加速受递归依赖限制的循环。该方法旨在减少固定处理单元带来的依赖串行化、中间结果寄存以及局部数据搬运。候选摘要在定量结果之前被截断,因此吞吐率与能效主张仍需结合完整评测确认。

Whisper Dot-Product Offload Cuts Projected Energy-Delay on a CGLA Whisper点积卸载降低CGLA的预计能量延迟

T. Ando, Y. Eto, A. Takeuchi, et al.

arXiv:2606.19913 · 2026-06-18T08:07:59Z

Profiling finds dot products consume 90.6% of FP16 and 87.1% of Q8_0 execution time for Whisper-tiny.en on Cortex-A72, motivating offload to the programmable IMAX coarse-grained linear array. An FPGA prototype plus a 28 nm, 840 MHz ASIC projection reports 11.58 J PDP for Q8_0, 2.35x below Jetson AGX Orin and 10.48x below RTX 4090 under a TDP-based comparison. The advantage narrows for larger Whisper models as 32 KB local-memory coverage falls, and the headline ASIC figures remain projected rather than measured silicon. 性能分析显示,在Cortex-A72运行Whisper-tiny.en时,点积占FP16执行时间的90.6%、占Q8_0执行时间的87.1%,因此研究将该内核卸载到可编程IMAX粗粒度线性阵列。基于FPGA原型和28 nm、840 MHz ASIC投影,Q8_0的功率延迟积为11.58 J;按TDP比较,分别比Jetson AGX Orin和RTX 4090低2.35倍与10.48倍。随着32 KB本地存储对更大Whisper模型的覆盖率下降,优势随之收窄,而且主要ASIC指标仍是投影而非实测硅结果。

AI Accelerators & Compute-in-Memory AI加速器与存算一体

A3C3 Jointly Searches AI Models and Accelerator Implementations A3C3联合搜索AI模型与加速器实现

S. Yildirim, Y. Huang, D. Chen

arXiv:2606.20869 · 2026-06-18T19:01:45Z

A3C3 parameterizes neural-network and accelerator design spaces together, then co-searches model-hardware pairs across accuracy, latency, throughput, energy, and utilization constraints. This directly addresses the inefficiency of designing an accurate model first and adapting it to hardware afterward. The work is a handbook chapter rather than a new silicon result, so it is most useful as a methodology and design-space reference. A3C3同时参数化神经网络与加速器设计空间,并围绕精度、延迟、吞吐率、能耗和利用率约束联合搜索模型—硬件组合。它直接针对先设计高精度模型、再事后适配硬件所造成的效率损失。该成果是手册章节而非新的流片结果,因此更适合作为方法论与设计空间参考。

A Reduced RISC-V Core Targets Programmable Tsetlin-Machine Inference 精简RISC-V核心面向可编程Tsetlin Machine推理

C. Gupta, S. Bhatia, S. Priyadarshi, et al.

arXiv:2606.19964 · 2026-06-18T09:05:41Z

The design profiles Tsetlin-machine workloads and removes unneeded RV32IM instructions and datapath/control complexity to build a domain-specific but programmable RISC-V inference processor. It avoids the tightly coupled host interfaces and microcode dependence common in dedicated Tsetlin accelerators, while retaining a software-visible instruction set. The candidate abstract provides no final speed or energy deltas, so the architecture should be judged as an early design study rather than a demonstrated efficiency lead. 该设计先分析Tsetlin Machine工作负载,再删减不需要的RV32IM指令以及数据通路和控制复杂度,构建领域专用但仍可编程的RISC-V推理处理器。它避免了专用Tsetlin加速器常见的紧耦合主机接口和微代码依赖,同时保留软件可见指令集。候选摘要未给出最终速度或能耗差异,因此现阶段更应视为早期架构研究,而非已证明的能效领先方案。

Hardware-Relevant AI Research 硬件相关AI研究

A Systems Pipeline for Machine Learning on Microcontroller-Class Devices 面向微控制器级设备的机器学习系统流程

M. Darvishi

arXiv:2606.18122 · 2026-06-16T16:22:24Z

This systems-oriented paper connects sampling, buffering, feature extraction, class-imbalance validation, quantization, scheduling, and streaming deployment for microcontroller-class inference. It works through inertial recognition using two-second three-axis windows and keyword spotting using MFCCs plus a compact 1D CNN, making hardware constraints explicit across the full pipeline. It is a synthesis and practical guide rather than a new accelerator benchmark, but it is useful for avoiding optimizations that shift cost between preprocessing, runtime, and memory. 这篇系统导向论文串联了采样、缓冲、特征提取、类别不平衡验证、量化、调度与流式部署,面向微控制器级推理给出完整流程。论文以两秒三轴窗口的惯性识别,以及采用MFCC和紧凑型一维CNN的关键词识别为例,把全流程硬件约束明确化。它属于综合性实践指南而非新的加速器基准,但有助于避免只在预处理、运行时和存储之间转移成本的局部优化。

Quantum & Photonics 量子与光子学

An Integrated Bright Source of Polarization-Entangled Photons on Lithium Niobate 铌酸锂芯片集成高亮度偏振纠缠光子源

H. Kwon, C. Kim, H. Kim, et al.

Optica · 2026-06-16

The paper reports an integrated bright source of polarization-entangled photons implemented on lithium-niobate photonic chips. Bringing the source on chip can reduce the size and stability overhead of quantum-photonic systems and improve integration with routing and detection. The supplied candidate metadata contains no brightness, fidelity, or loss figures, so quantitative comparison with other entangled-photon sources requires the full paper. 论文报道了在铌酸锂光子芯片上实现的集成高亮度偏振纠缠光子源。把光源集成到芯片上可降低量子光子系统的体积与稳定性负担,并有利于同路由和探测模块集成。候选资料未提供亮度、保真度或损耗数据,因此与其他纠缠光子源的定量比较仍需阅读原文。

EDA & Verification EDA与验证

Formal Verification Closes the Loop on LLM-Generated RTL Assertions 形式验证为LLM生成的RTL断言建立闭环

R. Krishnamurthy, D. Chitnis, T. Prodromakis

arXiv:2606.21451 · 2026-06-19T14:10:49Z

An open-source framework applies mutation-guided refinement to reject vacuous or weak LLM-generated RTL assertions, selects SMT solvers from RTL structure, and synthesizes counterexample-grounded failure narratives. Across multiple RTL designs, the authors report stronger assertion confidence and less runtime variability than fixed-solver flows. The contribution is the quality-control loop after generation, although the preprint abstract does not provide benchmark counts or absolute proof times. 这一开源框架利用突变引导优化剔除空洞或薄弱的LLM生成RTL断言,根据RTL结构选择SMT求解器,并生成以反例轨迹为依据的失败解释。作者在多种RTL设计上报告了更高的断言可信度和低于固定求解器流程的运行时间波动。其核心贡献是补上生成后的质量控制闭环,但预印本摘要未提供基准数量或绝对证明时间。

AI-Generated Pixelated Microwave Filter Is Validated by Field Measurements AI生成的像素化微波滤波器通过场测量验证

H. Zhou, R. Bannister, C. Pierce, et al.

arXiv:2606.18402 · 2026-06-16T18:51:02Z

A CNN-plus-genetic-algorithm flow synthesizes a pixelated low-pass microwave filter, then validates it with both S-parameters and electro-optical electric-field imaging. The fabricated design delivers a 7 GHz passband and more than 20 dB suppression above 9.5 GHz, with measured field patterns resembling coupled lines or stubs that make the generated topology more interpretable. The result is an experimental EDA demonstration, though broader value depends on extending the method beyond one filter class and frequency target. 研究以CNN结合遗传算法合成像素化低通微波滤波器,并通过S参数和电光电场成像进行双重验证。制成器件实现7 GHz通带,在9.5 GHz以上抑制度超过20 dB;实测场分布呈现类似耦合传输线或短截线的结构,使AI生成拓扑更易解释。这是一次实验性EDA验证,但更广泛价值仍取决于能否扩展到更多滤波器类型和频率目标。