semi·news
Headlines要闻 / Research研究 / /
Research digest · Monday, August 3, 2026 研究摘要 · 2026年8月3日 星期一

Hardware-aware design confronts data movement 硬件感知设计直面数据移动瓶颈

This week's selection shifts attention from nominal operations to the interfaces that make them useful: unstable clocks, KV-cache noise, chiplet PHYs, and HLS verification. Several papers offer measured or simulated gains, but their practical value hinges on device variation and system-level integration. 本周论文把关注点从名义运算能力转向决定其实用性的接口:不稳定时钟、KV缓存噪声、chiplet PHY以及HLS验证。多篇工作给出测量或仿真增益,但其实际价值仍取决于器件变异和系统级集成。

Look-back window: 7 days · 6 paper(s) 回溯窗口: 7天 · 6篇

Devices & Process 器件与工艺

Low-Power PLL-Based Clock Stabilization for Flexible IGZO Mixed-Signal Systems 面向柔性IGZO混合信号系统的低功耗PLL时钟稳定技术

Paula Carolina Lozano Duarte, Georgios Zervakis, Mehdi Tahoori

arXiv:2607.29357 · 2026-07-31T12:46:00Z

The authors propose a PLL tailored to n-type-only amorphous IGZO thin-film transistors, using low-bandwidth feedback to stabilize a ring-oscillator clock under large process, voltage, and temperature variation. It targets flexible mixed-signal electronics, where conventional clock sources can consume up to 90% of system power. This is a temporal stabilizer rather than a precision synthesizer, so its value depends on tolerance for residual jitter and drift. 作者提出面向仅含n型非晶IGZO薄膜晶体管的PLL,以低带宽反馈在较大的工艺、电压和温度变化下稳定环形振荡器时钟。其目标是柔性混合信号电子系统,传统时钟源在此类系统中可消耗高达90%的系统功耗。这更像时间稳定器而非高精度频率合成器,因此价值取决于应用对残余抖动和漂移的容忍度。

EDA & Design Automation EDA与设计自动化

RTLCurator: Label-Efficient Data Curation for RTL Generation RTLCurator:面向RTL生成的高标签效率数据筛选

Siyang Cai, Cangyuan Li, Wenjing Chang, et al.

arXiv:2607.29283 · 2026-07-31T10:55:01Z

RTLCurator learns a behavior-aware compatibility prior from a small set of validated examples to select training data for RTL-generating models. Only 24.4% and 53.5% of pairs in two common synthetic RTL datasets pass generated functional tests, the paper says. By balancing alignment, diversity, and complexity, it seeks to retain data beyond simple correct modules, but remains dependent on representative validation labels. RTLCurator从少量已验证样本中学习行为感知的兼容性先验,用于筛选RTL生成模型的训练数据。论文指出,两个常用合成RTL数据集分别仅有24.4%和53.5%的样本对能通过自动生成的功能测试。该方法平衡对齐度、多样性和复杂度,试图保留超越简单正确模块的数据,但仍依赖具有代表性的验证标签。

ContractHIL-HLS: Contract-Aligned Hardware-in-the-Loop Workflow for HLS ContractHIL-HLS:契约对齐、硬件在环的HLS工作流

Jingbo Zhang, Haoxiang Sun, Wenbo Wang, Wenbo Zhang

arXiv:2607.25283 · 2026-07-28T04:41:13Z

ContractHIL-HLS turns natural-language design requests into explicit interfaces, constraints, checks, and rollback rules, then feeds HLS, Vivado, PYNQ, power, and failure evidence back into generation. On 94 HLS-Eval tasks, its structured contract lifted estimated single-sample testbench pass rate from 64.0% to 70.2%. The feedback loop makes LLM-assisted HLS more verifiable, though board-level closure remains harder than benchmark tasks. ContractHIL-HLS将自然语言设计需求转化为明确的接口、约束、检查和回滚规则,再把HLS、Vivado、PYNQ、功耗与失败证据反馈到生成过程。在94个HLS-Eval任务上,结构化契约将估计的单样本测试平台通过率从64.0%提高至70.2%。该反馈循环使LLM辅助HLS更可验证,但板级收敛仍比基准任务更难。

AI Accelerators & Compute-in-Memory AI加速器与存算一体

Selective KV-Cache Protection for Analog Compute-in-Memory LLM Inference 用于模拟存算一体LLM推理的选择性KV缓存保护

Yuannuo Feng, Wenyong Zhou, Yuang Ma, et al.

arXiv:2607.29076 · 2026-07-31T06:56:20Z

The paper identifies initial and recent tokens as especially vulnerable to analog noise when KV-cache updates run on compute-in-memory arrays. Its scheme keeps those tokens on a higher-precision digital path while sending most cache data through analog hardware. Across nine LLMs, average perplexity falls from 33.91 to 11.95 under analog noise, but the trade-off depends on programming overhead and the noise model. 论文发现,当KV缓存更新在存算一体阵列上运行时,初始token和最近token对模拟噪声尤其敏感。其方案将这些token保留在高精度数字路径,同时把大部分缓存数据交由模拟硬件处理。作者在9个LLM上报告,模拟噪声下平均困惑度从33.91降至11.95,但权衡取决于编程开销和噪声模型。

Heterogeneous Accelerator for Multitask RF Signal Recognition 面向多任务射频信号识别的异构加速器

Zhifan Song, Haralampos-G. Stratigopoulos, Hassan Aboushady

arXiv:2607.24669 · 2026-07-27T17:10:06Z

This accelerator combines an attention-enhanced CNN, learnable streaming decimation, fused convolution-pooling, DMA streaming, and CPU co-execution for RF tasks. It reports 98 microseconds per frame, at least 99% average modulation-recognition accuracy above 4 dB SNR, and 99.5% for GNSS-jamming classification. Dataset accuracy is encouraging, but does not fully characterize field robustness. 该加速器将带注意力的CNN、可学习流式抽取、融合卷积—池化、DMA流和CPU协同执行用于射频任务。作者报告每帧98微秒延迟,在4 dB以上SNR的调制识别中平均准确率至少99%,GNSS干扰分类为99.5%。数据集准确率令人鼓舞,但尚不能完全说明现场鲁棒性。

Circuits & Architecture 电路与体系结构

DICE: End-to-End PHY Modeling for Chiplet Simulation DICE:用于chiplet仿真的端到端PHY建模

Rashid Aligholipour, Stefanos Kaxiras, Yuan Yao

arXiv:2607.24221 · 2026-07-27T09:53:10Z

DICE adds runtime-dependent PHY behavior to chiplet simulation, accounting for channel noise, crosstalk, jitter, iterative FEC decoding, retransmissions, and application traffic. The authors argue that fixed-latency links distort packet behavior as short-reach interconnects approach signal-integrity limits. The detail can improve architectural decisions, but credible early-stage device and channel parameters are needed. DICE将运行时相关的PHY行为引入chiplet仿真,纳入信道噪声、串扰、抖动、迭代FEC解码、重传和应用流量。作者认为,当短距互连接近信号完整性极限时,固定延迟链路会扭曲数据包行为。细节建模可改善架构决策,但需要可信的早期器件和信道参数。

Newsletter 邮件订阅

Daily semiconductor briefing. 每日半导体简报。