semi·news
Headlines要闻 / Research研究 / /
Research digest · Tuesday, July 28, 2026 研究摘要 · 2026年7月28日 星期二

Making AI hardware more efficient and verifiable 让AI硬件更高效、更可验证

This week's papers pair inference and memory-efficiency techniques with tools for building and validating hardware. Their common goal is to turn algorithmic flexibility into predictable behavior on real architectures. 本周论文一方面探索推理与存储效率,另一方面改进硬件构建和验证工具。共同目标是将算法灵活性转化为真实架构上可预测的行为。

Look-back window: 7 days · 7 paper(s) 回溯窗口: 7天 · 7篇

AI Accelerators & Compute-in-Memory AI加速器与存算一体

Task-conditional compute skipping for multi-task inference accelerators 面向多任务推理加速器的任务条件计算跳过

Afzal Ahmad, Gaoyu Mao, et al.

arXiv:2607.22038 · 2026-07-24

Sparse by Command uses the task command available before inference to predict hardware-aligned tile masks, letting a multi-task accelerator skip irrelevant work. It co-designs training, an ISA carrying tile bitmasks, and a scheduler, but real gains will depend on task mix and gate accuracy. Sparse by Command利用推理前已知的任务指令预测与硬件对齐的Tile掩码,使多任务加速器跳过无关计算。它协同设计训练、携带Tile位掩码的ISA和调度器,但实际收益仍取决于任务组合与门控准确性。

RED-PIM reduces Transformer attention data movement RED-PIM降低Transformer注意力的数据搬运

Zahra Yousefijamarani, Alaa Alameldeen

arXiv:2607.21731 · 2026-07-23

RED-PIM reorganizes Transformer attention for processing-in-memory so inter-bank movement falls from O(N²) to O(N), while intermediate matrices shrink from N×N to d×d. It targets a scaling weakness in prior PIM attention designs, although bank capacity and mapping overhead will determine practical benefits. RED-PIM为存内计算重组Transformer注意力计算,使跨Bank数据移动从O(N²)降至O(N),中间矩阵从N×N缩小到d×d。它针对先前PIM注意力设计的扩展性弱点,但Bank容量和映射开销将决定实际收益。

AI Inference Systems AI推理系统

Unified static and dynamic pruning for GPU LLM inference 用于GPU LLM推理的静态—动态统一剪枝

Jinhyeok Kim, Yejoon Lee, Jaeyoung Do

arXiv:2607.21985 · 2026-07-24

SPDP combines permanently pruned weights with input-adaptive activation skipping for LLM inference on GPUs. Its tiled bitmap format and separate CUDA-core and Tensor-Core kernels confront the metadata and irregular-access costs that have limited earlier sparse designs. SPDP将永久权重剪枝与输入自适应激活跳过结合,用于GPU上的LLM推理。其分块位图格式和独立的CUDA Core、Tensor Core内核,直接应对限制早期稀疏设计的元数据和不规则访问成本。

Circuits & Reliability 电路与可靠性

Warp divergence from Pascal through Blackwell 从Pascal到Blackwell的Warp分歧特性

Alpin Dale

arXiv:2607.23402 · 2026-07-26

This microbenchmark study finds that divergent GPU paths serialize roughly linearly with path count from Pascal through Blackwell, while warp efficiency falls as 32/k. It separates that stable programmer-visible cost model from changes in compiler reconvergence machinery, though the evidence comes from controlled kernels rather than full applications. 这项微基准研究发现,从Pascal到Blackwell,GPU分歧路径的串行化时间大致随路径数线性增长,Warp效率则按32/k下降。它将稳定的程序员可见成本模型与编译器汇合机制变化区分开来,但证据来自受控内核而非完整应用。

Distributed delay-based BIST for flexible mixed-signal circuits 用于柔性混合信号电路的分布式延迟BIST

Paula Carolina Lozano Duarte, Sule Ozev, et al.

arXiv:2607.19310 · 2026-07-21

The authors embed lightweight digital BIST into ring-oscillator stages for IGZO-TFT flexible electronics. They report 93% individual-defect coverage, 88% multiple-defect coverage, and 3% power overhead, addressing test constraints where external equipment and pins are scarce. 作者将轻量级数字BIST嵌入IGZO-TFT柔性电子的环形振荡器级中。他们报告单个缺陷覆盖率93%、多个缺陷覆盖率88%、功耗开销3%,面向外部测试设备和引脚稀缺的测试约束。

EDA & Design Verification EDA与设计验证

CircuitWeave aligns schematics and specifications for executable RTL CircuitWeave对齐原理图与规格以生成可执行RTL

Jiahao Feng, Haiyan Qin, et al.

arXiv:2607.23523 · 2026-07-26

CircuitWeave separates a schematic-derived topology contract from a text-derived behavior contract before generating RTL. Its 5,000-package dataset and explicit fusion aim to surface missing or conflicting evidence, a key failure mode in multimodal RTL generation that merits broader testing on human specifications. CircuitWeave在生成RTL前,先分离由原理图得到的拓扑契约和由文本得到的行为契约。其包含5000个软件包的数据集及显式融合机制旨在暴露缺失或冲突证据,这一多模态RTL生成的关键失败模式仍需在人工规格上更广泛测试。

AlphaRoute uses LLMs to tune multi-objective global routing AlphaRoute使用LLM调优多目标全局布线

Kabir Murjani, Mishri Bhavsar, et al.

arXiv:2607.19768 · 2026-07-22

AlphaRoute combines deterministic routing tools with LLM policy optimization to adapt congestion penalties during rip-up and reroute. On selected ISPD 2025 benchmarks it reports a 98.6% overflow reduction on MEMPOOL, though results need replication across designs and clear accounting of inference cost and determinism. AlphaRoute将确定性布线工具与LLM策略优化结合,在拆线重布过程中自适应调整拥塞惩罚。在选定的ISPD 2025基准上,它报告MEMPOOL溢出减少98.6%,但仍需在更多设计上复现,并清楚核算推理成本与确定性。

Newsletter 邮件订阅

Daily semiconductor briefing. 每日半导体简报。