semi·news
Headlines要闻 / Research研究 / /
Research digest · Thursday, July 9, 2026 研究摘要 · 2026年7月9日 星期四

Inference Efficiency Moves Into Memory and Tools 推理效率转向存储与工具链

This week's research queue is less about bigger models and more about the hardware-adjacent constraints around attention, memory hierarchy, FPGA automation, verification, RF circuits, and reconfigurable devices. Several papers trade brute-force scaling for structure, sparsity, or tool automation. 本周研究队列的重点不在更大模型,而在注意力、存储层级、FPGA自动化、验证、射频电路和可重构器件等贴近硬件的约束。多篇论文用结构化、稀疏化或工具自动化来替代单纯的暴力扩展。

Look-back window: 7 days · 10 paper(s) 回溯窗口: 7天 · 10篇

Devices & Process 器件与工艺

Neural-Operator Evolutionary Search for Nanophotonic Inverse Design 用于纳米光子逆向设计的神经算子进化搜索

Xiangming Huang, Guannan Zhang, Lu Lu, Raphaël Pestourie

arXiv:2607.07682 · 2026-07-08T17:41:56Z

The paper introduces NOTES, a neural-operator-enabled topology-informed evolutionary strategy for PDE-constrained inverse design. In a nanophotonic beam-deflector task governed by Maxwell's equations, it reduces the design dimension from 256 to 25 and reports over 95% efficiency. The hardware relevance is the search method: compact latent optimization could make photonic component design less dependent on expensive full-field sweeps. 论文提出NOTES,一种面向PDE约束逆向设计的神经算子与拓扑先验结合的进化策略。在由Maxwell方程约束的纳米光子束偏转器任务中,该方法把设计维度从256降到25,并报告超过95%的效率。其硬件意义在于搜索方法本身:紧凑潜空间优化可能降低光子器件设计对昂贵全场扫描的依赖。

Reconfigurable Non-Hermitian Metasurface With Multiple Singularities 具备多重奇点的可重构非厄米超表面

Xintong Shi, Yanjie Wu, Rui Zhou, Zuxing Lu, Tengyu Li, Tingting Liu, Hai Lin, Qiegen Liu, Shuyuan Xiao

arXiv:2607.07402 · 2026-07-08T13:35:30Z

This work uses a mirror-coupled metasurface to create a higher-dimensional effective parameter space without adding more physical resonators. A PIN-diode reconfigurable platform demonstrates coexistence and manipulation of an exceptional point and multiple reflection zeros. The result matters for tunable microwave devices because it links singularity control to functions such as coordinated absorption rather than treating the physics as a static curiosity. 这项工作通过镜面耦合超表面,在不增加物理谐振器数量的情况下构造更高维的等效参数空间。作者在集成PIN二极管的可重构平台上展示了exceptional point与多个反射零点的共存和调控。它对可调微波器件有意义,因为论文把奇点控制与协同吸收等功能联系起来,而不只是展示静态物理现象。

Circuits & Architecture 电路与架构

Six-Pole Dual-Band Microstrip Filter for WiMAX 面向WiMAX的六极点双频微带滤波器

Halimat Olamide Yusuf, Augustine O. Nwajana

arXiv:2607.07661 · 2026-07-08T17:17:17Z

The authors transform a third-order single-band filter into a sixth-order dual-band bandpass filter using a folded-arms square open-loop resonator microstrip structure. The design targets passbands around 2.2 GHz and 2.4 GHz, centered near 2.3 GHz with 7% fractional bandwidth, on Rogers RT/Duroid 6010LM. It is a practical RF-circuit contribution where the value is compact dual-band selectivity rather than a new semiconductor device. 作者使用folded-arms square open-loop resonator微带结构,将三阶单频带滤波器转换为六阶双频带带通滤波器。该设计在Rogers RT/Duroid 6010LM基板上实现约2.2 GHz和2.4 GHz两个通带,中心约2.3 GHz,分数带宽7%。这是一项偏实用的射频电路工作,其价值在于紧凑双频选择性,而不是新的半导体器件。

RISC-V Hardware-Software Co-Design for Embedded Blockchain Infrastructure 面向嵌入式区块链基础设施的RISC-V软硬件协同设计

Qinglin Yang, Yuan Liu, Yaoyao Zhang, Boya Wang, Zongjian You, Chunming Rong, Zhihong Tian

arXiv:2607.07625 · 2026-07-08T16:41:05Z

This survey frames embedded Blockchain Infrastructure Management as a RISC-V-based hardware-software co-design paradigm. The thesis is that open, low-level and verifiable computing substrates can reduce dependence on third-party blockchain-as-a-service operators. For chip architects, the useful part is the mapping from trust requirements to extensible RISC-V hardware, firmware and management layers. 这篇综述把embedded Blockchain Infrastructure Management定义为基于RISC-V的软硬件协同设计范式。其核心观点是,开放、底层且可验证的计算底座可以降低对第三方区块链即服务运营商的依赖。对芯片架构师而言,较有价值的是论文把信任需求映射到可扩展RISC-V硬件、固件和管理层的方式。

AI Accelerators & Compute-in-Memory AI加速器与存算一体

ATLAS Automates HLS for Deep-Learning FPGA Hardblocks ATLAS自动化面向深度学习FPGA硬块的HLS流程

Ruthwik Reddy Sunketa, Aman Arora

arXiv:2607.07643 · 2026-07-08T17:00:41Z

ATLAS is a flow that maps high-level deep-learning model descriptions to FPGA implementations using custom in-fabric DL hardblocks. The authors target a real usability gap: conventional HLS raises abstraction but still does not natively generate code for specialized hardblocks without manual wrappers and RTL libraries. If the flow generalizes, it could make DL-optimized FPGA fabrics easier to program beyond expert RTL teams. ATLAS是一套把高层深度学习模型描述映射到带有定制片上DL硬块的FPGA实现流程。作者瞄准的是一个真实易用性缺口:传统HLS提高了抽象层级,但仍无法原生面向专用硬块生成代码,通常需要人工包装和RTL库。如果该流程具备通用性,将降低DL优化FPGA织构的编程门槛,使其不再只依赖资深RTL团队。

TF-Engram Adds SSD-Backed External Memory to LLMs TF-Engram为LLM引入SSD支撑的外部记忆

Yutang Ma, Kecheng Huang, Xikun Jiang, Zili Shao

arXiv:2607.07388 · 2026-07-08T13:19:52Z

TF-Engram builds phrase-specific semantic memory offline, stores it across a GPU-DRAM-SSD hierarchy, and uses early-exit guided prefetching to hide external-memory latency during decoding. On Qwen3-0.6B, it reports an average downstream-score improvement from 57.6 to 59.4 while reducing GPU memory pressure. The paper is relevant to accelerator systems because it treats storage hierarchy and predictive data movement as part of the model-serving design. TF-Engram离线构建短语级语义记忆,将大规模记忆表分布在GPU、DRAM和SSD层级中,并用early-exit引导的预取来隐藏解码时的外部存储延迟。在Qwen3-0.6B上,论文报告平均下游分数从57.6提升到59.4,同时降低GPU显存压力。它对加速器系统有意义,因为论文把存储层级和预测式数据搬移视为模型服务设计的一部分。

AI Research & Inference Systems AI研究与推理系统

Analysis-Driven Linearization for Long-Context Transformers 面向长上下文Transformer的分析驱动线性化

Anna Kuzina, Paul N. Whatmough, Babak Ehteshami Bejnordi

arXiv:2607.07706 · 2026-07-08T17:59:09Z

The paper studies transformer linearization in a frozen-backbone regime and argues that softmax attention relies on key-dependent rank-1 orthogonal projections. It adds sink tokens, short convolutions and fixed-budget cache routing to reduce approximation error, then scales the approach across LLaMA and Qwen models up to 32B parameters. The hardware angle is clear: better post-hoc linearization can reduce long-context inference cost without retraining the whole model. 论文在冻结主干模型的设定下研究Transformer线性化,并指出softmax attention依赖key相关的rank-1正交投影。作者加入sink token、短卷积和固定预算cache routing来降低近似误差,并把方法扩展到最高32B参数的LLaMA和Qwen模型。其硬件意义很明确:更好的后训练线性化可以在不重训整个模型的情况下,降低长上下文推理成本。

DeLS-Spec Decouples Long and Short Contexts for Speculative Drafting DeLS-Spec解耦长短上下文以改进投机解码

Hong-Kai Zheng, Piji Li

arXiv:2607.07409 · 2026-07-08T13:41:52Z

DeLS-Spec improves speculative decoding by combining a fixed DFlash long-context expert with a lightweight local head trained independently for short-context causality. That avoids retraining the full draft model while addressing the lack of intra-block causal conditioning in block-parallel drafters. The method matters for inference systems because modular draft heads could be swapped and trained cheaply across serving stacks. DeLS-Spec通过把固定的DFlash长上下文专家与独立训练的轻量本地头结合,改进投机解码中的短上下文因果建模。这样既避免重训完整draft model,又缓解块并行草稿器缺少块内因果条件的问题。该方法对推理系统有意义,因为模块化draft head可能以较低成本在不同服务栈中替换和训练。

Quantum 量子

Quantum CNN With Rough-Path Signature Kernels 结合粗路径签名核的量子卷积神经网络

Leonardo Nogueira Falabella, Vasily Sazonov

arXiv:2607.07634 · 2026-07-08T16:50:46Z

The paper proposes a hybrid quantum-classical time-series classifier that uses path-signature kernels before a quantum convolutional neural network. The authors explore classical and variational quantum linear solvers inside the feature layers to handle time-reparameterization invariance. It is early-stage algorithmic work, but it gives a concrete workload for evaluating whether quantum kernels add value beyond classical preprocessing. 论文提出一种混合量子-经典时间序列分类器,在量子卷积神经网络之前使用路径签名核。作者在特征层中探索经典和变分量子线性求解器,以处理时间重参数化不变性问题。这仍是早期算法研究,但为评估量子核是否能超越经典预处理提供了具体工作负载。

EDA & Verification EDA与验证

LLM-Assisted SystemVerilog Assertion Generation LLM辅助生成SystemVerilog断言

Bhabesh Mali, Chandan Karfa

arXiv:2607.07444 · 2026-07-08T14:14:14Z

This paper reviews the use of LLMs for generating SystemVerilog Assertions from design specifications and frames the key problem as quality-aware methodology rather than prompt novelty. Assertion-based verification is labor-intensive, so even partial automation could change verification productivity if coverage and correctness are measurable. The authors' useful contribution is a challenge map for making LLM-generated assertions systematic enough for design-verification flows. 这篇论文综述了用LLM从设计规范生成SystemVerilog Assertions的方法,并把核心问题定位为质量可控的方法论,而不是单纯提示词创新。基于断言的验证高度依赖人工,因此如果覆盖率和正确性能够被度量,即使部分自动化也可能改变验证生产率。作者较有价值的贡献,是为LLM生成断言如何系统化进入设计验证流程绘制了问题地图。