semi·news
Headlines要闻 / Research研究 / /
Research digest · Tuesday, June 16, 2026 研究摘要 · 2026年6月16日 星期二

Memory Bottlenecks Shape Architecture Research 存储瓶颈塑造架构研究

This week's papers cluster around memory pressure, inference efficiency, and using LLMs inside hardware-design workflows. The strongest work is practical: measured DRAM behavior, memory-side execution, compression near the Shannon limit, and formal or power tools that fit real design loops. 本周论文集中在存储压力、推理效率,以及把LLM引入硬件设计流程。最值得关注的工作偏实践:实测DRAM行为、存储侧执行、接近Shannon极限的压缩,以及能嵌入真实设计循环的形式验证和功耗工具。

Look-back window: 7 days · 10 paper(s) 回溯窗口: 7天 · 10篇

Devices & Process 器件与工艺

In-DRAM Signatures from Simultaneous Multi-Row Activation 利用同时多行激活生成DRAM内签名

U. Baser, I. E. Yuksel, F. N. Bostanci, et al.

arXiv:2606.15470 · 2026-06-13T20:57:07Z

The paper demonstrates a DRAM PUF based on simultaneous multiple-row activation using 112 off-the-shelf DDR4 chips. The reported intra-Jaccard indices reach roughly 89% to 95% while inter-device indices stay near 2% to 4%, indicating repeatability within a chip and uniqueness across chips. The result is useful because it turns a memory disturbance mechanism into an authentication primitive, though deployment would still need controller support and temperature-aware calibration. 论文基于112颗商用DDR4芯片,展示了利用同时多行激活实现DRAM PUF的方法。文中报告的片内Jaccard指数约为89%到95%,片间指数约为2%到4%,说明同一芯片内可重复、不同芯片间可区分。该结果的意义在于把一种存储扰动机制转化为认证原语,但实际部署仍需要内存控制器支持和温度校准。

A Large-Scale Laboratory for Modern DRAM Characterization 面向现代DRAM表征的大规模实验平台

A. Olgun, H. Luo, I. E. Yuksel, et al.

arXiv:2606.13725 · 2026-06-11T08:49:31Z

The authors describe an updated DRAM characterization laboratory built around DRAM Bender, with broader experiment support, more interface standards, and easier large-scale operation. The contribution is infrastructure rather than a single device result, but that matters for finding real failure, refresh, timing, and security behavior in commodity memory. As memory scaling slows and reliability margins tighten, reproducible characterization platforms become a prerequisite for credible DRAM mechanisms. 作者介绍了围绕DRAM Bender更新的大规模DRAM表征实验室,扩展了实验类型、接口标准支持,并降低了大规模使用门槛。该工作的贡献更多是基础设施,而非单一器件结果,但它对发现商用内存中的失效、刷新、时序和安全行为很重要。随着存储缩放放缓、可靠性裕量收紧,可复现实测平台正成为可信DRAM机制研究的前提。

Circuits & Design 电路与设计

PANDA Links Analog Design Intent to Layout Generation PANDA把模拟设计意图连接到版图生成

H. Zhang, W. Fan, X. Gao, B. Liu, R. Wang, Y. Lin

arXiv:2606.15052 · 2026-06-13T01:48:52Z

PANDA is an LLM-enhanced analog design framework that carries high-level intent through topology synthesis, sizing, and constraint-driven layout. The authors claim the workflow can reduce analog design turnaround from days or weeks to hours while improving performance. The notable point is the cross-stage dependency handling; analog automation often fails when topology, sizing, and layout are optimized as separate islands. PANDA是一个LLM增强的模拟电路设计框架,把高层设计意图贯穿拓扑综合、尺寸优化和约束驱动版图生成。作者称该流程可把模拟设计周期从数天或数周缩短到数小时,并改善性能。其关键在于处理跨阶段依赖;模拟自动化常常失败于把拓扑、尺寸和版图当作彼此割裂的环节。

AI Accelerators & Compute-in-Memory AI加速器与存算一体

SupraSNN Exploits Synapse-Level Parallelism in SNN Accelerators SupraSNN在SNN加速器中利用突触级并行

S. S. Ghavami, M. H. Nikkhah, M. R. Roshanshah, S. Safari

arXiv:2606.13354 · 2026-06-11T13:41:41Z

SupraSNN proposes a hardware-software co-design for spiking neural network accelerators that treats synaptic events like parallel micro-operations. The architecture separates synaptic and neuronal computation with multicast and merge trees, aiming to expose more parallelism without duplicating expensive neuron-state logic. It is relevant because SNN efficiency claims often depend on whether irregular sparse events can be scheduled efficiently in real hardware. SupraSNN提出一种面向脉冲神经网络加速器的软硬件协同设计,把突触事件视为可并行执行的微操作。该架构通过多播树和归并树分离突触计算与神经元计算,试图在不重复昂贵神经元状态逻辑的情况下暴露更多并行性。该工作重要,因为SNN的能效优势往往取决于稀疏且不规则的事件能否在真实硬件中高效调度。

BenDi Uses Quasi-Stochastic Systolic Compute for Edge Bioelectronics BenDi用准随机脉动架构加速边缘生物电子

B. Ye, Y. Pan, S. Agwa, T. Prodromakis

arXiv:2606.12235 · 2026-06-10T15:42:21Z

BenDi is an edge bioelectronics accelerator that combines low supply voltage, quasi-stochastic multiplication, a systolic dataflow, and hardware-aware quantization. In a commercial 22nm implementation, the authors report 3.35x smaller area and 5x higher energy efficiency versus a reference design at 0.5V and 100MHz. The target is practical wearable signal processing, where energy and area budgets are often tighter than raw model accuracy. BenDi是一种面向边缘生物电子的加速器,结合低供电电压、准随机乘法、脉动数据流和硬件感知量化。作者在商用22nm工艺实现中报告,在0.5V、100MHz条件下,相比参考设计面积缩小3.35倍、能效提高5倍。其目标是实际可穿戴信号处理,这类场景的能耗和面积预算往往比单纯模型精度更苛刻。

Tiara Adds a Programmable Line-Rate ISA to Memory-Side NICs Tiara为存储侧NIC加入可编程线速ISA

B. Li

arXiv:2606.13708 · 2026-06-10T05:01:17Z

Tiara introduces a compact, statically verifiable ISA that runs on the memory-side NIC to resolve pointer indirection near remote memory. On an FPGA prototype, it reduces 10-hop graph traversal latency by 2.85x over one-sided RDMA and sustains 3.4x higher throughput. The idea is important for disaggregated memory and paged KV-cache serving, where dependent address lookups can turn network latency into the dominant cost. Tiara提出一种紧凑、可静态验证的ISA,在存储侧NIC上执行,以便靠近远端内存解析指针间接访问。FPGA原型显示,在10跳图遍历中,相比单边RDMA延迟降低2.85倍、吞吐提高3.4倍。该思路对内存解耦和分页KV-cache服务很重要,因为依赖性地址查找会让网络往返延迟成为主导成本。

AI Systems Research AI系统研究

Lossless LLM Weight Compression Approaches the Shannon Bound 无损LLM权重压缩接近Shannon极限

H. Tan, Y. Chen, G. Alonso, W.-F. Wong, B. He

arXiv:2606.15789 · 2026-06-14T12:43:47Z

This paper finds that LLM weights have effective entropy 2x to 10x lower than their stored bitwidth across models from 1.5B to 405B parameters. It then uses tile-level Asymmetric Numeral Systems decompression aligned to GEMM tiling and reports bit rates within 0.01 to 0.1 bits of the Shannon limit. Integrated into SGLang, the approach raises Qwen-14B maximum batch size from 47 to 75, making lossless compression a credible systems lever rather than only an archival trick. 该论文发现,在1.5B到405B参数模型中,LLM权重的有效熵比存储位宽低2倍到10倍。作者随后采用与GEMM分块对齐的tile级Asymmetric Numeral Systems解压,报告码率距离Shannon极限仅0.01到0.1 bit。集成到SGLang后,该方法把Qwen-14B最大批量从47提高到75,说明无损压缩可能成为推理系统手段,而不只是归档技巧。

Spatio-Temporal Expert Prefetching for MoE Inference 面向MoE推理的时空专家预取

Y. Zhao, R. Bunescu, A. Louri, A. Karanth, K. Wang

arXiv:2606.15453 · 2026-06-13T20:09:54Z

The paper targets MoE inference latency by predicting which experts will be needed across adjacent layers and consecutive decoding tokens. Its core observation is that expert requests are not fully random; they show spatio-temporal correlation within application domains. Prefetching matters because MoE models can shift bottlenecks from arithmetic to expert loading, especially when active experts sit across slower memory or devices. 该论文通过预测相邻层和连续解码token所需专家,来降低MoE推理延迟。其核心观察是专家请求并非完全随机,而是在特定应用域内存在时空相关性。预取之所以重要,是因为MoE模型可能把瓶颈从算术计算转移到专家加载,特别是当激活专家分布在较慢内存或不同设备上时。

EDA & Verification EDA与验证

BigPower Estimates CPU Module Power from Source-Level Design Data BigPower基于源级设计数据估算CPU模块功耗

H. Zhu, C. Luo, J. Zhan

arXiv:2606.13747 · 2026-06-11T15:05:13Z

BigPower uses LLM-based representations plus hierarchy, connectivity, configuration, and workload context to estimate module-level CPU power from source-level design information. The paper evaluates the approach on the open-source XiangShan processor family. If robust across designs, this kind of surrogate could reduce reliance on slow simulation loops during architectural power exploration. BigPower结合LLM表示、层次结构、模块连接、配置参数和工作负载上下文,从源级设计信息估算CPU模块级功耗。论文在开源XiangShan处理器家族上评估该方法。如果能跨设计保持稳健,这类替代模型可降低架构功耗探索对慢速仿真循环的依赖。

HierSVA Benchmarks LLM-Driven Hierarchical Hardware Formal Verification HierSVA评测LLM驱动的层次化硬件形式验证

M. Nie, J. Zhu, J. Zhang, Z. Zeng, J. Wang, et al.

arXiv:2606.13706 · 2026-06-09T22:41:44Z

HierSVA provides a data-synthesis pipeline, dataset, and benchmark for LLM-generated SystemVerilog Assertions on hierarchical RTL. On twelve recent LLMs, the paper reports a 67.1% module-level compile rate, 82.1% non-vacuous proof rate among evaluable runs, and 70.2% detection of eligible injected faults. The mixed results are useful: LLMs can help assertion generation, but syntax success and real fault coverage remain separate bars. HierSVA提供了面向层次化RTL的LLM生成SystemVerilog Assertions数据合成流程、数据集和基准。论文在12个近期LLM上报告模块级编译率67.1%,可评估运行中非空证明率82.1%,对可检测注入故障的覆盖率为70.2%。这些结果有价值,因为它们显示LLM确实能辅助断言生成,但语法成功和真实故障覆盖仍是不同门槛。