semi·news
Headlines要闻 / Research研究 / /
Research digest · Saturday, August 29, 2026 研究摘要 · 2026年8月29日 星期六

Inference Efficiency Meets Design-Flow Constraints 推理效率遇上设计流程约束

This week’s work ranges from low-cost model training and speculative decoding to VLM acceleration and EDA orchestration. Together, the papers frame hardware efficiency as a systems problem spanning algorithms, deployment, and design tools. 本周研究覆盖低成本模型训练、推测式解码、VLM加速和EDA编排。这些工作共同表明,硬件效率是横跨算法、部署和设计工具的系统性问题。

Look-back window: 7 days · 6 paper(s) 回溯窗口: 7天 · 6篇

AI Accelerators & Compute-in-Memory AI加速器与存算一体

Vision-Centric Generative AI Models: A Software-Hardware Perspective 以视觉为中心的生成式AI模型:软硬件视角

E. Tselepi, C. Sestito, S. Agwa, T. Prodromakis

arXiv:2608.27199 · 2026-08-27T14:41:50Z

This perspective maps vision-generative model families to deployment domains and compares their parameter cost and energy efficiency across accelerator platforms. It argues that edge uses such as vehicles, sensors, and mobile devices require software-hardware co-design from the outset, rather than adapting hardware after models have already grown. 该观点论文将视觉生成模型家族映射到不同部署领域,并比较其在多类加速器平台上的参数成本和能效。论文认为,车辆、传感器和移动设备等边缘应用需要从一开始就进行软硬件协同设计,而不是在模型规模增长后再被动调整硬件。

AI Systems & Hardware Co-Design AI系统与软硬件协同设计

Puro-2B: A Low-Cost Qwen2-1.5B Training Recipe on RTX 5090 Puro-2B:在RTX 5090上低成本训练Qwen2-1.5B的方案

K. Luo, J. Cui, Y. Yin, et al.

arXiv:2608.27370 · 2026-08-27T17:07:12Z

Puro-2B presents an open pretraining recipe that trains a 2B-parameter model on consumer RTX 5090 GPUs using FP8 and up to 1.4 trillion tokens. The authors report a best-model compute cost below $6.9K while approaching Qwen2.5-1.5B under their evaluation protocol, making the result a useful cost data point rather than a direct substitute for frontier-scale training. Puro-2B提出一套开放预训练方案,采用FP8并在消费级RTX 5090 GPU上处理最多1.4万亿token来训练20亿参数模型。作者报告最佳模型的计算成本低于6900美元,并在其评测协议下接近Qwen2.5-1.5B;这一结果提供了有价值的成本数据点,但并不能直接替代前沿规模训练。

Beyond Parallel Blindness: Information Floors and Model Gaps in Block Drafting 超越并行盲区:块级草稿生成中的信息下界与模型差距

X. Qiang, X. Fang, C. Chen, et al.

arXiv:2608.27339 · 2026-08-27T16:40:45Z

The paper separates speculative block-drafting rejections into an unavoidable information floor and a model-quality gap. Across four domains and several target models, one realized token removes 86–100% of the final-slot information floor, while current drafters remain substantially above that floor. 论文将推测式块级草稿生成的拒绝划分为不可避免的信息下界和模型质量差距。在四个领域及多个目标模型上,一个已实现token可消除最终位置86%至100%的信息下界,但现有草稿器仍明显高于这一界限。

PACE: Pixel-Adaptive Condense and Extract for Fast VLM Inference PACE:用于快速VLM推理的像素自适应压缩与提取

J. Liu, S. Ye, X. Chen

arXiv:2608.27206 · 2026-08-27T14:52:09Z

PACE is a training-free VLM inference method that reduces work before the vision encoder and then selectively retains visual tokens for the language model. By treating input pixels and post-encoder tokens as one pipeline, it targets latency that conventional token-pruning methods leave in visual encoding. PACE是一种免训练的VLM推理方法:先在视觉编码器之前减少计算量,再为语言模型选择性保留视觉token。它将输入像素和编码器后的token视为同一条流水线,从而瞄准传统token剪枝方法遗留在视觉编码阶段的延迟。

Electronic Design Automation 电子设计自动化

LLMs in Digital EDA: From Generation to Orchestration 数字EDA中的LLM:从生成走向编排

M. Youngman, C. Sestito, T. Prodromakis

arXiv:2608.27184 · 2026-08-27T14:29:35Z

This perspective classifies LLM roles in digital EDA as generators, agents, and orchestrators coordinating decisions across design stages. It warns that code which appears syntactically plausible can still be physically incorrect, while fragmented tools and lost context obstruct industrial-scale use. 该观点论文将LLM在数字EDA中的角色分为生成器、代理和跨设计阶段协调决策的编排器。论文警告,看似语法正确的代码仍可能在物理实现上错误,而工具碎片化和上下文丢失仍阻碍工业规模应用。

Quantum Technologies 量子技术

Diamond Quantum-Sensing Platform with Integrated Microwave Antenna and Thermometer 集成微波天线与温度计的金刚石量子传感平台

M. Ohkuma, R. Matsumoto, S. Adachi, et al.

arXiv:2608.27160 · 2026-08-27T14:11:17Z

The authors integrate an NV-center quantum-sensing layer, a boron-doped diamond microwave antenna, and a local thermometer on one diamond substrate. The platform detected laser-induced heating not clearly resolved by a stage thermometer while imaging superconductors at cryogenic temperature. 作者在同一金刚石衬底上集成NV中心量子传感层、硼掺杂金刚石微波天线和局部温度计。该平台在低温下对超导体进行成像时,检测到了台式温度计未能清晰分辨的激光局部升温。

Newsletter 邮件订阅

Daily semiconductor briefing. 每日半导体简报。