Devices & Process
器件与工艺
Y. Jung, C.-S. Hwang, S.-M. Yoon
ACS Applied Materials & Interfaces · 2026-08-21
An a-IGZO vertical charge-trap memory transistor uses an H2O-based Al2O3 tunneling layer and nanocrystalline ZnO traps to improve charge injection and erase behavior. Devices optimized at 60 nm and validated at a 40 nm channel length produce an approximately 10 V memory window. The result is relevant to tighter 3D NAND vertical scaling, although array-level endurance, uniformity, and manufacturability are not established in the candidate record.
该a-IGZO垂直电荷俘获存储晶体管采用H2O基Al2O3隧穿层与纳米晶ZnO俘获层,以改善电荷注入和擦除行为。器件在60 nm条件下优化,并在40 nm沟道长度下验证,获得约10 V存储窗口。结果与3D NAND更紧凑的垂直缩放相关,但候选资料尚未给出阵列级耐久性、一致性与可制造性。
P. Wang, B. Jiang, W. Xue, et al.
Nature Communications · 2026-08-19
A stacked dual-ferroelectric transistor switches polarization through electrical and optical inputs, combining electronic storage, multilevel optical memory, and visual-response emulation in one device. It reports more than 8-bit conductance states, low dark current, and a wide analog dynamic range. The artificial-vision demonstration is simulated, so array scalability and system energy remain open questions despite the device-level functionality.
该双铁电堆叠晶体管通过电输入与光输入翻转极化,在单个器件中结合电子存储、多级光存储和视觉响应模拟。器件报告了超过8 bit的电导状态、低暗电流和较宽的模拟动态范围。人工视觉演示基于仿真,因此尽管器件功能丰富,阵列扩展性与系统能耗仍待验证。
C. Huang, D. Kong, J. Liang, et al.
Measurement Science and Technology · 2026-08-20
A chromatic differential arrayed-confocal microscope reconstructs package-substrate topography without lateral or axial mechanical scanning. It cuts full-resolution acquisition from 20 seconds to 65 milliseconds, a 308x throughput gain, while step-height and surface-roughness deviations stay below 3% and 1% versus a commercial laser confocal tool. The comparison is on selected substrates, so robustness across reflective materials and production defect modes still needs broader validation.
该色差阵列共焦显微系统无需横向或轴向机械扫描即可重建封装基板三维形貌。其全分辨率采集时间从20秒降至65毫秒,吞吐量提升308倍;相较商用激光共焦设备,台阶高度与表面粗糙度偏差分别低于3%和1%。目前比较基于选定基板,面对不同反射材料与量产缺陷模式的稳健性仍需更广泛验证。
AI Accelerators & Compute-in-Memory
AI加速器与存算一体
D. B. Venkatesh, S. Radhakrishnan, R. Rakshit, et al.
arXiv:2608.19046 · 2026-08-19
APEX integrates the PASC-IF neuron into the LoAS accelerator framework so converted SNNs can preserve source-ANN accuracy at fewer timesteps. Its three-stage neuron datapath is fully combinational, adding no pipeline latency, while the architecture exploits sparsity in both activations and weights. This is an accelerator-design preprint rather than a measured-silicon report, so area and energy claims should be read as implementation results.
APEX把PASC-IF神经元集成到LoAS加速器框架中,使转换后的SNN能以更少时间步保持源ANN精度。其三级神经元数据通路采用全组合逻辑,不增加流水线延迟,同时利用激活与权重的双重稀疏性。该工作是加速器设计预印本而非实测芯片报告,因此面积与能效结论应视为实现结果。
A. Adair, J. Robertson, A. Tsiamis, et al.
arXiv:2608.16807 · 2026-08-17
A compact micro-LED array optically programs multiple form-free IGZO RRAM devices in parallel, using 450 nm light for SET and electrical pulses for RESET. Persistent photocurrent provides fading-memory behavior, and the prototype writes spatial patterns across several devices. The laboratory demonstration connects optical fan-out with neuromorphic state storage, but larger-array uniformity, endurance, and programming energy are not yet shown.
紧凑型微型LED阵列可并行光编程多个无需forming的IGZO RRAM器件,以450 nm蓝光执行SET、电脉冲执行RESET。持续光电流提供了衰减记忆行为,原型还在多个器件上写入空间图案。该实验把光学扇出与神经形态状态存储连接起来,但尚未展示更大阵列的一致性、耐久性与编程能耗。
Hardware-Relevant AI Research
面向硬件的AI研究
T. Ding, J. Liu, H. Xu
arXiv:2608.16477 · 2026-08-17
Pallas prepares LLM inference state at a predicted target base station before cellular handover, avoiding a cold transfer after the user moves. It reconstructs the stable prefix locally while streaming newly generated suffix KV blocks from the source, then assembles a current cache at handoff. The approach trades interruption time for prediction and duplicate-compute overhead, so its benefit depends on handover accuracy and available edge capacity.
Pallas在蜂窝切换发生前,把LLM推理状态预先准备到预测的目标基站,避免用户移动后再进行冷迁移。系统在目标端本地重建稳定前缀,同时从源端流式传输新生成后缀的KV块,并在切换时组装为最新cache。该方案以预测开销和重复计算换取更短中断时间,因此收益取决于切换预测准确率与边缘算力余量。
D. Z. Tootaghaj, K. Diab, B. Lantz, et al.
ACM Digital Library · 2026-08-17
DynamoServe pools stranded GPU memory to hold model weights and KV caches for multi-tenant LLM serving. Coordinated placement and demand-driven weight migration reduce fragmentation while trying to preserve locality and latency. The reported evaluation improves memory efficiency without a latency penalty, though the candidate record provides no workload-level figures for judging gains across cluster topologies.
DynamoServe汇聚分散在GPU上的闲置显存,用于存放多租户LLM服务的模型权重与KV cache。协同放置与按需权重迁移减少资源碎片,同时尽量保持局部性和延迟。论文报告显存效率提升且没有延迟惩罚,但候选资料未提供按工作负载拆分的数据,难以判断不同集群拓扑下的收益。
Q. Yang, R. Yang, H. Xiao, et al.
arXiv:2608.15602 · 2026-08-16
FluxBin pairs post-training binary decomposition with a CUDA lookup-table kernel to avoid floating-point arithmetic and runtime dequantization in ultra-low-bit LLM inference. Hessian-guided salient bases preserve critical weights, while virtual column mapping turns irregular work into denser execution; the paper reports up to 5.92x speedup. Results are tied to the tested CUDA kernels and models, so portability and accuracy at broader model scales remain the main caveats.
FluxBin把训练后二值分解与CUDA查找表内核结合,在超低比特LLM推理中避开浮点运算和运行时反量化。Hessian引导的显著性基底保留关键权重,虚拟列映射则把不规则计算转化为更稠密的执行;论文报告最高5.92倍加速。结果与测试所用CUDA内核和模型绑定,跨平台可移植性及更大模型上的精度仍是主要限制。