Tom's Hardware · 7h ago
Tom's Hardware · 7小时前
OpenAI used Hot Chips 2026 to present Jalapeño, its 700 W in-house AI ASIC co-developed with Broadcom. Its published comparison claims up to 1.9x higher throughput per kilowatt and 3.6x lower latency than NVIDIA's 1,400 W flagship GPU. The comparison is vendor-reported, but it makes custom inference silicon a more concrete part of OpenAI's infrastructure strategy.
OpenAI在Hot Chips 2026上展示了与Broadcom共同开发的700 W自研AI ASIC Jalapeño。其公布的对比结果宣称,相比NVIDIA的1,400 W旗舰GPU,Jalapeño每千瓦吞吐量最高提升1.9倍、时延降低3.6倍。这一对比属于厂商自报数据,但它使定制推理芯片成为OpenAI基础设施战略中更具体的一环。
ServeTheHome · 1h ago
ServeTheHome · 1小时前
Google discussed its eighth-generation TPU family at Hot Chips 2026, splitting the line into TPU 8t for training and TPU 8i for inference. The division signals a willingness to optimize silicon around distinct workload profiles rather than carry one general-purpose accelerator across both. It also keeps Google's internal hardware roadmap active alongside the rapid expansion of external AI platforms.
Google在Hot Chips 2026上介绍第八代TPU家族,并将产品线分为面向训练的TPU 8t和面向推理的TPU 8i。这种划分表明其正围绕不同负载特征优化芯片,而非用单一通用加速器覆盖两类任务。随着外部AI平台快速扩张,这也意味着Google仍在持续推进内部硬件路线图。
ServeTheHome · 2h ago
ServeTheHome · 2小时前
Microsoft provided further technical detail on Maia 200, its second-generation server AI inference processor, at Hot Chips 2026. The presentation is another indication that large cloud operators are treating inference hardware as a product-specific design problem. That approach can shift demand toward internally optimized systems even as GPU clusters remain central to training.
Microsoft在Hot Chips 2026上进一步介绍了第二代服务器AI推理处理器Maia 200的技术细节。这场演讲再次表明,大型云服务商正把推理硬件视为需要针对产品单独设计的问题。即使GPU集群仍是训练核心,这种路径也可能将需求转向内部优化的系统。
ServeTheHome · 3h ago
ServeTheHome · 3小时前
NVIDIA described Groq 3 LPU accelerators as part of heterogeneous AI compute in Vera Rubin clusters, following its Groq acquisition. The specialized LPUs are intended to reduce latency during inference decode, a stage that becomes more visible in interactive and agentic workloads. The move brings a purpose-built inference architecture into NVIDIA's broader rack-scale platform.
NVIDIA在Hot Chips上介绍了用于Vera Rubin集群异构AI计算的Groq 3 LPU加速器,此前其已收购Groq。该类专用LPU旨在降低推理解码阶段的时延,而这一阶段在交互式和智能体工作负载中愈发关键。此举将面向特定推理任务的架构纳入NVIDIA更广泛的机柜级平台。