Jalapeño's first results show industry-leading speed and efficiency in AI inference
OpenAI published the first independently-benchmarkable results for Jalapeño, its custom inference chip, using the public InferenceX benchmark from SemiAnalysis. Tested against commercial systems on GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T, Jalapeño delivered 1.5–1.9x more throughput per watt and 1.7–3.6x lower end-to-end latency — on Kimi K2.5, the largest model tested, that's roughly 3.4x lower latency at comparable throughput. The chip is rated at 700W (measured sustained draw stayed at or below 550W) versus 1,200–1,400W for the GB200/GB300 systems it's compared against. OpenAI plans to start deploying Jalapeño in its own infrastructure by year-end, with a second and third generation already in development.
semianalysis (via hn-algolia) · OpenAI Jalapeño: Better than Nvidia Blackwell — 2026-08-25 19:36 IST
Sources & depth
NVIDIA Announces Jetson Orin Nano 2 Robotics Computer to Redefine Entry-Level Edge AI
NVIDIA announced the Jetson Orin Nano 2, an update to its entry-level robotics computer, claiming 2x the inference performance of the original Orin Nano Super in the same form factor while using 40% less power at matched performance (78 TOPS, 8GB memory, 8-core Arm CPU). The module and dev kit are aimed at running Cosmos, Nemotron, Gemma 4, and Qwen 3 models directly on small robots, drones, and vision systems — NVIDIA says more than 3 million developers already build on its robotics stack, with early adopters including Cognex, Doosan Bobcat, Matic Robots, and Alphabet's Wing. Availability is set for the first half of 2027; NVIDIA hasn't disclosed pricing.
Sources & depth
Unsloth ships Auto Compaction and LAN Remote Access in a 170+ PR bug-fix release
Unsloth shipped a 170+ PR bug-fix release for its fine-tuning stack, alongside an experimental Auto Compaction feature that lets chats run past the model's context limit, and new LAN/remote access support including a keyless, passwordless local API mode. The release also fixes MLX runtime issues on Mac and AMD bugs affecting Strix Halo and other RDNA GPUs, following up on last week's Qwen3.8-27B and Unsloth Desktop support.
Sources & depth
Quantization-Aware Healing: a 4-bit model that beats its own full-precision original
Multiverse Computing published results for "quantization-aware healing" (QAH), a recovery technique for models that have been both structurally compressed and quantized to 4-bit. Tested on a GPT-OSS 120B model shrunk to 60B parameters and quantized to MXFP4, the QAH model beat its own bf16 (full-precision) source on 7 of 9 benchmarks — with the biggest gains on long-context reasoning (+7.4 points) and math (+5.6 points) — while using roughly a quarter of the weight memory. The technique distills from the original pre-compression model rather than the already-degraded compressed checkpoint, and reached peak accuracy in about 100 training steps versus 700+ for standard quantization-aware training before that approach started to degrade.
Sources & depth