Thursday, Sep 3, 2026

Top stories

Gemini 3.8 Flash and 3.8 Flash Cyber

Google shipped Gemini 3.8 Flash on 2026-09-02, the next entry in the Gemini 3 Flash line, together with a restricted specialized variant and a new access program built around it. Standard 3.8 Flash is generally available now via the Gemini API, AI Studio, Android Studio, and consumer apps at an introductory $0.75/M input and $3.75/M output tokens (through 2026-12-31), with a 1M-token input window; Google frames its gains over 3.7 Flash as concentrated in software engineering, agentic tasks, and multi-step reasoning. Gemini 3.8 Flash Cyber is a separate deployment aimed at vulnerability discovery and automated patching — Google claims over 70% success finding vulnerabilities across 20 programming languages — and ships exclusively through the newly-launched Fairwind Program, a limited-access initiative for governments, critical-infrastructure operators, and cybersecurity partners (650+ organizations already participating) to autonomously find and generate deployment-ready patches, paired with Google's existing CodeMender tool.

Sources & depth

Muse Spark 1.3

Meta released Muse Spark 1.3, a coding- and agentic-focused LLM (not a generative-media model despite the Muse branding) trained for long-horizon, multi-step agentic workflows with native multimodal perception across video, images, and documents. It ships alongside Muse Code, a terminal-based multi-agent coding tool, and is reachable through Meta's new self-serve Model API (OpenAI SDK-compatible) at $1.25/M input and $4.25/M output tokens standard, or a cheaper $0.10/$0.20 "Contributor" tier. Meta's launch post reports gains over Muse Spark 1.2 on agent, coding, instruction-following, and long-context evaluations and cites GPT-5.6 Sol and Claude Opus 5 as reference points in those comparisons — specific score deltas weren't visible in this session's fetch, so treat the competitive framing as directional rather than verified.

Sources & depth

2x faster Qwen3.8-Flash-Next + GLM-5.3-Flash via Unsloth MTP

Unsloth shipped v0.1.805-beta/v0.1.806-beta with multi-token prediction (MTP) enabled by default for the two locally-runnable models in this thread, delivering up to 2x faster generation on GLM-5.3-Flash and Qwen3.8-Flash-Next (MTP can still be disabled). The release bundles 170+ other changes: long Qwen chats on Apple Silicon get up to 30x faster follow-up turns via improved MLX batching, GLM tool-calling is now reliable across long multi-turn sessions, and new audio-model support (MiniMax-Music3, Higgs, MOSS) lands alongside GGUF export for GLM-5.3 MLX fine-tunes.

Sources & depth
Also today