Inference Brew

Google Upgrades Gemini Flash Series to Version 3.8

00:00 / --:--

← Back to home

Google Upgrades Gemini Flash Series to Version 3.8

1. Google Upgrades Gemini Flash Series to Version 3.8

Following the August 2026 release of Gemini 3.7 Flash, Google has introduced Gemini 3.8 Flash. This update improves reasoning and coding capabilities, though the model's increased use of iterative tool calls may raise total task costs by approximately 40%. Additionally, Google has launched Gemini 3.8 Flash Cyber, a specialized version for vulnerability detection restricted to government and trusted partners.

  • • Gemini 3.8 Flash succeeds the 3.7 Flash model with improved reasoning and coding performance.
  • • The model features a 1M token context window and is priced at $0.75 per million input tokens and $3.75 per million output tokens through the end of 2026.
  • • Increased reasoning steps and iterative tool calls can lead to a 40% higher total cost per task compared to previous versions.
  • • Google introduced Gemini 3.8 Flash Cyber, a restricted model variant tuned for automated vulnerability remediation.
  • • The model scored 59 on the Artificial Analysis Intelligence Index, a 3-point improvement over Gemini 3.7 Flash.

Developers can now access a more capable reasoning model in the Flash series, though they should account for potential cost increases due to the model's more intensive reasoning processes.

2. Meta Updates Muse Spark to Version 1.3 with Performance Gains

Meta has updated its Muse Spark model line to version 1.3, following the release of version 1.2 in August. The new version delivers performance gains in coding, agentic workflows, and scientific reasoning while maintaining the same pricing structure as its predecessor. The Muse Spark 1.3 (xhigh) variant is available now via Meta's API and Muse Code, with a more powerful 'max' variant currently in limited preview.

  • • Meta released Muse Spark 1.3, offering substantial performance improvements for coding, agentic tasks, and scientific reasoning.
  • • The Muse Spark 1.3 (xhigh) variant is available now, scoring 61 on the Artificial Analysis Intelligence Index.
  • • Pricing for the xhigh variant remains unchanged at $1.25 per million input tokens and $4.25 per million output tokens.
  • • A more powerful Muse Spark 1.3 (max) variant is currently in limited preview for partners.

Developers can now access a more capable reasoning and coding model through Meta's existing API and Muse Code ecosystem without any increase in token costs.

3. Google Launches Agentic Video Understanding for Gemini Models

Google has rolled out agentic video understanding capabilities across several models in its Gemini family. By combining native video processing tools with the models' core reasoning capabilities, this update enhances performance on complex video analysis tasks. Developers can leverage these features to build more robust applications for automated moment retrieval, anomaly detection, and precise object counting in video streams.

  • • Google launched agentic video understanding capabilities for several Gemini models.
  • • The feature combines native video tools with model reasoning.
  • • It improves performance in tasks such as moment retrieval, anomaly detection, and counting.

Developers can build applications that perform complex video tasks like moment retrieval, anomaly detection, and object counting with higher accuracy.

SOURCES

4. Multiverse Computing Releases Quasar 438B Flagship Reasoning Model

Multiverse Computing has entered the large model space with Quasar 438B, a flagship reasoning model supporting both English and Spanish. Designed for enterprise-scale agents, technical copilots, and workflow automation, the model achieved a score of 43 on the Artificial Analysis Intelligence Index v4.1.1, marking a high point for European models. It demonstrates strong capabilities in terminal environments (scoring 69.3 on Terminal-Bench v2.1) and long-document reasoning (scoring 75.0 on AA-LCR), and is accessible via the CompactifAI API.

  • • Multiverse Computing released Quasar 438B, its first large reasoning model supporting English and Spanish.
  • • The model scored 43 on the Artificial Analysis Intelligence Index v4.1.1, the highest result for a European model.
  • • It scored 75.0 on the AA-LCR long-document reasoning benchmark and 69.3 on Terminal-Bench v2.1 for terminal environments.
  • • Quasar 438B is available for testing and deployment through the CompactifAI API.
  • • The model completes a 500-token response in 15.3 seconds.

Developers have a new high-performance, bilingual option for complex terminal environments and long-document reasoning via the CompactifAI API.

SOURCES

5. Vercel Introduces Fluid Compute Layer for Dynamic Workload Scaling

Vercel has introduced Fluid, a unified compute layer engineered to dynamically configure infrastructure based on real-time workload demands. Designed to absorb sudden burst capacity, Fluid handles builds, sandboxes, and serverless functions seamlessly. The system is already operating at scale, processing over one trillion requests per month to ensure high availability and performance for modern web and AI applications.

  • • Vercel introduced Fluid, a unified compute layer that dynamically configures infrastructure for different workloads.
  • • The system is designed to absorb burst capacity for builds, sandboxes, and functions.
  • • Fluid currently scales to support more than one trillion requests per month.

Developers deploying AI apps on Vercel get more resilient, dynamically scaling infrastructure capable of handling massive traffic spikes without manual configuration.

SOURCES

6. WebLLM Enables High-Performance In-Browser LLM Inference via WebGPU

WebLLM, a companion project to MLC LLM, offers a high-performance in-browser inference engine powered by WebGPU. By running entirely client-side, it eliminates server costs and latency while maintaining compatibility with the OpenAI API standard, including support for streaming, JSON-mode, and function-calling. Developers can easily integrate the engine into web applications using NPM or CDN, offload heavy computations to Web Workers to keep the UI responsive, and deploy custom models compiled in the MLC format.

  • • WebLLM runs language models entirely within the browser using WebGPU for hardware acceleration.
  • • The engine is compatible with the OpenAI API, supporting streaming, JSON-mode, and function-calling.
  • • It natively supports popular open models including Llama 3, Phi 3, Gemma, Mistral, and Qwen.
  • • Developers can integrate WebLLM via NPM, Yarn, or CDN, and offload computations to Web Workers or Service Workers.
  • • The engine provides optional integrity verification for model artifacts using Subresource Integrity (SRI) hashes.

Developers can build fully local, private web applications that run models directly in the user's browser with OpenAI-compatible APIs.

SOURCES

7. Perplexity Open-Sources Lily for Optimized On-Device Apple Silicon Inference

Perplexity AI has open-sourced Lily, a lightweight Mac inference server designed to maximize on-device LLM performance on Apple Silicon. Specifically tuned for the Qwen3.6-35B-A3B model—which features sparse MoE routing and Gated DeltaNet layers—Lily leverages unified memory and Apple's specialized hardware to deliver prefill and decode throughput that surpasses MLX-LM. The repository is available on GitHub, enabling developers to host highly efficient local inference on a single Mac.

  • • Perplexity AI open-sourced the 'lily' repository on GitHub, designed as a Mac inference server.
  • • The engine is specifically optimized for the Qwen3.6-35B-A3B model to achieve peak performance on Apple Silicon.
  • • Lily leverages unified memory and specialized hardware to deliver higher prefill and decode throughput than MLX-LM.
  • • The Qwen3.6-35B-A3B model utilizes sparse MoE routing and Gated DeltaNet layers.

Developers can run large, sparse MoE models locally on a single Mac with prefill and decode throughput that outperforms MLX-LM.

SOURCES

8. Worlds via Code Generates Explorable 3D Scenes Using Claude Fable 5.1

The "Worlds via code" project has demonstrated a novel pipeline that uses autonomous Claude Fable 5.1 agent swarms to generate browser-native, explorable 3D reconstructions of real-world locations. Built entirely as Three.js applications without proprietary game engines, the current demo reconstructs a detailed San Francisco square featuring 129 storefronts, working traffic lights, and interactive interiors. The entire pipeline—which spans reconnaissance, Blender asset generation, runtime assembly, and Playwright-based quality assurance—is open-sourced under the MIT license.

  • • The "Worlds via code" project creates browser-native, explorable reconstructions of real locations using Claude Fable 5.1 agent swarms.
  • • Reconstructions are built as Three.js applications without game engines or proprietary 3D tiles.
  • • The current demo reconstructs a San Francisco square with 129 storefronts, working traffic lights, and explorable interiors.
  • • The pipeline uses a four-stage process: reconnaissance, offline asset generation in Blender, runtime assembly, and camera-match QA via Playwright.
  • • The project's code and assets are licensed under the MIT license, using OpenStreetMap and USGS 3DEP data.

Developers can study or adopt an MIT-licensed pipeline that translates real-world spatial data into interactive Three.js applications without relying on proprietary game engines.

SOURCES

9. Switch Runs AI Agents Locally Across Slack, Teams, and Discord

Switch has launched as a local software tool designed to connect AI models and agents directly to Slack, Microsoft Teams, and Discord. Operating entirely locally and requiring no account creation, Switch allows developers to test and run agentic workflows within their existing chat channels without platform lock-in or the need to migrate their underlying infrastructure.

  • • Switch allows users to run AI models, infrastructure, and agents locally.
  • • The tool integrates agents with Slack, Microsoft Teams, and Discord in minutes.
  • • It requires no account to use and avoids platform lock-in or migration.

Developers can quickly deploy and test conversational agents in Slack, Teams, or Discord without setting up complex cloud infrastructure or creating external accounts.

SOURCES

10. Hugging Face Releases Optimized WebGPU Kernels for Browser Inference

Hugging Face has launched @huggingface/kernels, a dedicated library containing 207 highly optimized WebGPU kernels. The library is built to accelerate client-side AI model inference, allowing web developers to run complex neural network operations directly within the browser with minimal latency and maximum hardware utilization.

  • • Hugging Face introduced the @huggingface/kernels library.
  • • The library contains 207 optimized WebGPU kernels.
  • • It is designed to accelerate AI model inference directly within web browsers.

Web developers can leverage these pre-optimized kernels to significantly speed up client-side model execution and build more responsive local AI features.

SOURCES

11. Anthropic Launches C2PA File-Checking Tool and Detection API Preview

Building on its August commitment to implement invisible watermarking, Anthropic has now launched a tool that reads cryptographically signed C2PA metadata to verify if files were processed by Claude. The tool operates locally on user devices. Additionally, the company has opened a private preview of its Detection API for text watermarks, enabling organizations to verify content provenance in compliance with EU AI regulations.

  • • Anthropic has released a local tool to verify C2PA metadata in .png, .jpg, and .svg files.
  • • A private preview of the Detection API for text watermarks is now available for eligible organizations.
  • • These releases fulfill the company's prior commitment to provide technical means for detecting Claude-generated content.
  • • The file-checking tool ensures privacy by processing metadata locally without sending files to Anthropic.

These tools provide the practical implementation of the provenance standards Anthropic previously announced, allowing developers and organizations to verify Claude-generated content.

SOURCES

12. Unsloth Enables GGUF Quantization for DeepSeek-V4-Flash-Vision-Exp

Building on the August 31 release of DeepSeek-V4-Flash-Vision-Exp weights on Hugging Face, Unsloth has now merged GGUF format support for the model. This update enables developers to run quantized versions of the experimental vision-capable model locally, facilitating integration into resource-constrained environments.

  • • Unsloth added GGUF support for the DeepSeek-V4-Flash-Vision-Exp model.
  • • This follows the August 31 release of the model's weights on Hugging Face.
  • • The update enables local, quantized inference for the experimental vision model.

This development allows developers to move from simply hosting the raw weights to running optimized, quantized versions of the DeepSeek-V4-Flash-Vision-Exp model locally.

SOURCES

13. llama.cpp Officially Adds Metal Support for Apple M5 Neural Accelerators

The llama.cpp project has updated its Metal backend to natively support M5 matrix multiplication and neural accelerators. This official integration addresses the performance gap previously requiring custom w8a8 kernels, pushing prefill performance for Unsloth's Q_8 GGUF models to 300-350 tokens per second on M5 hardware. With generation speeds reaching 19 tokens per second for a Qwen3.8 27B model in the Q8 range on an M5 Pro, llama.cpp now achieves performance parity with native MLX models.

  • • llama.cpp now includes native Metal support for M5 matmul and neural accelerators.
  • • This update replaces the need for previously reported custom w8a8 kernels to unlock M5 performance.
  • • Prefill performance for Unsloth's Q_8 GGUF models on M5 hardware now reaches 300-350 tokens per second.
  • • Generation speeds of approximately 19 tokens per second were observed for the Qwen3.8 27b model in the Q8 range on an M5 Pro.

Developers can now leverage M5 hardware acceleration directly within the main llama.cpp project without relying on third-party custom kernels to bypass default 16-bit activation limits.

SOURCES

14. VoxGen Releases Lightweight, AMD-Optimized TTS Inference Engine

VoxGen has been released as a lightweight, native text-to-speech inference engine designed specifically for VoxCPM2 models. Written in Rust and powered by Vulkan compute, VoxGen bypasses Python, PyTorch, and CUDA entirely, offering optimized performance on AMD graphics cards (including a dedicated mode for the XTX 7900). The engine runs directly from the shell, making it straightforward for developers to integrate local, low-latency speech synthesis into their applications and scripts.

  • • VoxGen is a lightweight native inference engine for VoxCPM2 models written in Rust.
  • • The engine utilizes Vulkan compute instead of Python, PyTorch, or CUDA, optimizing performance for AMD GPUs.
  • • It features a specific optimization mode for the AMD Radeon XTX 7900.
  • • The application runs from a shell, allowing easy integration with external scripts and programs.
  • • Installation requires the VoxGen executable and two model files from Hugging Face.

Developers can run local, high-performance speech synthesis on AMD hardware without relying on heavy Python or PyTorch dependencies.

SOURCES

15. Mistral Details Data Training Opt-Out Settings for Vibe and API Users

Mistral has outlined the specific steps required for users and developers to opt out of having their input and output data used for model training. While Vibe Enterprise accounts are opted out by default, standard Vibe, mobile, and API users must manually adjust their settings. Developers using Mistral Studio or the API can secure their data by disabling the "Anonymous improvement data" toggle in the Admin panel, keeping in mind that Vibe and API opt-out configurations are managed separately.

  • • Mistral may use user conversations and uploaded documents for model training unless users opt out.
  • • Vibe Enterprise customers are opted out of training by default, but standard Vibe users must manually opt out via settings.
  • • API and Mistral Studio users can opt out by disabling the "Anonymous improvement data" toggle in the Admin panel.
  • • Mobile app users on iOS and Android can opt out by deselecting "Enable data sharing" in settings.
  • • Opt-out settings for Vibe and the API are separate and must be configured individually.

Developers and teams must manually configure their privacy settings in the Mistral Admin panel to prevent their proprietary inputs and outputs from being used for model training.

SOURCES

Inference Brew in your inbox

5 minutes a day. Free, unsubscribe anytime.

Inference Brew in your inbox

5 minutes a day. Free, unsubscribe anytime.