Inference Brew

Mercury 2.5 Released: Faster Speeds and Expanded Context Window

00:00 / --:--

← Back to home

Mercury 2.5 Released: Faster Speeds and Expanded Context Window

1. Mercury 2.5 Released: Faster Speeds and Expanded Context Window

Building on the foundation of the Mercury 2 model released in June, the new 2.5 version offers increased throughput, a significantly larger context window, and new features like tunable reasoning and native schema-aligned JSON output. This update provides developers with enhanced performance for real-time voice agents and complex agentic workflows.

  • • Mercury 2.5 improves generation speeds to 1,107 tokens per second and adds a 260K token context window.
  • • New features include tunable reasoning, parallel tool calls, and native schema-aligned JSON output.
  • • Standard pricing is $0.20 per million input tokens and $0.75 per million output tokens, with temporary launch discounts.
  • • Early adopters report median latencies of 170ms and 90% reductions in context compaction costs.
  • • The model is available immediately through the Inception API, Baseten, and OpenRouter.

Developers can now leverage higher performance and expanded context capabilities for real-time voice agents and complex agentic workflows while benefiting from competitive pricing.

SOURCES

2. OpenAI Releases ChatGPT Images 2.5 with New API Models

ChatGPT Images 2.5 is available across desktop, mobile, and web platforms for ChatGPT, ChatGPT Work, and Codex users. The release represents a significant step forward in multimodal speed and control, giving developers more granular tools to guide image generation via sketches and localized edits.

  • • OpenAI released two new API models: GPT-Image-2.5 Flare for general use and GPT-Image-2.5 Sunburst for high-precision creative workflows.
  • • The new image engine reduces image generation latency by up to 50% compared to the previous Images 2.0 version.
  • • The model features improved capabilities for preserving the likeness of people and pets, moving away from a singular AI aesthetic.
  • • New editing features include Sketch for drawing reference guides directly and tools to modify specific image areas while preserving previous edits.
  • • The models incorporate safety safeguards, including automatic invisible watermarking and C2PA metadata to identify AI-generated content.

Developers can integrate faster, higher-fidelity image generation and precise, sketch-guided editing workflows directly into their applications.

3. inclusionAI Expands Ling-3.0-flash Family with Multimodal VL Model

Building on the Ling-3.0-flash architecture released in early August, inclusionAI has introduced the Ling-3.0-flash-VL. This new multimodal variant adds native image and video understanding to the existing model family, utilizing a sparse MoE design and a specialized VideoRoPE component for spatial-temporal reasoning.

  • • Ling-3.0-flash-VL is a 124B parameter sparse MoE model, activating 5.5B parameters per token.
  • • The model supports a 1-million-token context window.
  • • It integrates a ViT visual encoder and a two-layer MLP projector for native multimodal processing.
  • • A specialized VideoRoPE component enables precise event localization and video editing.
  • • The model is now available on Hugging Face.

This release extends the Ling-3.0-flash ecosystem, providing developers with a high-capacity multimodal model capable of processing massive video and image inputs for complex reasoning tasks.

SOURCES

4. DeepSeek Upgrades Flash API to V4.1 with Native Multimodal Support

Building on the V4-Flash series, DeepSeek has introduced the V4.1 Flash model in beta. This update advances the previous V4-Flash API by integrating native multimodal support while maintaining the existing pricing structure. Developers can now access this upgraded architecture by updating their API model string, continuing the evolution of the Flash ecosystem.

  • • DeepSeek V4.1 Flash introduces a new architecture with native multimodal support.
  • • The beta is accessible by setting the API model name to deepseek-v4.1-flash-expires-on-0910.
  • • Pricing remains identical to the previous V4-Flash model.
  • • The beta includes a rate limit of 20 concurrent requests per account.

Developers can now access an upgraded, natively multimodal flash model by updating their API model string, building on the capabilities established in the V4-Flash ecosystem.

SOURCES

5. Reducto Releases r-1 Single-Pass Document Parser

By consolidating text, table structure, layout, reading order, and formatting into a single pass, r-1 bypasses the latency and compounding errors of multi-stage pipelines. Reducto plans to expand the lineup in the future with an r-1 mini model and automatic per-page routing.

  • • The r-1 model processes documents in a single full-page pass, replacing complex multi-stage agentic OCR pipelines.
  • • Reducto reports a 20% reduction in error rates compared to its legacy agentic document parsing pipelines.
  • • Pricing is set at a flat 1 cent per page, down from the 3 to 6 cents per page charged for legacy models.
  • • The model is available in preview through Reducto's V3 Parse API and requires a specific configuration flag to enable.
  • • Reducto is offering up to $5,000 in credits for organizations to conduct side-by-side comparisons against other parsers.
  • • There are currently no open-weights or local self-hosting options available for the r-1 model.

Developers can significantly lower their document ingestion costs while improving the accuracy of text, table, and layout extraction for RAG pipelines.

SOURCES

6. Anthropic Expands Security Warning to Include API and Subscription Token Theft

Building on the previous security measures taken against infostealer malware, Anthropic has now alerted subscribers to ongoing attempts to steal Claude API and subscription tokens. This development indicates that unauthorized access attempts have evolved beyond session hijacking, requiring developers to audit their API keys and monitor billing dashboards for anomalous usage.

  • • Anthropic issued a new warning regarding the theft of Claude API and subscription tokens.
  • • This follows previous efforts to mitigate session hijacking caused by infostealer malware.
  • • Developers are advised to rotate API keys and monitor billing dashboards for unauthorized usage spikes.

Developers must now extend their security audits beyond session management to include active API keys and subscription credentials to prevent unauthorized billing charges.

SOURCES

7. NVIDIA Expands Rust-for-GPU Initiative with New SIMT Compiler

NVIDIA has formalized its Rust-for-GPU efforts under the 'CUDA Rust' initiative. This expands upon the previously released cutile-rs (a tile-based system) by introducing cuda-oxide, a new track designed for the SIMT execution model. While cutile-rs remains available for JIT-compiled tile-based kernels, cuda-oxide provides a path for compiling Rust MIR to PTX, requiring a pinned nightly toolchain and compute capability 8.0+ hardware. Both projects are currently in alpha.

  • • The 'CUDA Rust' initiative now encompasses both the existing cutile-rs and the new cuda-oxide project.
  • • cuda-oxide targets the SIMT execution model by compiling Rust MIR to PTX.
  • • cuda-oxide requires a pinned nightly toolchain, Linux, and NVIDIA GPUs with compute capability 8.0 or higher.
  • • cutile-rs continues to support stable Rust 1.89+ and CUDA 13.3 for tile-based kernel development.
  • • Both tracks are in alpha and not yet recommended for production.

This initiative consolidates NVIDIA's Rust support, offering developers two distinct paths—SIMT and Tile—to achieve memory-safe GPU kernel development.

SOURCES

8. Infercat Shares Local AI Models via Encrypted P2P Tunnels

Infercat effectively turns local hardware into a private mini-cloud. By bypassing the need for public-facing ports or complex network configurations, it offers a highly secure and frictionless way for developers to collaborate or test local models across different devices.

  • • Infercat utilizes tailcat to establish secure, end-to-end encrypted peer-to-peer tunnels between host machines and clients.
  • • The tool supports popular local inference engines including llama.cpp, vLLM, Ollama, and LM Studio.
  • • It provides an OpenAI-compatible API, enabling seamless integration with Cursor, Claude Code, and Open WebUI.
  • • Users can generate simple invite codes to grant access without requiring external accounts or system-level VPN configurations.
  • • Host machines record basic usage statistics but do not store private conversation transcripts.

It allows developers to expose their local GPU setups to external clients, collaborators, or coding tools without setting up complex VPNs or cloud hosting.

SOURCES

9. I-Have-ADHD Plugin Strips Conversational Filler from Claude Code

By stripping away the polite but time-consuming conversational fluff typical of modern LLMs, the i-have-adhd plugin optimizes terminal-based coding assistants for pure speed and utility. It provides a highly practical way to customize local developer environments for maximum focus.

  • • The plugin enforces 10 specific rules for LLM responses, including leading with the next action and capping lists at five items.
  • • It strictly prohibits conversational elements such as preambles, recaps, and closing remarks.
  • • The tool is open-source, hosted on GitHub under the MIT license, and loosely based on ADHD coping strategies.
  • • Installation for Claude Code requires cloning the repository and using CLI commands to add the plugin to the local marketplace.

Developers can speed up their interactions with Claude Code by eliminating repetitive conversational filler and forcing the model to lead directly with actionable code.

SOURCES

10. Hip-Agent: A Lightweight 200-Line Python Agent Harness

Designed to fit directly within a prompt, hip-agent avoids the complexity of larger agent frameworks. By treating actions as shell commands and subagents as child processes, it offers a clean, Unix-like approach to building and orchestrating LLM-driven automation.

  • • The core loop of hip-agent is implemented in approximately 200 lines of Python code.
  • • The harness uses standard environment variables for configuration and executes actions via shell commands.
  • • Subagents within the harness are implemented as standard child processes.
  • • The tool includes a dedicated module for integrating with the Codex API.
  • • It relies on existing protocols and formats for its remaining agent functionality.

It gives developers a simple, highly transparent alternative to bloated agent frameworks, making it easy to inspect, customize, and debug agent execution loops.

SOURCES

11. Deltafin Runs 2.8T Kimi K3 Model Locally via SSD Streaming

Deltafin operates as an independent, open-source project with no affiliation to Moonshot AI. By avoiding pruning and instead streaming the full 2.8-trillion-parameter model from high-speed SSDs, it provides a novel pathway for local execution of massive mixture-of-experts models on accessible hardware.

  • • Deltafin achieved a steady decode speed of 1.0 token per second for a 512-token answer on an M1 Max MacBook Pro with 128 GB of RAM.
  • • The tool streams model experts from SSDs, with performance scaling from 52% on a single drive up to 90% on a four-drive configuration.
  • • A current prefill limitation causes a 512-token prompt to take 6.3 minutes to generate the first token due to repetitive expert reading.
  • • Developers can integrate Qwen models to accelerate raw text completion via speculative decoding, where the smaller model proposes tokens for Kimi K3 to verify.
  • • Deltafin includes an OpenAI-compatible server supporting standard chat and completion endpoints under an MIT license.

It allows developers to run massive frontier-class models locally on standard workstations without needing enterprise-grade GPU clusters.

SOURCES

12. Edge Browser Agent Runs Sub-1B Models on Legacy Mobile Hardware

The experiment highlights the massive performance gains of optimizing the input representation for edge agents. While the tiny models still struggle with judgment-based tasks like identifying topical decoys, the structured perception layer makes basic web scraping and data extraction highly viable on decade-old hardware.

  • • Researchers successfully ran models like Qwen3-0.6B on a 2017 Samsung Galaxy Note 8 using llama.cpp inside Termux.
  • • The system feeds models a structured webpage representation of about 200 tokens, rather than raw HTML or heavy screenshots.
  • • Sub-1B and 2B models, including Qwen3-0.6B and Llama-3.2-3B, achieved a 10/10 success rate on a sandbox book-scraping task.
  • • Using raw HTML instead of structured perception increased task execution time from 80 seconds to 22 minutes and caused task failures.
  • • The project is open-source and available on GitHub at github.com/e2llm/edge-browser-agent.

It proves that developers can build highly efficient, low-latency browser agents that run entirely on-device without requiring expensive cloud APIs or high-end mobile GPUs.

SOURCES

13. Inference Engine Benchmarks for Qwen3.8-Flash-Next at 262K Context

The benchmarks were conducted on a workstation equipped with an NVIDIA RTX PRO 6000 Blackwell GPU and an AMD Ryzen 9 9950X CPU. While accuracy on standard benchmarks like GSM8K and MATH-500 remained statistically identical across all engines, the dramatic differences in prefill times and energy consumption highlight the importance of engine selection for production-level local deployments.

  • • At a 262K context window, SGLang achieved a time-to-first-token of 35.4 seconds, compared to 80.4 seconds for FreeToken and 258.4 seconds for baseline llama.cpp.
  • • FreeToken maintained flat decode speeds of 100.1 to 94.8 tokens per second across context sweeps, while baseline llama.cpp degraded from 101.9 to 20.2 tokens per second.
  • • The llama.cpp Multi-Token Prediction fork improved decode speeds by up to 1.69x on coding tests, but only when all model experts remained on the GPU.
  • • SGLang consumed approximately 13 kJ of GPU energy for a full-window request, compared to 116 kJ for baseline llama.cpp, due to shorter request durations.
  • • Startup times to reach the first answer favored llama.cpp at 16 seconds, compared to 108 seconds for SGLang and 126 seconds for FreeToken.

Developers can drastically reduce latency and energy costs for long-context local models by selecting the optimal inference engine for their hardware.

SOURCES

14. Qwen3.8 27B Overcomes Prior Quantization Degradation Issues

Building on earlier reports that identified significant math and reasoning degradation in 27B models using Q4_K_M quantization, new evaluations of the Qwen3.8 27B model confirm that this version preserves full-precision performance on key benchmarks like GPQA Diamond and Terminal-Bench 2.1. This indicates that the latest model architecture is more resilient to standard 4-bit compression than its predecessors.

  • • Qwen3.8 27B's Q4_K_M quantization matches full-precision BF16 performance on GPQA Diamond and Terminal-Bench 2.1.
  • • This contrasts with earlier 27B model evaluations that showed a 9% drop in math accuracy under the same Q4_K_M quantization.
  • • The model fits on a single 24GB GPU while maintaining approximately 64k tokens of context.
  • • Extreme 1-bit quantization remains unusable, consistent with findings from previous model generations.

Developers can now deploy the Qwen3.8 27B model using 4-bit quantization on consumer hardware without the accuracy trade-offs previously observed in earlier 27B iterations.

SOURCES

15. Benchmarking Local Agent Concurrency on Dual RTX 4090 GPUs

The benchmarks demonstrate that aggregate prefill throughput remains constant at roughly 1,500 tokens per second regardless of concurrent agent slots. For developers looking to integrate these local models into their active workflows, the setup can be exposed as an MCP server to serve as subagents within Claude Code.

  • • A dual RTX 4090 system with 128GB RAM hit a soft cap at 5 concurrent agents with 64k context, beyond which latency increased without throughput gains.
  • • The hard cap for the system was 9 concurrent agents at 64k context, strictly limited by VRAM availability.
  • • The Qwen 27B model achieved 3.40 seconds per tool-call, significantly outperforming a 122B MoE model which took 11.91 seconds.
  • • No measurable accuracy differences were found between Q4_K_M, Q6_K_XL, and Q8_K_XL quantizations for long-range retrieval up to 251,557 tokens.
  • • Using q8_0 for the KV cache saved 1.4 GiB of VRAM per slot with no loss in quality or speed compared to f16.

Developers can optimize their local agent architectures by matching model sizes and quantization levels to specific VRAM and concurrency limits.

SOURCES

Inference Brew in your inbox

5 minutes a day. Free, unsubscribe anytime.

Inference Brew in your inbox

5 minutes a day. Free, unsubscribe anytime.