1. Meta Expands Muse Ecosystem with Consumer AI Agent Platform
Following the release of the Muse Spark 1.3 model and the Muse Code terminal agent, Meta has launched the Muse personal AI agent platform. Available on iOS, Android, and web, the platform uses Muse Spark 1.3 to automate personal tasks like email management and travel booking. To ensure security, each user is assigned a dedicated Muse Secure VM, with a Sentinel agent overseeing network egress and connector actions.
- • Meta launched the Muse personal AI agent platform for iOS, Android, and web.
- • The platform is powered by the Muse Spark 1.3 model, previously released for developer use.
- • Each user receives a dedicated Muse Secure VM to isolate browser and credential data.
- • A security agent, Sentinel, manages network egress and connector actions using surrogate tokens.
- • The platform automates personal tasks including email, travel booking, and bill negotiation.
This launch brings Meta's agentic capabilities to consumer devices, moving beyond the developer-focused Muse Code terminal agent and raw model API.
2. Mercury 2.5 Pricing Detailed: 80% Discount Now Active
Building on the release of Mercury 2.5 announced on September 8, Inception Labs has now detailed the model's pricing structure. The model is available at $0.04 per million input tokens and $0.15 per million output tokens, representing an 80% discount from the previously noted standard rates. Mercury 2.5 continues to offer a 260K-token context window and generation speeds of 1,107 tokens per second.
- • Mercury 2.5 is now available at $0.04 per million input tokens and $0.15 per million output tokens.
- • This pricing represents an 80% discount on the model's standard rates.
- • The model maintains its 260K-token context window and 1,107 tokens per second generation speed.
Developers can now calculate the specific cost-savings for integrating the new 260K-context model into their production workflows.
3. Suno Releases v6 AI Music Models Trained on Licensed Data
Suno has launched its v6 AI music model suite, developed in collaboration with major industry partners including Warner Music Group, BMG, and Believe following the settlement of legal disputes. The release includes three models: the flagship v6, the experimental v6-wild, and a free v6-mini tier. The new models introduce advanced capabilities such as plain-language editing of specific song parts, track element isolation, and the ability to use images, video, or audio as reference prompts.
- • Suno released its v6 AI music model suite, consisting of the flagship v6, experimental v6-wild, and free v6-mini models.
- • The models were trained from the ground up on a new dataset including licensed content from Warner Music Group, BMG, and Believe.
- • New features include plain-language editing of specific song parts, combining elements from a user's library, and prompting via text, audio, images, or video.
- • Access to the flagship v6 and v6-wild models requires a Pro ($8/month) or Premier ($24/month) subscription.
- • The models demonstrate improved genre understanding but consistently output harmonic and rhythmic perfection, unable to produce natural imperfections.
Developers building audio and creative applications can leverage Suno's new v6 models to enable plain-language editing, track combining, and multimodal reference prompts.
4. Desert Ant Labs Launches 18 On-Device Intelligence Models
European startup Desert Ant Labs has launched a suite of 18 specialized on-device intelligence models designed to run locally on hardware as old as five-year-old phones. Accessible via a unified SDK for Swift, Kotlin, and JavaScript, the suite includes Voz (a transcription model 4.7x faster than Whisper), Clear (a 9MB audio enhancer), Redact (a real-time PII masking tool), and Clips (a 284MB video clip generator). The models are free to deploy for up to 100,000 monthly active devices, offering a zero-token-cost alternative to cloud APIs.
- • Desert Ant Labs launched 18 specialized on-device models (12 stable, 6 beta) designed to run locally on devices, including older phones.
- • The models are accessible via an SDK supporting Swift, Kotlin, and JavaScript.
- • Key models include Voz (transcription 4.7x faster than Whisper), Clear (9MB studio-quality audio enhancer), and Redact (real-time PII masking in 27 languages).
- • Clips is a 284MB model that generates video clips 10x faster and with 470x less energy than Claude Sonnet.
- • The models are free to use for up to 100,000 monthly active devices.
Developers can integrate fast, local audio, video, and text processing features into mobile and web apps without cloud API costs or latency.
5. Tencent Releases AuK Speech Generation and Editing Model
Tencent has released AuK, a unified speech generation and editing model that operates via a natural-language interface. AuK supports zero-shot text-to-speech (TTS), audio enhancement, separation, and instruction-driven editing of content, acoustics, and paralinguistics. For latency-sensitive applications, Tencent has also released AuK-Flash, a distilled version that runs 4.5 times faster than the standard model.
- • Tencent released AuK, a unified model for speech generation and editing.
- • The model supports zero-shot text-to-speech (TTS), instruction-driven editing of content, acoustics, and paralinguistics, and audio enhancement.
- • AuK utilizes a natural-language interface for all operations.
- • A distilled version called AuK-Flash operates 4.5 times faster than the standard version.
Developers can build advanced voice applications that edit existing speech, enhance audio, and perform zero-shot TTS using a natural-language interface.
6. Gradium Launches Voice Design API for Synthetic Voice Generation
Gradium has launched Voice Design, a tool available via API and Studio that generates brand-new synthetic voices from written descriptions without requiring reference audio. Supporting English, French, Spanish, Portuguese, and German, the system processes prompts of 1 to 500 characters to return up to 5 non-deterministic voice candidates in 3 to 5 seconds. The tool is available on Gradium's free tier, with unconverted candidates automatically deleted after 30 days.
- • Gradium's Voice Design generates new synthetic voices from written descriptions of 1 to 500 characters without requiring reference audio.
- • The tool is available in the Gradium API and Studio, including on a free tier, and supports English, French, Spanish, Portuguese, and German.
- • The system returns 1 to 5 non-deterministic voice candidates in 3 to 5 seconds.
- • Unconverted voice candidates are deleted after 30 days and have restricted functionality.
- • Gradium achieved a 72.6% win rate in blind pairwise listening tests against five other voice design systems.
Developers can programmatically generate custom, non-deterministic synthetic voices in seconds using simple text prompts instead of reference audio.
7. Google Open-Sources Mantis Security Toolkit
Google has officially open-sourced the Mantis security toolkit under an Apache 2.0 license, building on the Mantis Skills framework introduced in July. The toolkit provides modular slash commands and rules compatible with frameworks like Gemini CLI and Google ADK, enabling agents to manage the full vulnerability lifecycle—from detection and sandboxed reproduction to patching and re-attack verification.
- • Google has open-sourced the Mantis security toolkit under an Apache 2.0 license.
- • The release follows the initial introduction of the Mantis Skills framework in July 2026.
- • The toolkit supports end-to-end vulnerability management, including sandboxed reproduction and patch re-attack.
- • Compatible with Gemini CLI, Antigravity CLI, and Google ADK.
- • Features a hierarchical summary tree to reduce token overhead by over 85 percent.
Developers can now integrate the previously announced Mantis framework into their production pipelines to automate security reviews with sandboxed reproduction and patch verification.
8. OtoDock 1.6.0 Launches Self-Hosted Agent OS with Sandboxed Kernels
OtoDock 1.6.0 has been released as a self-hosted company OS designed to run persistent AI agents in secure environments. The platform integrates capabilities similar to Claude Code and supports Anthropic, OpenAI, or local models. To ensure security, agents run within a kernel sandbox using bubblewrap and network isolation via pasta, and can connect to remote machines via outbound WebSockets to eliminate the need for inbound ports or VPNs. OtoDock is available on GitHub under a Fair Source license and is free for up to 5 users.
- • OtoDock is a self-hosted company OS for business management and multi-tenant collaboration.
- • It integrates capabilities similar to Claude Code, Cowork, and cloud sessions into a single application.
- • Agents run as persistent processes within a kernel sandbox using bubblewrap and network isolation via pasta.
- • Agents can run on remote computers using outbound WebSockets, eliminating the need for inbound ports or VPNs.
- • OtoDock is licensed under a Fair Source license, is free for up to 5 users, and is installed via Docker Compose.
Developers can deploy and manage persistent, multi-tenant coding and business agents locally or remotely without exposing inbound ports.
9. Apple Introduces Reference Image API to Authenticate Photos
Apple has announced a hardware-based photo authentication tool called Apple Reference Image, arriving with the iPhone 18 Pro lineup. When set to Reference mode, a new camera sensor signs every captured pixel, which Apple's Private Cloud Compute then processes into an unalterable reference image. To help developers integrate this verification, Apple is releasing a Reference Image API across iOS, iPadOS, and macOS, enabling third-party applications to inspect signed photos and compare them against edited versions to detect AI manipulation.
- • Apple's Reference Image feature uses a new camera sensor to sign every pixel captured in Reference mode.
- • Apple's Private Cloud Compute processes the signed sensor data into an unalterable reference image viewable in the Photos app.
- • A Reference Image API will allow developers to inspect signed photos in third-party apps across iOS, iPadOS, and macOS.
- • The feature is built directly into hardware, distinguishing it from standards like SynthID, C2PA, and Meta's Content Seal.
- • The feature will launch with the iPhone 18 Pro lineup later this month, though EU users will receive development and viewing support in iOS 27.
Developers can use the new Reference Image API to programmatically inspect signed, unalterable photos in iOS, iPadOS, and macOS apps to detect AI edits.
10. New Megakernel Serving Engine Accelerates North Mini Code
Following the June release of the North Mini Code model, a new serving engine has been developed to optimize its deployment. By leveraging a decode megakernel, the engine achieves 1.58x faster performance than vLLM across all batch sizes. It supports continuous batching, paged attention, and context lengths up to 256K, while providing an OpenAI-compatible endpoint with tool calling support.
- • The new engine is specifically optimized for North Mini Code, the 30B MoE model released in June.
- • Performance is 1.58x faster than vLLM across various batch sizes.
- • Achieves 292 tokens per second at batch size 1.
- • Supports continuous batching, paged attention, and 256K context lengths.
- • Includes an OpenAI-compatible endpoint with tool calling support.
Developers can now deploy the North Mini Code model with significantly reduced latency and improved throughput compared to standard vLLM implementations.
11. Artificial Analysis Launches Model Release Pages for Effort Levels
Artificial Analysis has launched Model Release pages designed to help developers navigate the complex performance and cost profiles of frontier models. With models like GPT-6 Astra and Claude Fable 5.1 now offering up to six distinct effort levels, these pages provide side-by-side comparisons of cost per task, output speed, latency, and domain-specific Capability Index scores. This tool allows developers to make data-driven decisions when configuring effort levels for their specific application workloads.
- • Artificial Analysis launched Model Release pages to compare intelligence, cost, and speed across different model effort levels.
- • Frontier models now feature up to six different effort levels, which significantly alter performance and cost profiles.
- • The pages provide side-by-side comparisons of effort levels, including cost per task, output speed, and latency.
- • Capability Index scores are provided for domains including Engineering, Finance, Legal, and Healthcare.
- • GPT-6 Astra ranges from 46–53 on the Intelligence Index, while Claude Fable 5.1 ranges from 47–53.
Developers can evaluate and optimize the trade-offs between cost, latency, and accuracy for models like GPT-6 Astra and Claude Fable 5.1.
12. Run DeepSeek-V4-Flash-Vision-Exp Locally on Consumer GPUs
A new open-source project enables running the 285B DeepSeek-V4-Flash-Vision-Exp MoE model locally on consumer hardware using 10 to 12 RTX 3090 GPUs. Utilizing FP4 experts and FP8 attention, the setup achieves decoding speeds of over 120 tokens per second on 12 GPUs with speculative decoding. The project provides a pre-built Docker image alongside runtime patches that resolve memory and scheduling issues, supporting up to a 1M token context window entirely in GPU memory.
- • DeepSeek-V4-Flash-Vision-Exp is a 285B MoE model utilizing FP4 experts and FP8 attention with 157 GB of weights.
- • The model can run on consumer hardware using 10 to 12 RTX 3090 GPUs with an SM86-compatible vLLM build.
- • Performance reaches over 60 tokens per second on 10 GPUs and over 120 tokens per second on 12 GPUs using speculative decoding.
- • The system supports a 1M token context window without offload and up to 4M tokens with RAM offload.
- • The project provides a pre-built Docker image, build guides, start scripts, and runtime patches on GitHub.
Developers can self-host a massive 285B vision-capable MoE model on 10 to 12 consumer GPUs, achieving up to 120 tokens per second decoding.
13. Qwen3.8-Flash-Next Released on MLX-serve with 1M Context
The co-creator of the Qwen3.8-Flash-Next engine support in MLX-serve has announced its release, bringing 1-million-token context capabilities to Apple Silicon. Tested on an M5 Max 128GB system, the engine utilizes an 8-bit KV cache, 8-bit dense layer quantization, and 4-bit expert layer quantization to maintain high output quality. Running the full 1 million context requires setting the GPU wired limit to 120,000 MB, delivering generation speeds of up to 75 tokens per second for coding.
- • The Qwen3.8-Flash-Next engine support in MLX-serve supports up to a 1 million token context length.
- • The engine utilizes an 8-bit KV cache and requires approximately 117GB of peak memory on an M5 Max 128GB system.
- • The model uses 8-bit quantization for dense layers and 4-bit quantization for expert layers to preserve quality.
- • Generation speeds reach 40 tokens per second for prose and 75 tokens per second for coding at 1 million context.
- • Source code and model weights are available on GitHub and Hugging Face, including an MLX Serve Monitor plugin.
Developers can run local inference with a 1-million-token context window on Apple Silicon Macs, achieving up to 75 tokens per second for coding tasks.
14. Mentria.ai Enables 1-Bit 27B Model Inference in the Browser
Mentria.ai has introduced a browser-based inference engine built on WebGPU and WGSL that runs large models entirely client-side. By optimizing the 1-bit matrix-by-vector kernel to eliminate scratch memory bank conflicts, the engine can run Prism ML's 1-bit Bonsai-27B model at 25 to 30 tokens per second on an RTX 3060 Laptop GPU. The platform requires no installation or server-side processing, and supports hot-swappable LoRA adapters, a vision tower, and smaller Qwen3.5 models.
- • Mentria.ai is a browser-based inference engine built using WebGPU and WGSL.
- • The engine runs the 1-bit Bonsai-27B model (requiring 3.8 GB of GPU memory) at 25 to 30 tokens per second on an RTX 3060 Laptop.
- • Performance gains were achieved by optimizing the 1-bit matrix-by-vector kernel to resolve bank conflicts in scratch memory.
- • The engine operates entirely within the browser without requiring installations or server-side processing.
- • The platform supports hot-swappable LoRA adapters, a vision tower for image input, and smaller Qwen3.5 models.
Developers can deploy large, highly quantized models directly to client browsers with zero server-side costs or installations.
15. Cosmos3 (64B) INT4 Quantization Implemented for CUDA and MLX
An open-source implementation of the 64-billion-parameter Cosmos3 model brings INT4 text-to-image and image-to-video capabilities to local hardware. Compatible with both CUDA and Apple Silicon MLX runtimes, the project hosts quantized weights on Hugging Face and code on GitHub. On an M4 Max 128 GB Mac, generating a single video clip takes approximately 5 minutes.
- • The 64B Cosmos3 model has been implemented for INT4 text-to-image and image-to-video tasks.
- • Code is available on GitHub, and model weights are hosted on Hugging Face.
- • Generating a single video clip takes approximately 5 minutes on an M4 Max 128 GB Mac.
- • The implementation supports both CUDA and Apple Silicon MLX runtimes.
Developers can run high-fidelity local image and video generation tasks on Apple Silicon and Nvidia hardware using optimized INT4 weights.
16. Foundation-1 Audio Model Enables Timbre-Locked Keybeds
An independent researcher has released Foundation-1, a custom audio model designed to generate infinite one-shots and turn text prompts into playable synths. By separating timbre control from the instrument, the model achieves timbre-locked keybeds that remain consistent across multiple diffusion calls. The model weights are hosted on Hugging Face, and the complete inference pipeline and tools have been open-sourced on GitHub under the RC-stable-audio-tools repository.
- • Foundation-1 is a custom AI model that controls timbre as a separate element from the instrument.
- • The model achieves timbre-locked keybeds that remain consistent across multiple diffusion calls.
- • The model is available on Hugging Face, and the inferencing pipeline is open-sourced on GitHub under RC-stable-audio-tools.
- • The system can generate infinite one-shots for music production and turn text prompts into fully playable synths.
Developers building music production tools can leverage open-weights models to generate consistent, playable synths from text prompts.