1. OpenAI Launches GPT-6 Astra with 1.05M Context and Computer-Use Capabilities
OpenAI has released GPT-6 Astra, its next-generation model optimized for computer-use and software engineering tasks. Astra introduces a 1.05M-token context window that manages state using a new note-keeping method across windows instead of compaction. Built on a looped transformer architecture that reuses layers to save hosting RAM, the model is priced at $10/M input and $50/M output. It is rolling out to Daybreak Access members, ChatGPT paid tiers, the OpenAI API, and AWS.
- • GPT-6 Astra features a 1,050,000-token context window and a 128,000 maximum output token limit.
- • API pricing is set at $10.00 per million input tokens and $50.00 per million output tokens, with a 90% discount for cached reads.
- • The model achieves 74.1% on the DeepSWE v1.1 benchmark and 99.9% on ARC-AGI-3 using the Provider Adapter harness.
- • Astra utilizes a looped transformer architecture that reuses layers to increase capacity without increasing storage or RAM requirements.
- • The model is designated as reaching OpenAI's Critical cybersecurity threshold, restricting advanced exploit discovery for standard users.
Developers gain access to a highly advanced computer-use and coding model that drastically reduces token usage per task, despite a higher base API price.
2. Meta Details Muse Spark 1.3 Performance and Contributor Pricing
Meta has provided further details on the Muse Spark 1.3 model released yesterday, highlighting a 1M-token context window and improved efficiency. The model reduces tool calls by 20% and token usage by 25% compared to version 1.2. Meta is also offering a 95% discount to users who opt to share their prompts and outputs to help train future models.
- • Muse Spark 1.3 features a 1M-token context window.
- • Efficiency improvements include a 20% reduction in tool calls and 25% lower token usage.
- • The model achieved a 75.4 score on the DeepSWE v1.1 benchmark.
- • A 95% discount is available for users who opt to share data for model training.
- • The highest reasoning mode remains gated for safety testing.
Developers can now evaluate the model's efficiency gains and cost-saving opportunities through the newly disclosed contributor program.
3. IFM Releases K2 Horizon Open-Weight Model Family with 512K Context
The Institute of Foundation Models (IFM) at MBZUAI has released K2 Horizon, a connected fleet of six open-weight models. The release is highly transparent, providing the complete training lifecycle, intermediate checkpoints, training data recipes, and the xLLM training infrastructure. The models range from a lightweight 0.9B variant to a massive 375B-A23B sparse Mixture-of-Experts (MoE) model. The 36B-A4B model introduces a Mixture-of-Value-Attention (MoVA) mechanism to optimize active parameters, while the larger models support a native 512K token context window.
- • The K2 Horizon family includes six models: 375B-A23B, 36B-A4B, 32B, 7B, 3.7B, and 0.9B variants.
- • Models are released under the Apache 2.0 license, with intermediate checkpoints, training data recipes, and code fully open-sourced.
- • The 36B-A4B and 375B-A23B models utilize Mixture-of-Experts (MoE) architectures, with the 36B model featuring a new Mixture-of-Value-Attention (MoVA) mechanism.
- • The models support a native context window of up to 524,288 tokens.
- • Day-zero deployment support is available for NVIDIA, AMD, and Cerebras hardware via vLLM, SGLang, and Ollama.
Developers get access to a highly capable, fully open-weight model family with a massive 512K context window and day-zero support in vLLM, SGLang, and Ollama.
4. Microsoft Releases MAI-Transcribe-2 with 72% Price Cut
Building on the MAI-Transcribe-1.5 model introduced in June 2026, Microsoft has released MAI-Transcribe-2. The new model offers a 72% price reduction to $0.10 per hour and increases processing speed to 411x real-time, while improving accuracy to a 2.0% word error rate. It is now available via Microsoft Foundry and the MAI Playground.
- • MAI-Transcribe-2 is priced at $0.10 per hour, a 72% reduction from the $6 per 1,000 minutes cost of MAI-Transcribe-1.5.
- • The model achieves a 2.0% word error rate, improving on the 2.4% rate of its predecessor.
- • Processing speed has increased to 411x real-time, up from 276x in the previous version.
- • New features include support for 60 languages, speaker diarization, and word-level timestamps.
- • Available now through Microsoft Foundry and the MAI Playground.
The update significantly lowers the cost and latency for developers using Microsoft's proprietary transcription models compared to the previous generation.
5. Inworld Moves Realtime TTS-2 to General Availability
Inworld has transitioned its Realtime TTS-2 model from a research preview to a full commercial release. The model, which supports over 100 languages and plain-text style instructions, has now secured the #1 position on the Artificial Analysis Controlled Voice Arena with an Elo score of 1,123. It is available to developers at a price of $20.83 per 1 million characters.
- • Realtime TTS-2 is now generally available following its initial research preview.
- • The model is priced at $20.83 per 1 million characters.
- • It holds the #1 rank on the Artificial Analysis Controlled Voice Arena with an Elo score of 1,123.
- • Supports over 100 languages with automatic language detection.
- • Achieves a generation throughput of 106 characters per second.
Developers can now access the production-ready version of the model, which offers high-quality, multi-lingual speech generation at a throughput of 106 characters per second.
6. Alibaba Releases Wan 3.0 Video Generation and Editing Model
Alibaba has released Wan 3.0, a major upgrade to its video generation model family. The model is capable of generating up to 30 seconds of 1080p video complete with native audio. It supports a wide range of creative inputs, including text, images, audio, and web pages, and allows for instruction-led editing of visuals, plot, dialogue, and sound. Wan 3.0 is currently in public preview on Alibaba Cloud Model Studio.
- • Wan 3.0 generates up to 30 seconds of 1080p video with native audio.
- • The model supports text-to-video, image-to-video, reference-based generation, and instruction-led editing.
- • It ranks #1 in Video Editing with Audio and #2 in Text to Video with Audio on the Artificial Analysis Video Arena.
- • Pricing starts at $0.05/sec for 480p, $0.10/sec for 720p, and $0.20/sec for 1080p.
- • The model is available in public preview through Alibaba Cloud Model Studio.
Developers can integrate advanced video editing and generation capabilities, including style transfer and motion preservation, directly via Alibaba Cloud Model Studio.
7. sanoTTS Releases Ultra-Lightweight Neural TTS Stack for Microcontrollers
A developer has released sanoTTS, an ultra-lightweight neural text-to-speech stack designed for low-resource hardware. The model family ranges from 294k to 2.2 million parameters, with the smallest model occupying just 337kb when quantized to int8. Despite its tiny footprint, the 1.5M model can run on a $3 ESP32 microcontroller with a Real-Time Factor of 0.225, generating four seconds of audio in under a second.
- • sanoTTS supports 11 voices and 6 languages with model sizes ranging from 294k to 2.2M parameters.
- • The 294k model occupies just 337kb when quantized to int8 and runs on $3 chips with 512kb of SRAM.
- • The 1.5M model achieves a Real-Time Factor of 0.225 on an ESP32 microcontroller, generating 4 seconds of audio in 1 second.
- • The model family is available for web applications via the `sanotts-web` npm package.
- • The developer claims sanoTTS is the smallest neural TTS model ever created.
Developers can deploy high-quality, local text-to-speech capabilities directly on low-cost edge hardware and microcontrollers without an NPU.
8. Anthropic Releases Claude Commerce Agents Reference Blueprint
Anthropic has released `anthropics/commerce-agents`, an open-source reference blueprint for building shopping and merchant agents. The repository provides runnable implementations for four verticals: retail, travel, telecom, and entertainment. To optimize performance, Anthropic recommends an "agent skills" architecture over traditional intent routers, which helps minimize latency and API costs. The system also utilizes an approval gate where the model proposes actions and the harness applies them.
- • The `anthropics/commerce-agents` repository is released under the Apache 2.0 license.
- • The blueprint includes runnable verticals for retail, travel, telecom, and entertainment, running on Python 3.11+ and Node 22.
- • It supports deployment via the Claude API, Amazon Bedrock, Microsoft Foundry, and Google Cloud Vertex AI.
- • The architecture uses an "agent skills" design over intent routers to minimize latency and token costs.
- • The release includes a Claude Code plugin for scaffolding and reviewing agent implementations.
Developers can quickly scaffold and deploy production-ready commerce agents using a pre-built, optimized skills architecture that minimizes latency and token costs.
9. Cursor Cloud Agents Now Support Execution on Private Networks
Following the August 20th introduction of event-driven cloud agents, Cursor has expanded its infrastructure support to allow these agents to run on private networks. While agents were previously restricted to Cursor's standard cloud environment, they can now be dynamically scheduled onto machine pools managed directly by development teams, enabling secure access to internal services and private source control.
- • Cursor cloud agents now support execution on private, user-managed machine pools.
- • This expands upon the event-driven agent capabilities introduced in August.
- • Agents can now directly access internal services and private source control systems.
- • Agents continue to be initiated and managed through the standard Cursor platform.
This update enables teams to integrate Cursor's agentic workflows with internal databases, custom build pipelines, and private source control systems that were previously inaccessible to cloud-hosted agents.
10. Restless Launches to Help AI Agents Create and Debug API Interactions
Restless has launched a developer tool designed to streamline how AI agents interact with APIs. The platform automatically generates API documentation from existing endpoints, making it easier for agents to understand how to call them. Crucially, Restless includes an error-handling mechanism that intercepts failed API calls, generating corrective instructions that allow the agent to self-correct and retry the request in real time.
- • Restless automatically generates API documentation from existing endpoints to assist agent integration.
- • The tool features an error-handling loop that catches failed API calls and provides agents with instructions to retry.
- • Real-time API logs allow developers to monitor customer activity call by call.
- • Developers can initialize the tool locally by running `npx restless init`.
Developers can build more resilient agent integrations by letting Restless catch failed API calls and feed corrective instructions back to the agent.
11. Nvidia Launches PAIR to Link Idle Computers for Local AI Inference
Nvidia has launched Personal AI Router (PAIR), a free and open-source tool designed to network idle computers into a unified local AI inference pool. Compatible with Windows, Linux, and macOS, PAIR supports Nvidia RTX GPUs, DGX Spark systems, and Apple M4 chips. The software uses mutual TLS (mTLS) to secure communications and dynamically scales as devices join or leave the network, allowing developers to run local models across multiple machines.
- • PAIR is a free, open-source tool that links idle home computers to perform local AI inference in parallel.
- • The software is compatible with Nvidia GeForce RTX 20-series or newer, RTX Pro, DGX Spark, and Apple M4 chips or newer.
- • PAIR automatically adjusts to devices joining or leaving the network to avoid interfering with active tasks.
- • Connections are secured using a six-digit pairing code and mutual TLS (mTLS) encryption.
- • Nvidia also announced simplified local setup for agent applications like Perplexity Portable Computer, Hermes Agent, and OpenClaw.
Developers can harness the compute of multiple local machines to run larger models or speed up local inference without paying for cloud GPUs.
12. ik_llama.cpp Adds Multi-Token Prediction Support for Qwen4exp
The ik_llama.cpp project, previously noted for its MTP performance gains on Qwen3.6, has expanded its capabilities by merging MTP support for Qwen4exp models. This update allows the use of a 2.6B MTP head to draft and verify tokens, doubling inference speeds on compatible hardware while maintaining output parity.
- • PR #2369 in ik_llama.cpp adds MTP support for Qwen4exp.
- • Uses a 2.6B MTP head for drafting and verification.
- • Draft acceptance rates reach 93-99% for code and 60-65% for prose.
- • Performance on an RTX 5090 reaches 90 tokens per second.
- • Limited to single-slot operations without multi-GPU support.
Developers can now leverage MTP-driven speedups for the newer Qwen4exp architecture, extending the performance benefits previously observed in the project.