Inference Brew

Archive.

All past insights, collected.

Anthropic Releases Claude Opus 5.5 with Lower Pricing and Enhanced Safety →

SpaceXAI Updates Frontier Model to Grok 4.7 →

Alibaba Releases Qwen-Image-2.1 Open-Weight 7B Model →

TypeSafe AI Details Jev API Pricing and Functionality →

Alibaba Launches Hosted Qwen3.8-Omni-Flash API →

PrismML Updates Bonsai 27B Line with Qwen3.8-based Ternary Bonsai 2 →

Google Moves Gemini 3.5 Transcribe to General Availability and Launches Gemini 3.8 Live →

UkisAI Optimizes Qwen 3.8 27B to Reduce Overthinking by 58% →

Shanghai AI Lab Releases Intern-S2-397B Multimodal Model →

Agnes-3.0-Flash 33B Multimodal Preview Released on Hugging Face →

OpenAI Opens GPT-Live-1 to API Developers →

OpenAI Launches Agents API in Public Beta with Managed Codex Harness →

Meta Expands Muse Ecosystem with Consumer AI Agent Platform →

Mercury 2.5 Released: Faster Speeds and Expanded Context Window →

OpenBMB Releases MiniCPM5-2B On-Device Model →

H Company Releases NeoMME Single-Tower Multimodal Encoders →

OKF Agent Memory Launches Git-Native Persistent Memory →

Microsoft Launches MAI-Image-2.6-Flash; Meta Expands Muse Image Capabilities →

OpenAI Launches GPT-6 Astra with 1.05M Context and Computer-Use Capabilities →

Google Upgrades Gemini Flash Series to Version 3.8 →

Anthropic Upgrades Claude 5 Models to 5.1 with 75% Cheaper Cache Reads →

Tencent Releases Hy4 Preview, a 770B Parameter Open-Weight Model →

MirroS Releases Code-as-World Open-Weight Models for Video-to-Physics Translation →

Tencent Releases Hy4 Preview with 1M Context Window →

Z.ai Releases Open Weights for GLM-5.3 Model →

Ox Alpha Revealed as GLM-5.3-Flash with Ultra-Low-Cost API →

Z.ai Expands GLM-5.3 Family with Multimodal Flash Model →

NVIDIA Groq 3 LPX Inference Accelerators Enter Full Production →

Anthropic's Cheaper Opus 5 Overtakes Fable 5 in Corporate Spending →

Qwen 3.8 27B Successfully Reverse-Engineers arm64 Code in 30 Minutes →

Mysterious 'Ox Alpha' Model Debuts with 100 Trillion Free Tokens Daily →

OpenAI Adjusts GPT-5.6 Sol Pricing Following Initial Reduction →

Google Updates DiffusionGemma Performance to 1,500 Tokens Per Second →

DeepReinforce Updates Ornith Model Family with 1.5 Release →

Alibaba Releases Open-Weights Qwen3.8-27B Model →

Cartesia Releases Sonic 3.6, Succeeding Sonic 3.5 as Top-Ranked TTS Model →

Claude System Prompts Updated for Web and Mobile Apps →

End-to-End Guide Released for Fine-Tuning Tool-Calling LLMs →

Qwen3.8-27B Open Weights Released Following Preview →

OpenAI Adds 'Ultrafast' Tier to GPT-5.6 Sol Powered by Cerebras →

SpaceXAI Upgrades Frontier Model to Grok 4.6 →

Nvidia Launches NeMo Switchyard and Nemotron 3.5 Lightning →

Meta Releases Muse Glimmer, a 30B Open-Weight Agentic Model →

Zerank 2 and F2LLM V2 Achieve Top Local Multilingual Retrieval Benchmarks →

Claude Code Introduces Cross-Session Messaging for Multi-Agent Coordination →

OpenAI Implements Unlimited Text Chats and GPT-5.6 Luna for Free Users →

OpenAI Integrates GPT-5.6 into ChatGPT with New 'Think' Slider and Unlimited Free Text →

Meta Launches Muse Code Terminal Agent with Muse Spark 1.2 →

Mistral Releases Shieldstral 3B Multimodal Moderation Model →

Alibaba Launches Qwen3.8-Max API and Sets Open-Weight Release Date →

DeepSeek-V4-Flash Reasoning Effort Modes Exhibit Verbosity and API Discrepancies →

ByteDance Launches Seedance 2.5 with 30-Second Generation and Multimodal Controls →

DeepSeek-V4-Flash API Enters Public Beta →

Google DeepMind Releases Gemini Robotics 2 →

Google Integrates Gemini 3.6 Flash into Managed Agents with New Developer Controls →

Google Launches Gemini Distillation Service →

Moonshot AI Releases Kimi K3 Open-Weights Model →

KAT-Coder-V2.5 Benchmark Results and AutoBuilder Details Released →

Anthropic Updates Context Engineering Guidelines for Claude 5 Models →

Anthropic Releases Claude Opus 5 with Adjustable Thinking and Relaxed Code Safety →

Microsoft Integrates In-House MAI Models into Core Products, Replacing OpenAI →

Microsoft Releases Fara1.5-27B Browser Agent Model →

Google Officially Launches Gemini 3.6 Flash and 3.5 Flash-Lite →

DeepSeek-V4 API Moves to Production Release →

Hugging Face Breach Highlights How API Guardrails Block Incident Response →

Kimi K3 Open Model Matches Claude Coding Performance at Lower Cost →

NVIDIA Releases Nemotron 3 Embed Open Model Collection →

Moonshot AI Announces Kimi K3 2.8-Trillion-Parameter MoE Model →

Thinking Machines Lab Releases Inkling, Its First Public Multimodal Model →

Mistral AI Releases Robostral Navigate for Vision-Based Robot Navigation →

PrismML Compresses Qwen 3.6 27B Model to Run Locally on iPhone →

Moondream 3.1 Released with 9B Parameter Mixture-of-Experts Architecture →

Grok Build CLI Uploads Entire Repositories and Secrets to xAI Cloud →

Ant Group Releases LingBot-World-Infinity Causal Video World Model →

OpenAI Launches GPT-5.6 Models and ChatGPT Work Agent into General Availability →

SpaceXAI Graduates Coding Model to Grok 4.5 →

OpenAI Updates Realtime API with GPT-Realtime-2.1 and 2.1-mini →

Tencent Releases Apache 2.0-Licensed Hy3 MoE Model →

Distilled LivePortrait Model Runs at 25 FPS in the Browser via WebGPU →

GPT-5.5 Codex Exhibits Reasoning-Token Clustering Anomaly →

Mistral AI Releases Leanstral 1.5 for Lean 4 Code Verification →

Anthropic Implements Usage Quotas and Expands Access for Redeployed Fable 5 and Mythos 5 →

Anthropic Restores Global Access to Claude Fable 5 →

Anthropic Releases Claude Sonnet 5 →

Meituan Officially Releases LongCat-2.0 1.6T MoE Model →

OpenRouter's Owl Alpha Revealed as Meituan LongCat-2.0-Preview →

OpenAI Fully Releases Sol, Terra, and Luna Cybersecurity Models →

OpenAI Launches Limited Preview of GPT-5.6 Model Family →

OpenAI Upgrades GPT-5.5 Instant with Enhanced Intent Recognition →

Google Integrates Native Computer Use into Gemini 3.5 Flash →

Anthropic Launches Claude Tag Agentic Slack Teammate →

Z.ai Officially Releases GLM-5.2 Open-Weights Model →

Qwen 3.6 27B Abliterated Released to Reduce Refusal Rates →

GLM 5.2 Token Optimization and Local Performance Metrics →

OpenAI Prepares GPT-5.6 Launch with 1.5M Context Window →

Poolside Releases Laguna M.1 225B Mixture-of-Experts Model →

Z AI Officially Releases GLM-5.2 Open-Weights Model →

Z.ai Releases GLM-5.2 Open-Weights Model with 1M Context Window →

OpenMythos Open-Weights Cybersecurity Model Released on Hugging Face →

Fable-5 and Kimi-K2.7-Code Top Autoresearch Benchmarks →

Anthropic Suspends Claude Fable 5 and Mythos 5 Globally Following US Export Control Order →

Moonshot AI Releases Kimi K2.7-Code with 30% Thinking Token Reduction →

Researchers Introduce Latent Context Language Models for 16x Input Compression →

Google Releases DiffusionGemma, a 26B MoE Model Generating Text 4x Faster →

Anthropic Releases Claude Fable 5 and Mythos 5 →

Apple Unveils Siri AI and Foundation Models Framework at WWDC 2026 →

Harness-1 20B Retrieval Subagent Released with Stateful Search Harness →

Google Releases Colab CLI for Remote GPU and TPU Execution →

Google Releases Gemma 4 Quantization-Aware Training Checkpoints →

Stanford and Lambda Labs Release OpenJarvis Local Agent Framework →

Google Releases Gemma 4 12B with Encoder-Free Multimodal Architecture →

Microsoft Launches MAI Model Family Led by MAI-Thinking-1 Reasoning Model →

MiniMax Releases M3 Model with 1M-Token Context and Desktop Control →

GitHub Copilot Transitions to Token-Based Billing Model →

Hermes Agent Introduces Tool Search to Handle Large MCP Catalogs →

Undocumented Configurations Uncovered in Claude Code v2.1.87 →

Anthropic Launches Claude Opus 4.8 and Claude Code Dynamic Workflows →

EAGLE 3.1 Speculative Decoding Integrates into vLLM →

Critical BadHost Vulnerability Discovered in Starlette Package →

Model Context Protocol Release Candidate Introduces Stateless HTTP Core →

Build Complete LLM Observability Pipelines with Langfuse →

llama.cpp Server Adds Native Agentic Tool Execution →

GBrain: Open-Source MCP Memory Layer for AI Agents →

Alibaba Launches Qwen3.7-Max with Anthropic API Compatibility →

Cohere Releases Command A+ Under Apache 2.0 →

Google Releases Gemini 3.5 Flash with High-Speed Agentic Capabilities →

AI Supply Chain Vulnerabilities →

BitLocker Bypass Vulnerability Disclosed →

Multi-Token Prediction Merged into llama.cpp →

OpenAI Consolidates Product Teams for Agentic Future →

Cerebras Systems IPO Debut →

Anthropic Surpasses OpenAI in Enterprise Adoption →

Shai-Hulud Worm Targets AI Coding Agents →

OpenAI Launches 'Daybreak' Cybersecurity Initiative →

Gemini API File Search Adds Multimodal Support →

GitHub Releases Spec-Kit for AI Coding Agents →

Anthropic Reports Progress on Agentic Misalignment →

OpenAI Releases GPT-Realtime-2 →

Anthropic Partners with SpaceX for Massive Compute Expansion →

OpenAI Releases GPT-5.5 Instant →

Microsoft Agent 365 Reaches General Availability →

Sakana AI introduces KAME tandem speech-to-speech architecture →

Stash: An open-source, continuous memory layer for AI agents →

Researchers Identify Command Execution Flaw in Model Context Protocol →

PyTorch Lightning package compromised in supply chain attack →

Mistral Releases Medium 3.5 and Vibe Remote Agents →

Amazon Bedrock adds OpenAI models in limited preview →

OpenAI and Microsoft end exclusive cloud partnership →

xAI Releases grok-voice-think-fast-1.0 Model →

Dive into Claude Code: The Design Space of Today's and Future AI Agent Systems →

DeepSeek previews DeepSeek-V4 model series →

OpenAI releases GPT-5.5 and GPT-5.5 Pro →

OpenAI releases Privacy Filter model on Hugging Face →

Google releases Deep Research and Deep Research Max agents →

GitHub pauses Copilot signups and shifts to token-based billing →

Vercel Confirms Security Breach via Compromised Third-Party AI Tool →

Open Agents cloud-based coding framework →

iTerm2 vulnerability discovery allows arbitrary code execution →

Claude Opus 4.7: Anthropic's new frontier model for agentic coding and complex tasks →

Gas Town framework secretly consumes user LLM credits for upstream bug fixes →

Claude Code Routines (Research Preview) →

MiniMax M2.7 Open-Sourced: Self-Evolving Agent Model with 56.22% SWE-Pro Score →

Claude Mythos Preview on Vertex AI →

Anthropic Claude API adds Advisor tool for hybrid model routing →

Vercel plugin for Claude Code injects hidden telemetry prompts →

Meta previews proprietary Muse Spark model with multimodal and reasoning capabilities →

Zhipu AI Releases GLM-5.1 Open-Source 754B Model →

OpenAI Apps SDK: Third-Party Integrations in ChatGPT via MCP →

TermHub: Open-Source Terminal Control Gateway for AI Agents →

TurboQuant-WASM Brings Google's Vector Quantization to the Browser →

OpenClaw CVE-2026-33579: Privilege escalation vulnerability patched →

Google Releases Gemma 4 Under Apache 2.0 License →

Veo 3.1 Lite: Lower-Cost Video Generation via Gemini API →

Urgent Security: Malicious Axios versions drop remote access trojan →

Microsoft Copilot injects advertisements into GitHub pull requests →

Claude Code Bug Silently Overwrites Local Repositories →

NVIDIA Releases ProRL Agent for Multi-Turn LLM Training →

Telnyx PyPI Package Compromised in Supply Chain Attack →

OpenTelemetry Profiles Enters Public Alpha →

GitHub Copilot Updates Training Data Policy →

Critical Supply Chain Attack Compromises LiteLLM PyPI Package →

Anthropic Launches Computer Use for Claude →

Flash-MoE Runs 397B Parameter Model on a MacBook Pro →

Building Uncertainty-Aware LLM Systems →

Anthropic Introduces Claude Code Channels →

Cursor Launches Composer 2 Coding Model →

Xiaomi Releases 1T-Parameter MiMo-V2-Pro LLM →

OpenAI Releases GPT-5.4 Mini and Nano for Coding Workloads →

Nvidia Agent Toolkit and NemoClaw Platform →

LangChain Releases Deep Agents for Multi-Step Stateful Workflows →

AWS to Deploy Cerebras Wafer-Scale Engine for AI Inference →

Anthropic Makes 1M Context Window Generally Available for Claude 4.6 →

Amazon Imposes 90-Day Code Safety Reset Following AI-Assisted Outages →

NVIDIA Releases Nemotron 3 Super 120B Hybrid Model →

Yann LeCun's AMI Labs Raises $1.03B for World Models →

Anthropic Launches "Code Review" for Claude Code →

FlashAttention-4 Achieves 1605 TFLOPs/s on Blackwell GPUs →

Google Releases TensorFlow 2.21 and LiteRT →

Anthropic Challenges Pentagon 'Supply-Chain Risk' Label in Court →

OpenAI Launches GPT-5.4 with 1M Context and Computer Use Mode →

Google Launches Gemini 3.1 Flash-Lite for High-Volume Workloads →

Anthropic vs. Pentagon Lawsuit and Claude Service Outages →

OpenAI Raises $110B at $730B Valuation with Amazon and Nvidia →

OpenAI Faces Financial Strain with $15M Daily Burn and Ad Rollout →

Alphabet Folds Intrinsic Robotics Back into Google →

Google Launches Nano Banana 2 Image Model →

Inference Brew in your inbox

5 minutes a day. Free, unsubscribe anytime.

Inference Brew in your inbox

5 minutes a day. Free, unsubscribe anytime.