Inference Brew

Z.ai Officially Releases GLM-5.2 Open-Weights Model

00:00 / --:--

← Back to home

Z.ai Officially Releases GLM-5.2 Open-Weights Model

1. Z.ai Officially Releases GLM-5.2 Open-Weights Model

Building on the announcement from June 13, Z.ai has now made the GLM-5.2 model weights publicly available. The 744-billion-parameter model, which features a 1-million-token context window, is now accessible for self-hosting via frameworks like vLLM, SGLang, Transformers, and llama.cpp. While the model excels at long-horizon reasoning and agentic tasks, it remains a text-only model without multimodal capabilities.

  • GLM-5.2 is now available as an open-weights model under the MIT license.
  • The model features 744 billion total parameters and a 1-million-token context window.
  • It is optimized for long-horizon reasoning and agentic tasks, supporting three thinking effort modes.
  • Local execution is supported via Unsloth Dynamic GGUFs, with quantization options available to reduce disk space requirements.

Developers can now move from API-based testing to self-hosting the full reasoning model for local agentic workflows at a fraction of the cost of proprietary alternatives.

2. Gemma 4 Models Leverage QAT for High-Quality 4-Bit Local Inference

While Gemma 4 offers significant technical advantages—such as native MTP and high-quality low-bit quantization—the increased training time has slowed the adoption of custom fine-tunes. However, early community releases like Equinox and MeroMero demonstrate growing developer interest in the architecture.

  • Gemma 4 models are available in 12B, 26B-A4B, and 31B sizes under the Apache 2.0 license, featuring native image and video understanding.
  • Quantization-Aware Training (QAT) allows 4-bit quantized versions to maintain quality levels comparable to the BF16 base models.
  • The 12B model fits into 8GB of VRAM, while the 31B model fits into 20-24GB of VRAM.
  • Finetuning Gemma 4 models can take up to twice as long because it requires training on both the original BF16 and unquantized QAT versions.
  • All models feature global Multi-Token Prediction (MTP) support.

Developers can run highly capable, Apache 2.0-licensed multimodal models locally on consumer GPUs with as little as 8GB of VRAM.

SOURCES

3. Inception Labs Launches Mercury 2 Diffusion-Based Reasoning Model

By shifting from traditional autoregressive generation to a diffusion-based architecture, Mercury 2 achieves massive throughput gains. This makes it a specialized tool for developers who need to run high-frequency, multi-step agent loops where latency has historically been a limiting factor.

  • Mercury 2 generates approximately 1,000 tokens per second, outperforming Google's DiffusionGemma.
  • The model utilizes diffusion technology, adapting the generation processes typically used in image models to text reasoning.
  • It is optimized specifically for high-volume, speed-sensitive workflows rather than complex, frontier-level reasoning tasks.
  • Mercury 2 is available exclusively through API or cloud services.

Developers can build ultra-low-latency workflows and high-throughput agent pipelines that require rapid, structured reasoning.

SOURCES

4. Alibaba Releases HappyHorse 1.1 Video Generation Model with Multi-Image Reference

Alibaba's upgrade represents a significant step in multimodal generation, consolidating multiple assets into a single generation pass. The release coincides with Alibaba's massive $52.7 billion global cloud network expansion, though the company was recently added to the Pentagon's list of Chinese military companies.

  • HappyHorse 1.1 is available on Alibaba Cloud Model Studio with full API access and a 40% launch discount for the first two weeks.
  • The model features R2V multi-image reference capability for character consistency, improved motion quality, and zero-drift lip sync.
  • It is built on a 15-billion-parameter unified self-attention Transformer that processes text, image, video, and audio in a single generation pass.
  • The previous version, HappyHorse 1.0, holds the No. 2 position on the Artificial Analysis Video Arena leaderboards.

Developers can build advanced video generation and editing features with consistent characters and synchronized audio using a unified, API-accessible model.

SOURCES

5. InclusionAI Releases Ling-2.6 and Ultra-Fast Ling-mini-2.0 Models

The release provides a scalable range of options for local deployment, from the massive trillion-parameter flagship down to highly optimized edge models. The performance of the mini model on standard CPU and low-end GPU hardware makes it a strong candidate for on-device agent tasks.

  • InclusionAI released base models for the 1-trillion-parameter Ling-2.6-1T and the 100-billion-parameter Ling-2.6-flash.
  • Ling-2.6-flash can run on 24GB or 32GB of VRAM when using Q4 quantization.
  • The smaller Ling-mini-2.0-IQ4_XS (a 16B-A1.4B model) achieves 160 tokens per second on a single 8GB VRAM GPU.
  • On CPU-only inference with 32GB of RAM, Ling-mini-2.0-IQ4_XS delivers 50 to 70 tokens per second.

Developers can run fast, low-latency local inference on highly constrained hardware, including 8GB VRAM GPUs and standard CPUs.

SOURCES

6. Moebius: A 0.2B Parameter Image Inpainting Model with 10B-Level Performance

Detailed in arXiv paper 2606.19195, Moebius addresses the capacity drops typically caused by extreme structural compression. By combining LCG with its restructured U-Net, the framework delivers high-quality image inpainting without the massive computational footprint of traditional 10B-parameter diffusion models.

  • Moebius uses a Latent Diffusion Model (LDM) architecture equipped with Latent Categories Guidance (LCG) to achieve high-capacity performance.
  • The framework runs 15x faster than rival 10-billion-parameter models while utilizing only 0.2 billion parameters.
  • An adaptive multi-granularity distillation strategy aligns the lightweight model with a high-capacity teacher to prevent capacity drops.
  • The denoising U-Net is restructured using LλM I blocks to maximize architectural efficiency.

Developers can deploy highly efficient, low-latency image editing and inpainting features directly to edge devices or low-cost cloud instances.

SOURCES

7. Sakana AI Launches Fugu Multi-Agent Orchestration API

Fugu's architecture is based on Sakana's TRINITY and Conductor research papers, hiding the specific model selection and coordination strategies from the user by design. It offers a high-speed version for everyday tasks and the flagship Fugu Ultra for complex tasks like research and cybersecurity. However, early user reports indicate that Fugu Ultra can be slow, with some complex coding tests taking up to 30 minutes to execute.

  • Fugu and Fugu Ultra operate as a single model through an OpenAI-compatible API, managing model selection, delegation, and synthesis.
  • Fugu Ultra is priced at $5 per million input tokens and $30 per million output tokens, with higher rates for workloads exceeding 272K tokens.
  • Fugu Ultra achieved a score of 73.7 on SWE-Bench Pro, outperforming Claude Opus 4.8 but trailing Claude Fable 5.
  • Users can opt out of training data usage and exclude specific models or providers from their routing pool.
  • The service is currently unavailable in the European Union and European Economic Area due to GDPR compliance alignment.

Developers can access frontier-level performance and mitigate vendor lock-in without managing complex multi-model routing infrastructure themselves.

8. Ai2 Releases TMax Terminal Agent Models Up to 27B Parameters

The TMax training recipe utilizes GRPO with stability fixes to train models ranging from 2B to 27B parameters. By leveraging the newly released TMax-15k dataset—which is over 2.5 times larger than the next-largest open terminal dataset—these models deliver strong command-line execution capabilities. The Qwen 3.5-based 9B variant was built using DPPO on the OpenThoughts dataset, leading the TMax ablations.

  • Ai2 released TMax 9B and TMax 27B terminal agent models on Hugging Face.
  • TMax 27B achieves a 42.7% score on Terminal Bench 2.0, approaching the performance of much larger models like Kimi K2.5.
  • The TMax-9B model achieves 27.2% on Terminal Bench 2.0 and 53.0% on Terminal Bench Lite.
  • The models were trained using the TMax-15k dataset, which contains 14,600 reinforcement learning environments.

Developers can integrate highly specialized, open-weights terminal agents into their CLI tools and local automation workflows.

9. Codex Merges Fix for SQLite Logging Bug Causing Excessive SSD Wear

The bug exhibited a 10,000x gap between retained rows and historical inserted row IDs, indicating severe write amplification. Developers running Codex locally should pull the latest updates to apply these logging filters and protect their hardware's lifespan.

  • A global TRACE logging default in Codex's SQLite feedback database caused excessive disk writes, reaching up to 37 TB over 21 days of uptime.
  • The issue stemmed from persisting internal logs, dependency noise, and raw WebSocket payloads, combined with high write amplification from insert-and-prune cycles.
  • On June 22, 2026, two pull requests were merged to filter noisy targets and stop logging every WebSocket event.
  • The merged fixes are expected to reduce Codex's local log volume by 85%.

Developers using Codex must update their installation immediately to prevent severe write amplification and premature SSD failure.

SOURCES

10. OpenAI Expands Access to GPT-5.5-Cyber and Codex Security via 'Patch the Planet'

Building on the May announcement of the Daybreak security agent, OpenAI is now opening access to its Codex Security scanner and GPT-5.5-Cyber model through the 'Patch the Planet' initiative. This program provides open-source maintainers with subsidized security scanning and expert patching support, moving the technology beyond the initial sales-gated preview.

  • OpenAI released its Codex Security scanner as an app plug-in, subsidizing 20 trillion tokens for open-source and private code.
  • The 'Patch the Planet' initiative, partnering with Trail of Bits, HackerOne, and Calif, offers free security consulting and vulnerability patching.
  • GPT-5.5-Cyber scored 85.6% on the CyberGym benchmark, outperforming Anthropic's Mythos 5 model.
  • Program participants receive six months of free ChatGPT Pro and six months of access to the Codex Security scanner.
  • An initial five-day sprint by Trail of Bits engineers uncovered hundreds of bugs and produced dozens of patches across participating projects.

Open-source developers can now leverage subsidized security scanning and expert patching to secure their codebases against emerging AI-driven cyber threats, expanding on the capabilities previously limited to enterprise partners.

SOURCES

11. Security Risks Rise in 'Vibe-Coded' Applications

The ease of generating functional software with AI has led to a surge in applications deployed without proper security audits. While tools like Claude Code can assist in identifying flaws, developers must remain vigilant, perform regular manual security reviews, and consider hiring professional security engineers when handling sensitive business information.

  • A report by Red Access found nearly 2,000 out of 5,000 publicly accessible 'vibe-coded' apps leaking sensitive medical, financial, or strategic data.
  • The AI-agent social network Moltbook exposed tens of thousands of private messages and emails due to an open production database.
  • AI coding tools like Claude Code offer built-in security commands, such as `/security-review`, but require manual invocation by the developer.
  • Security experts advise treating local-only apps as low risk, but warn that transitioning to shared, hosted data requires professional security standards.

Developers using AI coding assistants must actively run security reviews and establish threat models to avoid leaking sensitive user data in production.

SOURCES

12. Self-Harness Framework Lets AI Agents Rewrite Their Own Operating Rules

Developed by the Shanghai Artificial Intelligence Laboratory, Self-Harness is optimized for environments with measurable, deterministic outcomes. By automating the refinement of execution rules, the framework reduces the need for manual prompt engineering, though developers must weigh the performance gains against the increased token costs and latency.

  • Self-Harness uses a three-stage loop of weakness mining, harness proposal, and proposal validation to update agent protocols.
  • The framework achieved relative performance gains of 33% to 60% on Terminal-Bench-2.0 tasks using models like GLM-5 and Qwen-3.5.
  • It allows agents to adapt to model-specific weaknesses without requiring human engineering or stronger external models.
  • The system introduces significant computational overhead, including increased API token usage, latency, and regression testing infrastructure.
  • Researchers advise against using Self-Harness in high-stakes or subjective fields like medicine or safety-critical systems.

Developers can implement self-improving agent loops to boost performance by up to 60% in deterministic environments like coding and DevOps.

SOURCES

13. Oak: A Git Alternative Designed for AI Agents

By eliminating the need for full repository clones, Oak addresses a major bottleneck in agentic workflows where sandboxed environments must quickly spin up and modify code. Although it is in early development and relies on GitHub Actions for its own builds, it represents a shift toward agent-native developer tooling.

  • Oak utilizes virtual mounts to allow agents to work locally or in the cloud without requiring a full copy of a repository.
  • The system enables parallel task execution without the need for full downloads or Git worktrees.
  • The project is in early development, currently lacking features like CI, issue tracking, comments, and Windows support.
  • The project has been fully bootstrapped on its own system without Git backups for several months.

Developers building coding agents can use Oak to enable fast, parallel task execution and reduce disk space overhead in agent sandboxes.

SOURCES

14. Llama.cpp PR Delivers 50% Speedup for Gemma 4 Local Inference

The redundant sorting operation was identified as a bottleneck in specific sampler chains. While the author of the pull request noted some uncertainty regarding how the change might affect other sampler configurations that rely on the legacy behavior, the performance gains for Gemma 4 models are substantial.

  • Pull Request #22645 in llama.cpp removes a redundant unconditional softmax and sort operation from the Top-N-Sigma sampler.
  • The optimization is effective when the Top-N-Sigma sampler is followed by a Dist sampler.
  • Testing on an M3 Max MacBook Pro with a Gemma 4 model showed speed increasing from 30 to 45 tokens per second.
  • The change reduces latency by 10ms per token on the tested hardware.

Developers running local Gemma 4 models can achieve significantly lower latency by adopting the optimized sampler chain.

SOURCES

15. New Techniques Optimize Local Code Generation and GPU Interconnects

These developments target the primary bottlenecks of local model execution. By optimizing software layers—such as tailoring speculative decoding to specific domains and automating low-level GPU kernel tuning—developers can extract enterprise-grade performance from affordable, consumer-grade hardware configurations.

  • Morph LLM achieves a 3.07x speedup in code generation by training a speculative decoding drafter specifically on coding output.
  • Autoresearch automates kernel tuning for low-demand NVIDIA and AMD GPUs, enabling warp-decode kernels to reach 162 tokens per second.
  • A new PCIe interconnect method replaces expensive NVLink setups, using custom kernels to share caches via TCP and reducing time-to-first-token by 84%.

Developers can run local coding models and multi-GPU setups significantly faster on consumer-grade or low-demand hardware without expensive enterprise interconnects.

SOURCES

16. Claude Code Encrypts Local Thinking Logs, Restricting User Access

This design choice ensures that the actual step-by-step reasoning that drives the model's actions remains proprietary. While developers can see the final actions and a summarized logic path, they cannot audit the exact local logs generated during complex agent sessions.

  • Claude Code session logs saved to disk contain thinking blocks with a 600-character encrypted signature.
  • Anthropic holds the encryption keys for these blocks, preventing users from decrypting their local logs.
  • The Claude API and Claude Code's extended-thinking output provide a summary of the model's logic rather than the raw reasoning.
  • Access to the full, unencrypted thinking output requires an enterprise agreement with Anthropic.

Developers analyzing agent behavior must rely on high-level summaries rather than raw reasoning logs, as Anthropic holds the exclusive decryption keys.

SOURCES

17. Apostate Adds Contrastive Co-Vector Operator to Bypass Model Refusals

This approach minimizes KL divergence (achieving 0.081 nats on Granite) to ensure the model's general capabilities remain intact after editing. By targeting the specific refusal component rather than applying broad fine-tuning, developers can steer local models more precisely for specialized tasks.

  • Apostate added a contrastive co-vector edit operator, defined as E = I − R Dᵀ, to isolate and remove refusal directions.
  • The technique preserves harmless variance by fitting a predictor to safe prompts and suppressing it on harmful ones.
  • Testing on the granite-3.3-8b model reduced the refusal rate from 96.0% to 5.0% while maintaining a 95.0% compliance rate.
  • The method is highly effective on architectures with residual or embedding scaling multipliers.

Developers can bypass stubborn safety refusals in open-weights models like Granite without degrading the model's general performance or causing severe divergence.

SOURCES

18. US Export Ban Restricts Foreign Access to Anthropic's Mythos and Fable

The restrictions highlight the growing geopolitical risks associated with relying on single-provider proprietary APIs. As access to top-tier US models becomes politically constrained, international developers are increasingly looking to open-weights models and non-US providers to secure their application pipelines.

  • The US government has barred foreign nationals from accessing Anthropic's latest AI models, Mythos and Fable.
  • An FT analysis found Anthropic discussed risk and regulation at a rate eight times higher than OpenAI in 2026.
  • Critics, including Yann LeCun, attribute the export ban to Anthropic's frequent public warnings regarding AI risks.
  • The ban has raised widespread concerns in Europe and Silicon Valley about future US restrictions on non-US access to frontier models.

Non-US developers face immediate restrictions on building with Anthropic's newest frontier models, forcing them to seek alternative providers or open-weights architectures.

SOURCES

Inference Brew in your inbox

5 minutes a day. Free, unsubscribe anytime.

Inference Brew in your inbox

5 minutes a day. Free, unsubscribe anytime.