Inference Brew

TypeSafe AI Details Jev API Pricing and Functionality

00:00 / --:--

← Back to home

TypeSafe AI Details Jev API Pricing and Functionality

1. TypeSafe AI Details Jev API Pricing and Functionality

TypeSafe AI has provided further details on its Jev 'System One' model, which is now available via a hosted API waitlist. The service is priced at $42 per billion input tokens, with output tokens provided at no cost. The API enables developers to receive typed decisions and confidence scores (0 to 1) rather than text, supporting parallel execution of Choice, Score, and Noul question types against a single state.

  • • Jev API is priced at $42 per billion input tokens, with no cost for output tokens.
  • • The API supports typed decisions with confidence scores (0 to 1).
  • • Developers can execute Choice, Score, and Noul question types in parallel against a single state.
  • • Access remains limited to a waitlist for the hosted API.

Developers can now evaluate the cost and technical integration requirements for incorporating Jev's structured decision-making into their software workflows.

SOURCES

2. Linkup Research Releases SPARSEUP Open-Source Sparse Embedding Model

SPARSEUP is a 149-million-parameter sparse embedding model built on a ModernBERT backbone. Released under the Apache 2.0 license, the model utilizes technical optimizations such as logit shifting and per-position top-k expansion to reduce output dimensions. It is fully compatible with the Transformers and Sentence Transformers libraries for easy integration.

  • • SPARSEUP is an open-source sparse embedding model with 149 million parameters, released under the Apache 2.0 license.
  • • The model is built on a ModernBERT backbone and is compatible with the Transformers and Sentence Transformers libraries.
  • • It achieved an average nDCG@10 score of 56.4 on the BEIR-13 benchmark.
  • • Using the Seismic inverted index on MS MARCO, the model achieves over 97% recall in approximately 380 microseconds per query.
  • • Technical optimizations include a logit shift of 15, a per-position top-k expansion of 12, and case folding to reduce output dimensions.

It gives developers a highly efficient, Apache 2.0-licensed sparse embedding model for building fast, high-recall hybrid search systems.

SOURCES

3. Von Released as an Open-Source, CPU-Friendly Alternative to Jev

Von has been released as an open-source alternative to TypeSafe's Jev model. Designed to run efficiently on standard CPUs with minimal memory requirements, Von offers rapid response times without requiring dedicated GPU hardware. The project is fully open-source and hosted on GitHub and Hugging Face, targeting developers looking for local, structured decision-making capabilities.

  • • Von is a 395M-parameter open-source model hosted on GitHub and Hugging Face.
  • • The model serves as a drop-in replacement for TypeSafe's Jev, running entirely on a CPU with 1–2 GB of memory.
  • • Response times range between 25 and 300 milliseconds, with GPU acceleration optional but not required.
  • • The developer claims Von outperforms Jev across all benchmarks.

It gives developers a lightweight, open-source alternative for structured decision-making that runs locally on a CPU with minimal memory overhead.

SOURCES

4. CUA-S1-FORMS Released for Ultra-Fast Computer Use Decisions

CUA-S1-FORMS is a highly specialized 706k-parameter decision model designed specifically for computer use tasks. Trained on synthetic data in under 30 minutes, the model predicts form actions based on structured inputs without processing screenshots or generating text. The project is fully open-sourced under the MIT license, providing synthetic data generation, training, evaluation, and driver integration code.

  • • CUA-S1-FORMS is an open-source, MIT-licensed model containing 706k parameters with a 2.8 MB checkpoint size.
  • • The model predicts whether to CHECK, CLICK, or SKIP form elements based on structured input without generating text or processing screenshots.
  • • Local scoring takes only 7 to 9 milliseconds, compared to 260 to 280 milliseconds for hosted Jev calls.
  • • In internal evaluations, the model achieved 99.7% accuracy on the full decision set compared to 83.6% for hosted Jev.
  • • The release includes synthetic data generation, training, evaluation, and driver integration code.

It enables agent developers to execute lightning-fast, highly accurate form-filling decisions locally in under 10 milliseconds without relying on heavy LLMs or screenshot processing.

SOURCES

5. TIN Full-Text Search Extension for Postgres Released

TIN (Text INdex) is a GA-released full-text search extension for Postgres and Neki databases. By using native Postgres ctid values as document identifiers, TIN avoids document renumbering during segment merging, which reduces write amplification and CPU costs. The extension utilizes two-level bitmap encoding to enable efficient vectorization on modern CPUs.

  • • TIN is a general availability (GA) full-text search extension for Postgres and Neki databases.
  • • The extension uses Postgres' native ctid values as document identifiers, eliminating the need for separate mapping structures.
  • • It utilizes two-level bitmap encoding to enable efficient CPU vectorization.
  • • Benchmarks show TIN is at least 8x faster than alternatives like GIN, ParadeDB, and pg_textsearch.
  • • It supports boolean expressions, phrase queries, span queries, fuzzy/regex matching, and case/accent folding.

It allows developers to implement ultra-fast full-text search directly inside Postgres without the write amplification or complex mapping structures of external search indexes.

SOURCES

6. Inco Splash Inference Engine Accelerates Local Models on Apple Silicon

Inco Splash is an open-source inference engine optimized for Apple Silicon that significantly outperforms Ollama and oMLX in decode speeds. The engine is particularly efficient when handling multi-agent workflows that fan out into sub-agents. It integrates directly with popular agents and can be configured within the LM Studio Bionic app runtime settings.

  • • Inco Splash is an open-source inference engine built specifically for Apple Silicon.
  • • The engine provides up to 3 times the decode speed of Ollama and 2 times that of oMLX.
  • • Performance increases to nearly 4 times faster when an agent fans out into sub-agents.
  • • System requirements include an M3 chip or newer, macOS 26.4 or later, and 36 GB of memory.
  • • It can be installed via Homebrew and is compatible with agents like Claude Code, OpenCode, Codex, and Hermes.

It significantly accelerates local model testing and agent execution on Mac hardware, reducing latency when running multi-agent workflows.

SOURCES

7. OpenClaw 2026.9.5 Adds Atomic Updates and Hot Reloading to 2.0 Platform

Building on the 2.0 architecture, OpenClaw 2026.9.5 introduces Atomic Updates, which keep the Gateway operational during updates and provide automatic rollback on failure. The release also adds plugin hot reloading, read-only conversation sharing, and an expanded GPT Live feature. Users should note that database migrations in this version require a backup before upgrading.

  • • OpenClaw 2026.9.5 is an open-source, MIT-licensed personal AI agent requiring Node 24.16+ or 26.1+.
  • • The new Atomic Updates feature keeps the Gateway operational during updates and automatically rolls back if the update fails.
  • • The release introduces plugin hot reloading, read-only conversation sharing (Session Share), and guided specialist-agent setup.
  • • Database migrations are modified in this update, requiring a verified backup before upgrading as Atomic Updates do not revert migrations.
  • • A new FreeBSD CLI install route using system Node and npm has been added.

It enhances the stability and developer experience of the 2.0 platform by enabling safer production updates and faster iteration through hot reloading.

SOURCES

Inference Brew in your inbox

5 minutes a day. Free, unsubscribe anytime.

Inference Brew in your inbox

5 minutes a day. Free, unsubscribe anytime.