SemIf
Semantic ifs from open models, on a 3090 at home. Independent; not affiliated with Jev or TypeSafe.
Benchmarks, evals, calibration studies, and open replicas. Browse the collection in Jev Library.
129 entries in the current community snapshotSemantic ifs from open models, on a 3090 at home. Independent; not affiliated with Jev or TypeSafe.
Personal-assistant agent built on Vercel's eve with 100 mocked tools, measuring how many steps it takes when Jev picks the tool versus the LLM.
This is a LLM Gateway that mimics typesafe ai structured output. Like an imposter Jev.
Unofficial study: Jev-style parallel typed decisions on stock 1.5B-8B models on an Apple Silicon laptop. Benchmarks, research notes, and a Hugging Face Space demo.
Typed JSON inference with DiffusionGemma, with Every and Jev benchmark results
openvons (open-Jev): 有限選択肢に確率で答える判断層 — テキスト / 画像 / 日本語音声コマンド
Probability-aware evaluation for typed decision models: calibration, selective risk, latency, and reproducible benchmarks.
One-pass option scoring with a local Gemma 3 4B on Apple silicon via MLX, inspired by jevlike, with a Doom demo.
A word-level language model whose output layer is Jev: n-gram drafter, Noul chunk verification, bits-per-token eval
One-pass typed decisions with calibrated probabilities (System One style model), fine-tuned from Qwen3.5-2B
LegalForecast-MTD benchmark alpha and official evaluation workflows
Reproducible early-access evaluation of Jev on Korean understanding and medical text, with runtime and cost evidence
Independent Jev 1.13.0 behavior study: report, controlled prompt experiments, raw results, and offline verification.
JEV-inspired parallel decisions for CUDA LLMs. One context, many decisions. vLLM API, game-agent examples, and reproducible benchmarks.
Local bilingual probability decisions from context, questions, and candidate answers. Independent research preview inspired by TypeSafe Jev.
An experimental JEV-powered framework for forecasting short-term stock price direction from structured market data.
Eight minimal working examples of TypeSafe's Jev (a System One model) applied to mechanical and electrical engineering: CAD/CAE/CAM routing, FEM result triage, DFM screening, BOM alignment, hallucination-proof extraction. Zero dependencies.
Backtest Jev (TypeSafe) as a BUY/SELL/HOLD trader on NQ L10 order-book data
A show-and-tell capability study for Jev, TypeSafe's System One decision model.
AI benchmark on Japan's 2026 Common Test: Jev vs luna-none vs luna-low (static dashboard)
Open replica of TypeSafe's Jev: typed calibrated decisions in one forward pass, on Gemma 4 E2B / Gemma 3 270M (Modal)
A Jev-inspired decision interface for existing LLMs. Explicit choices, scores, calibration, and review thresholds.
Typed-decision benchmark from PadFlow (land development SaaS): schemas, anonymized labeled rows, and a runner for confidence-calibrated models like TypeSafe Jev.
Discriminative Monte Carlo Tree Search using TypeSafe Jev System One Primitives and Gemini
Can a decision model beat dedicated rerankers? TypeSafe Jev vs Cohere Rerank 4 vs ZeroEntropy zerank-2 vs a chat-model baseline: 14 datasets, every raw API response, bootstrap ranges on every gap.
Benchmarks and a playground for TypeSafe's Jev (System One) model: chess, and who-is-the-player-talking-to for speech-to-text game NPCs
Blind security benchmarks for Jev, TypeSafe's System One model: prompt injection and vulnerable code detection, built on jev-go
Benchmarks Jev against other evaluation models in games with explicit states, legal actions, and measurable outcomes.
Using Jev as an evaluator.
Benchmarking Jev (Typesafe.ai) against a strong LLM on the Who&When Pro agent-failure-attribution benchmark (text subset).
发明 RLHF 的人,这次做了个不会说话的模型:Jev 独立研究报告。52 页 PDF + 50 条中文实测复现包 + 143 条可回溯数据表
Jev-style parallel constrained decisions for any MLX model on Apple Silicon. Typed, schema-valid JSON in one forward pass.
Reproducible Jev Ultrafast research-browser eval harness + field note (QC’d cases, suite runner, report generator). Not investment advice.
Jev (TypeSafe) exploratory thread: claim audit, live demos, and runnable code
Measures how well TypeSafe's RLCD-Jev model spots real secret credentials in file snippets
Jev (TypeSafe) vs Claude Haiku 4.5 on 2 000 phishing emails: accuracy, calibration, latency, cost. Reproducible benchmark.
Open-source Jev-style System One decision model. Gemma 3 270M with a scoring head — fast, calibrated decisions in a single forward pass. No text generation. Inspired by TypeSafe.ai's Jev.
Benchmarking TypeSafe's Jev decision model as a cost-efficient LLM router on RouterArena
Jev (TypeSafe) 性能評価プロジェクト — 日本郵便 KEN_ALL をマスタに、AI SDK 経由の Jev が住所のあいまい一致にどこまで使えるかを検証
Calibration and confidence-based routing measured on Banking77: 80.2% accuracy at $0.103 per 500 decisions
Experiments with openjev, an open Jev-style option-logit runner, on local models.
Evaluating TypeSafe's Jev as a fast monitor and action gate for agent sabotage in SHADE-Arena, compared with Gemini 2.5 Flash/Pro.
Charts: TypeSafe Jev evaluated on Thai standardized exams vs 110 other models
Jev vs Gemini 3.8 Flash: labelling 1,000 app reviews, 4.1× faster and 7× cheaper
Jev-style calibrated decision model (Choice/Score/Noul) on Qwen3.5-0.8B
Application of TypeSafe Jev (noul judgment primitive) on the collusion.wiki corpus: agent vs human page authorship, head-to-head vs local Qwen3.8-Flash-Next
A chatbot built on a model that cannot generate text (TypeSafe AI's Jev, driven autoregressively)
A small second eval for shadcn-ui/lint that uses TypeSafe's Jev to judge the linter's own output.
Position paper: the Hidden-Markov and fuzzy primitives missing from TypeSafe AI's Jev and System-One decision models. Two lemmas, one principle (Deferred Crispification), one architecture (BSF-S1).
I tortured Jev into being a RISC-V CPU.
typesafe.ai model jev finance benchmark
A small reproducible MuJoCo pilot comparing Jev, Claude Haiku, and reactive rules for pick-and-place.
Zero-shot spam filtering with TypeSafe Jev Noul questions, compared with TF-IDF baselines
A chatbot from typed Jev decisions: hierarchical speculative decoding over System One probabilities.
Jev (TypeSafe) vs. Gemini 3.8 Flash vs. GPT-5.6 Luna na anotação estruturada de sentenças do TJSP: qualidade, tempo e custo
Jev (TypeSafe System One) × ASReview SYNERGY abstract screening demo — Choice/Noul vs gold labels
TypeScript experiments, evaluations, and latency benchmarks for TypeSafe's Jev model
Can Jev pick the winner of a real headline A/B test? 64.5% across 10,984 Upworthy randomized experiments, 74.7% when the difference was decisive.
Does Jev predict stock returns from news? It reads the news well; there is no tradeable alpha. Three arms separate reading from recall.
Using Jev to test how well it predicts financial markets(just like most llms as of september 2026, it doesnt do that good)
An observable raw-character chat experiment powered entirely by TypeSafe Jev Choice
An evaluation of typesafe AI chess. As it turns out, the AI isn't doing really well even though chess is not a particularly open-ended game. Still, it's only a prototype and this probably wasn't optimzied for games.
Batched single-token choice inference for open language models, compatible with TypeSafe
Evaluating TypeSafe's System One primitives (Choice/Score/Noul) — where a typed oracle beats an LLM call
Rust port of TypeSafe system-one-adapter (LLM-backed system_one evaluations)
On-device iPhone visual decision tool using MLX and Qwen3-VL direct option logits.
Hugging Face Space exploring open-source parallel constrained decoding as an alternative to Jev.
Ongoing Japanese research deck on Jev and System One models, maintained as Markdown slides.
Type-safe one-decision-per-token decoding engine for autoregressive LLMs, inspired by Jev.
Train a small model that chooses among a changing list of text options, one probability per option in a single pass. Includes Doom, chess, and Wikispeedia demos.
A stronger one-pass scorer over a variable list of text options: hashed n-gram encoder, rival-aware attention, gated head, temperature scaling, benchmarked against jevlike.
An educational Jev-like visual inference experiment on Apple Silicon: shared context, direct candidate scoring, and local visual demos.
A small open decision model: state + typed questions -> calibrated probabilities. A Jev / System One re-creation on Qwen3.5.
mini-Jev: what a Jev-style typed-decision interface looks like on a frozen Qwen3-4B — read the option letter's logits instead of generating JSON. Preregistered experiment, results, teaching bench.
Open, Jev-compatible System One decision server on DiffusionGemma
看看 Jev 能做什么:用中英文讲清热门应用、工作原理和各自优缺点。Explore Jev apps with plain-language examples, explanations, and comparisons.
Inspired by TypeSafe Ai, Ask a local LLM typed questions, get calibrated probabilities instead of text. Structured output without generation or parsing. MLX / Apple Silicon.
daf-jev: composable Python toolkit for TypeSafe's Jev (System One) decision API — question builders, confidence gates, evaluator, calibration, CLI, MCP server, agent skill
Does a TypeSafe Jev rerank beat embedding search? Graded relevance eval (9,831 pairs, 164 zh/en queries) over the Agent Skills Hub catalog, with the judge-circularity bias measured.
Non-autoregressive decision engine on ModernBERT (151M) with calibrated uncertainty (RLCD), TypeSafe AI Jev benchmark audit, and in-browser WebGPU playground
Decision harness for TypeSafe Jev — confidence gates, shadow mode, recipes, and evals. Claude CLI 48.9s → Jev 1.3s on the same row-filter job.
Stop guessing confidence thresholds: calibrate, threshold, and drift-check typed decision models (TypeSafe Jev) against an LLM teacher.
Not every coding task needs your best model. Experimental Jev-powered model routing for Claude Code — V3 prototype runs today, V4 routes at the task boundary.
Open alternative to Jev: typed, calibrated decisions from any open-weights LLM in one forward pass (HF + vLLM), with benchmarks
High-throughput synthetic & pretraining dataset sifter powered by TypeSafe AI Jev (api.typesafe.ai). Stream, filter, and score Parquet & JSONL datasets at 1,500+ rows/sec using System One typed decisions (Choice, Score, Noul).
An agent skill to discover TypeSafe Jev opportunities, design typed questions, and learn from recent community experiments.
A reproduction of Jev that turns any Qwen model into a fast decision model, serving the same /v1/systemone schema (Choice, Score, Noul) with no training and no generated answer text.
Replicating Jev with a local LLM
Independent, evidence-based map of when TypeSafe's Jev actually holds up vs. breaks down — real API-call receipts, not a leaderboard. 中文為主的雙語 repo。
The open-source System One decision model. Sub-15ms, non-autoregressive, local drop-in alternative to TypeSafe Jev.
Nushell module for the TypeSafe System One API: typed decisions with calibrated probabilities
Fast, cheap judgment for AI coding agents: semantic search, focused reads and list picking in ~2s. CLI + MCP server on TypeSafe Jev. Benchmarked on SWE-bench.
I kept watching coding agents burn context on decisions that aren't hard - triage 400 tickets, tag 600 files, route to one of six teams. jev-mode moves those verdicts to a typed-judgment model. I A/B'd it: 78% fewer tokens, 16x less work-attributable input, accuracy 96.1% vs 93.7%. Python, no deps, MIT.
Autonomous System-One Triage Engine & Benchmark powered by TypeSafe AI (Jev). 75ms inference, $0 output tokens, and RLCD epistemic safety gates.
MCP server and agent skill for the TypeSafe AI System One API (Jev): decompose a judgment into Choice / Score / Noul questions, lint them, measure on labelled data, and put calibrated thresholds in code
TypeSafe AI System One (Jev) task plugin for QuantumNous/new-api — native /v1/systemone, synchronous evaluation, token billing
TypeSafe Jev demonstration for new analyzation — experimenting with Jev for fast analysis of news and tickers
Claude Code plugin that scores review findings, debug hypotheses and design options with TypeSafe's Jev — calibrated probabilities instead of one more opinion.
Calibrated alignment verifier for LLM responses and agent plans — powered by Jev
What your last session knew, scored against what this one is doing. MCP server: a per-project ledger written as things happen, recalled per task with TypeSafe's Jev evaluation model via Vercel AI Gateway.
Chess moves, evaluations, persona opponents, and game classification with TypeSafe AI System One
Curated catalog of System One / Decision Models — contributions for modelsystem.one
TypeSafe (Jev) vs DeepSeek-flash: side-by-side speed/token/cost/accuracy comparison across invoice extraction, email classification, and reranking
Calibrated 151M Non-Autoregressive Decision Engine beating TypeSafe Jev & Laya on LocalLLaMA/typed-decisions (77.10% acc, 0.0636 Brier, 0.0144 ECE)
JevBench v1 - a benchmark for Jev-class typed decision models: smart, cheap, fast, reliable, open.
if you're experimenting with jev it will be easier from here
sort by meaning: order lines along a plain-English dimension, from pairwise comparisons judged by TypeSafe's Jev model
Reproducible benchmark for measuring Jev reranking quality, latency, and cost in RAG
Find, design, and evaluate TypeSafe Jev decision loops.
Agent skill: design judgment-assisted systems with TypeSafe Jev (System One). Maps Choice/Score/Noul onto decision theory, reranking, and routing. Composition algebra, question design, validation gates. MIT.
Command-line tool for TypeSafe AI's Jev model. Ask yes/no, multiple-choice and rubric questions about any text and get calibrated probabilities back. Answers become exit codes for shells and CI, JSON for scripts, and MCP tools for AI agents.
Unofficial TypeSafe Jev showcase — System One decisions, not chat.
Semantic test matchers for Vitest and Jest, powered by TypeSafe's Jev model. Write expectations in plain English, get calibrated probabilities back.
OpenJev: an independent Jev-inspired System One decision API based on TypeSafe.ai concepts. Choice, score and noul primitives, local mock server, Python and TypeScript SDKs. Real inference planned; not affiliated with TypeSafe AI.
Reward-hack radar for coding agents: structural denies + TypeSafe Jev System One sidecar for Claude Code & Cursor hooks
Catch breaking API behavior hidden in OpenAPI prose with deterministic checks and TypeSafe JEV System One semantic review.
FUn little experiment with Typesafe AI Jev Model playing chess against stockfish :)
Open auto mode for AI agents — a calibrated tool-call firewall powered by TypeSafe Jev. Ships as a Claude Code hook
Experimental semantic line search with TypeSafe Jev via OpenRouter. Python CLI with no runtime dependencies.
AI life-and-civilization simulation: TypeSafe Jev makes every decision (typed, probabilistic, auditable); LLMs plan — OpenAI-compatible APIs, local models (Ollama, LM Studio), Claude Code, Codex.
Independent calibration test of TypeSafe's Jev on a task it cannot have seen: 900 rule-generated support tickets (choice / score / boolean) plus 3 public benchmarks via Vercel AI Gateway. Raw responses, ECE with noise floor, temperature refit, per-type sign of miscalibration. Reproducible for ~$0.06.
Audit your git diff against YAML coding-standards packs using TypeSafe's Jev model, from a CLI or your AI agent's command/skill.
Typed, policy-driven decision workflows on top of TypeSafe AI Jev: confidence routing, fallbacks, evaluation, and RAG patterns for TypeScript apps.
Jev-compatible System 开源Jev
Jev, TypeSafe's System One model, plays chess against any OpenRouter LLM, Stockfish and you. One-page web app with live moves, Jev's move probabilities, saved games and win rates.
Never confidently wrong: a TLA+-verified consensus kernel around TypeSafe's Jev, run through 1,680 chaos-tested pharmacy decisions with zero wrong verdicts. Film, code, and every captured call.
Open-source BYOK arena for Jev and other AI judges. Find failures, compare quality, cost, and latency.
Jev 1.13 reward-model evaluation across 8 benchmark tracks, with an interactive report and 54-row SOTA comparison
See what Jev thinks about your SaaS website — powered by ReplyNodes web context and Vercel AI Gateway.