Back to benchmarks & research
Benchmarks & research

jev-ood-calibration

Independent calibration test of TypeSafe's Jev on a task it cannot have seen: 900 rule-generated support tickets (choice / score / boolean) plus 3 public benchmarks via Vercel AI Gateway. Raw responses, ECE with noise floor, temperature refit, per-type sign of miscalibration. Reproducible for ~$0.06.

Added
2026-09-20
Language
Python
Stars at snapshot
2

Source: community catalog. Details and metrics reflect the published community snapshot and are not a performance endorsement.