Back to benchmarks & research
jev-benchmark (themsquared) project preview
Benchmarks & research

jev-benchmark (themsquared)

Reproducible benchmark for TypeSafe AI's Jev on agent tool-call risk classification: accuracy, latency, and whether the confidence score is worth routing on.

How it uses Jev

📖 Read the write-up: I Benchmarked Jev on Agent Tool-Call Risk. Calibration Held.

Language
Python 100%
License
Apache-2.0
Last activity
Sep 2026
Added
2026-09-21
Stars at snapshot
1
Status
Community listing
agentgatewayai-agentsbenchmarkcalibrationllmmcp

Source: community catalog. Details and metrics reflect the published community snapshot and are not a performance endorsement. Documentation details extracted from GitHub · the README · the project website.