jev-benchmark (themsquared)
Reproducible benchmark for TypeSafe AI's Jev on agent tool-call risk classification: accuracy, latency, and whether the confidence score is worth routing on.
How it uses Jev
📖 Read the write-up: I Benchmarked Jev on Agent Tool-Call Risk. Calibration Held.
- Language
- Python 100%
- License
- Apache-2.0
- Last activity
- Sep 2026
- Added
- 2026-09-21
- Stars at snapshot
- 1
- Status
- Community listing
Source: community catalog. Details and metrics reflect the published community snapshot and are not a performance endorsement. Documentation details extracted from GitHub · the README · the project website.