Back to benchmarks & research
Benchmarks & research

jev-agent-failure-benchmark

Benchmarking Jev (Typesafe.ai) against a strong LLM on the Who&When Pro agent-failure-attribution benchmark (text subset).

Added
2026-09-17
Language
Python
Stars at snapshot
1

Source: community catalog. Details and metrics reflect the published community snapshot and are not a performance endorsement.