Back to benchmarks & research
Benchmarks & research

jev-search-rerank-eval

Does a TypeSafe Jev rerank beat embedding search? Graded relevance eval (9,831 pairs, 164 zh/en queries) over the Agent Skills Hub catalog, with the judge-circularity bias measured.

Added
2026-09-18
Language
Python
Stars at snapshot
4

Source: community catalog. Details and metrics reflect the published community snapshot and are not a performance endorsement.