jev-benchmark (YidiDev)
Rubric-Based Zero-Shot Classification Benchmark: Jev vs Claude Haiku 4.5 vs OpenJev on rubric-conditioned classification, chained decision execution, and exam grading -- with full price tracking.
Rubric-Based Zero-Shot Classification Benchmark: Jev vs Claude Haiku 4.5 vs Claude Sonnet 5 vs OpenJev on rubric-conditioned classification, chained decision execution, and exam grading -- with full price tracking.
How it uses Jev
Jev Benchmark: Rubric-Based Zero-Shot Classification, Chained Execution & Exam Grading
Highlights
- Seeded RNG discipline. Every choice that should be uninfluenced by semantics — opaque folder
- The shuffle control (Part 1) is the load-bearing methodological device of this whole
- Price tracked as it was spent , not estimated afterward — every API call, successful or not,
- Negative results are reported, not hidden. CT7 and CT8 (hard mode) found no weakness for
- Decomposition frozen up front (test-plan.md §8): rubric-clause granularity is a large free
- A comparison arm added after the fact, reasoned about openly. Sonnet was proposed, declined,
- Language
- Python 100%
- License
- MIT
- Last activity
- Sep 2026
- Added
- 2026-09-23
- Stars at snapshot
- 3
- Status
- Community listing
Source: community catalog. Details and metrics reflect the published community snapshot and are not a performance endorsement. Documentation details extracted from GitHub · the README.





