Back to benchmarks & research
jev-benchmark (YidiDev) project preview
Benchmarks & research

jev-benchmark (YidiDev)

Rubric-Based Zero-Shot Classification Benchmark: Jev vs Claude Haiku 4.5 vs OpenJev on rubric-conditioned classification, chained decision execution, and exam grading -- with full price tracking.

Rubric-Based Zero-Shot Classification Benchmark: Jev vs Claude Haiku 4.5 vs Claude Sonnet 5 vs OpenJev on rubric-conditioned classification, chained decision execution, and exam grading -- with full price tracking.

How it uses Jev

Jev Benchmark: Rubric-Based Zero-Shot Classification, Chained Execution & Exam Grading

Highlights

  • Seeded RNG discipline. Every choice that should be uninfluenced by semantics — opaque folder
  • The shuffle control (Part 1) is the load-bearing methodological device of this whole
  • Price tracked as it was spent , not estimated afterward — every API call, successful or not,
  • Negative results are reported, not hidden. CT7 and CT8 (hard mode) found no weakness for
  • Decomposition frozen up front (test-plan.md §8): rubric-clause granularity is a large free
  • A comparison arm added after the fact, reasoned about openly. Sonnet was proposed, declined,
Language
Python 100%
License
MIT
Last activity
Sep 2026
Added
2026-09-23
Stars at snapshot
3
Status
Community listing

Source: community catalog. Details and metrics reflect the published community snapshot and are not a performance endorsement. Documentation details extracted from GitHub · the README.

Listed in Jev Library

If this is your project, the badge below links back to this page. Paste it near the top of your README.

[![Listed in Jev Library](https://jevlibrary.dev/api/badge/yididev-jev-benchmark)](https://jevlibrary.dev/projects/yididev-jev-benchmark)

Something wrong on this page, or would you rather not be listed? Tell us and it will be changed.