Back to benchmarks & research
jev-aita project preview
Benchmarks & research

jev-aita

Benchmark of TypeSafe's Jev against Sonnet 5, GPT-5 nano and local LLMs on 770 Reddit AITA verdicts: Brier scores, latency and cost

How it uses Jev

Jev vs. LLMs on Reddit's "Am I the Asshole?"

Highlights

  • Who is actually right. Only how well each model predicts Reddit's verdict
  • Label-only use. Every model had to give four probabilities, Jev natively
  • Full-size local models. Qwen and Gemma ran as 4-bit builds with reasoning
Language
Python 100%
Last activity
Sep 2026
Added
2026-09-24
Stars at snapshot
2
Status
Community listing
aitabenchmarkcalibrationllmllm-evaluationreddit

Source: community catalog. Details and metrics reflect the published community snapshot and are not a performance endorsement. Documentation details extracted from GitHub · the README.

Listed in Jev Library

If this is your project, the badge below links back to this page. Paste it near the top of your README.

[![Listed in Jev Library](https://jevlibrary.dev/api/badge/dchristopoulos-jev-aita)](https://jevlibrary.dev/projects/dchristopoulos-jev-aita)

Something wrong on this page, or would you rather not be listed? Tell us and it will be changed.