RLCDAlignBench
Code and benchmark data for a paper evaluating Jev as a zero-shot detector across 44 alignment-failure benchmarks. The authors report a median AUROC of 0.886 for one generic question; results are their research findings, not an independent replication.
RLCDAlignBench: 44 alignment-failure detection benchmarks and code for 'Just Ask Jev' (RLCD zero-shot detector of AI alignment failures)
How it uses Jev
Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures
Source: Project README and arXiv paper. Details and metrics reflect the published community snapshot and are not a performance endorsement. Documentation details extracted from GitHub · the README · the project website.
