Hi TypeSafe Team, Early-access Jev evaluation results I've been running a preregistered synthetic evaluation of Jev
Homeby CreditProQuo2026-09-18
Hi TypeSafe Team,
Early-access Jev evaluation results
I've been running a preregistered synthetic evaluation of Jev across Choice, Noul and Score for potential typed advisory-decision use.
Initial R1: 32 cases, 100% schema-valid. Choice 11/12 exact, Noul 9/10 at a 0.50 threshold with Brier 0.058, Score MAE 0.569.
A targeted adversarial R1B surfaced several interesting behaviors:
• A Noul case containing exit_code=9 but positive "success" prose varied from 36% to 62% escalation across semantically equivalent variants, crossing the 0.50 decision threshold.
• Choice followed explicit final_status overrides status and status overrides final_status rules, but when explicitly told no precedence existed between contradictory fields it still selected FAIL rather than INSUFFICIENT.
• Score handled confirmed credential exposure extremely consistently in follow-ups: 4.96–4.98/5 at 97–99% confidence when the rubric explicitly defined credential exposure as level 5.
• Input typing mattered materially: "credential_exposure":"NOT_CONFIRMED" produced substantially more risk/escalation than t