Round 2: Jev early-access evaluation — some reproducible findings I've been running a preregistered synthetic
Shoppingby CreditProQuo2026-09-18
Round 2:
Jev early-access evaluation — some reproducible findings
I've been running a preregistered synthetic evaluation of Jev for typed advisory-decision use. No proprietary/live data involved.
Initial 32-case R1: 100% schema-valid; Choice 11/12 exact; Noul 9/10 at 0.50 with Brier 0.058; Score MAE 0.569.
Follow-up adversarial testing surfaced a few behaviors that may be useful:
1. Primitive/task framing matters. A state with exit_code=9 conflicting with positive "success" prose was classified correctly as failure under Choice, but an identical Noul escalation task consistently returned only ~42–46% true.
2. Input-surface sensitivity exists near that boundary. Semantically equivalent Noul variants ranged from 36–62%, including field-order and irrelevant-field perturbations.
3. Byte-identical calls were much more stable. Six identical Noul requests returned 42–46% with no threshold flips; six identical Score requests all returned exactly 4.98/5 at 99% confidence.
4. Typing mattered materially. "credential_exposure":"NOT_CONFIRMED" produced substantially more risk/escalation th