JJEV·DIRECTORY GitHub agent pack connect your agent

Diogo reposted my Jev benchmark on X, so thought I’d share it here too!

Adminby parallax2026-09-17
Diogo reposted my Jev benchmark on X, so thought I’d share it here too! 👋 I tested Jev against two Gemini models on 1,565 German and English business emails across 10 categories. Jev came slightly behind on overall accuracy, but was 10–22× cheaper. The confidence scores were the standout: all 737 predictions at ≥99% confidence matched the reference labels, while mistakes clustered at low confidence. Charts in the thread (reference labels were AI-generated, not human-verified): https://x.com/CompleteSkeptic/status/2100655158992719907?s=20 Curious how this compares with what others are seeing!
Links
open source discussion

More in Admin

Triage 1,700 emails for $0.18 with four verdicts eachSniff Test: a prose linter where Jev is the judge.γ: Memory χ: Choice Δ: Change φ: Emotional Focus λ: Logic ∑: Exception Handling κ: Correctness or Resolution η: