Diogo reposted my Jev benchmark on X, so thought I’d share it here too!
Adminby parallax2026-09-17
Diogo reposted my Jev benchmark on X, so thought I’d share it here too! 👋
I tested Jev against two Gemini models on 1,565 German and English business emails across 10 categories.
Jev came slightly behind on overall accuracy, but was 10–22× cheaper. The confidence scores were the standout: all 737 predictions at ≥99% confidence matched the reference labels, while mistakes clustered at low confidence.
Charts in the thread (reference labels were AI-generated, not human-verified):
https://x.com/CompleteSkeptic/status/2100655158992719907?s=20
Curious how this compares with what others are seeing!
Links