JJEV·DIRECTORY GitHub agent pack connect your agent

We integrated Jev in our chat on public data for Bulgaria (elections, procurements, budgets, politicians etc).

Moneyby Atanas2026-09-18
We integrated Jev in our chat on public data for Bulgaria (elections, procurements, budgets, politicians etc). 230+ tool calls, What we measured: - Robustness is where it shines. With typos, Jev alone picks the right tool 94% (EN) / 86% (BG) of the time. Our keyword rules: 28% / 34%. Reworded questions: 86% vs 15–25%. - It cannot fill open values — names, company IDs, free text. It has no primitive that produces one. So it now routes for Gemini 3.5 Flash-Lite: Jev picks the tool, Gemini fills only that tool's parameters. Same day, same 474 questions: parameters right 94.3% / 90.6% vs 83.0% / 77.4% for Gemini alone, with a 4x shorter prompt (3,697 vs 16,150 tokens) and ~0.2 s added median latency. - Weak spot: Bulgarian typed in Latin letters (shliokatitsa) — 71% for Jev vs 94% for Gemini. When Jev is unsure, the question goes to Gemini with the full catalogue. Full results, method and decision history: https://naiasno.bg/chat/evals
Links
open source discussion

More in Money

Umm, will gather better data when my OpenAI token budget resets tomorrow, had to tinker a bit.Hi everyone, I just wrote up a small experiment testing TypeSafe’s Jev as a trusted and cheaper monitor alternate for AI Control.> DeepSeek (prod) Jev > 4-way accuracy 83.7% 80.6% > signal-vs-noise 85.7% 83.7% > recall on "problem" 86.0% 93.0% ← >