JJEV·DIRECTORY GitHub agent pack connect your agent

This is a lot of fun...my results from the first few tests: The hard limits Where it's excellent (95 to 100% right) -

Eventsby ben2026-09-17
This is a lot of fun...my results from the first few tests: The hard limits Where it's excellent (95 to 100% right) - Logic: reasoning tests using made-up words, so memory can't help. - Reading other languages: Chinese, Arabic, Hindi, Japanese, Swahili and mixed-language slang all scored 100%. - Tricks in the text: it ignored every attempt to force its answer, such as "ignore your instructions". - Long documents: it found a buried fact every time, and used a later correction over the original. - Routing: it picked the right handler out of 255, even when requests were reworded. - Checking claims: it caught "proposed" vs "enacted", "most" vs "all", and "pending approval" vs "approved". - Duplicate news and sarcasm: it could tell when two headlines described the same event, and when a comment was sarcastic. Where it breaks - Arithmetic: it gets simple sums wrong 25% of the time and totals of five numbers wrong 62% of the time. It's often confidently wrong, which makes this the one area where you can't trust it. - Currency conversion, weekdays, and date offsets like "6 weeks after":
open source

More in Events

I catalogued seventeen years of work as 743 skills and asked Jev, once per skill, whether a given job posting callsHello. Building an operating system by-the-way. Aiming to run Windows games eventually.Well I took the whole "Jev is a computer" thing a lot further lol.