JJEV·DIRECTORY GitHub agent pack connect your agent

Context compaction evaluation

Evaluationmeasuredby Teknium 🪽2026-09-19
Okay, everyone wants us to give the unbiased facts. Jev compaction here has a lot of problems. The biggest problems: - When I ran it on our public compaction eval (that you can run too), it resulted in a programmatic rule, that you dont need any Jev or other model for - it simply removed all tool calls from the chat history - This means you could do it for free, first of all, but second of all, it means you run into a viscious cycle. Every compaction, there are less and less tool calls to remove from the chat history, meaning it compacts less and less tokens, until there is no room for compaction, and you hit a hard stop, and can't compact anymore. - And, if you are cycling like this, every compaction breaks your cache, so you're paying 10x the price on input tokens, and keeping more input tokens each round means a higher baseline cost after compaction as well. Here's the full reproducible eval that you can run against Hermes Agent between the two (Hermes' vs Jev's): https://t.co/0vJV5u10Cn Via DAIR.AI Jev Field Notes (research paraphrase, snapshot 2026-09-20).
Claim · reportedOriginal post by @Teknium.
CaveatAuthor-reported example. Results have not been independently verified.
Links
open source

More in Evaluation

Return probabilities, not labels — 80/10/10 beats "orange"I’m building Aurapunk, an open-source multi-agent IDE.Tested jev against deepseek v4.1 flash on norwegian text.