JJEV·DIRECTORY GitHub agent pack connect your agent

Jev + Mercury broke the WebMCP benchmark

Evaluationby @hackgoofer2026-09-18
This is the result of not benchmaxxing Quote idan levin @0xidanlevin · Sep 18 We just ran Jev on our WebMCP benchmark. The result: basically broke the benchmark. Jev + Mercury 2.5 (a fast, low-cost LLM) using WebMCP solved 100% of the tasks at roughly 112× lower model cost than GPT-6 Astra using computer use with code execution. Compared to Astra using…
open source

More in Evaluation

Return probabilities, not labels — 80/10/10 beats "orange"I’m building Aurapunk, an open-source multi-agent IDE.Tested jev against deepseek v4.1 flash on norwegian text.