WebMCP browser benchmark
Evaluationmeasuredby idan levin2026-09-20
A reported 49-task benchmark combines Jev for tool selection, Mercury 2.5 for arguments and WebMCP for browser actions. The author reports 49/49 tasks solved versus 25/49 for their modified browser-control baseline; results are specific to this setup.
Via DAIR.AI Jev Field Notes (research paraphrase, snapshot 2026-09-20).
Claim · reportedOriginal post by @0xidanlevin.
CaveatAuthor-reported example. Results have not been independently verified.
Links