I want to use Jev as a high-throughput security triage layer for bug bounty / vuln research — scoring thousands of code paths and hypotheses for attacker control, trust-boundary crossing, duplicate
Evaluationby no meaning2026-09-17
I want to use Jev as a high-throughput security triage layer for bug bounty / vuln research — scoring thousands of code paths and hypotheses for attacker control, trust-boundary crossing, duplicate risk, evidence quality, and whether they’re worth escalating to the more classic and expensive LLM reasoning agents like Claude or Daybreak Blue
I already have a small eval set from real security research with PROMOTE / HOLD / KILL outcomes, so I’d love to benchmark Jev on whether it can rank the same cases correctly.