I worked on a harness in the past that used a per-turn evaluation of relevant past context, but daily driving and using it at scale was obviously prohibitively expensive, and "cheap, dumb" LLMs can
Adminby Dr. Fumes2026-09-17
I worked on a harness in the past that used a per-turn evaluation of relevant past context, but daily driving and using it at scale was obviously prohibitively expensive, and "cheap, dumb" LLMs can make a lot of mistakes and have inconsistent performance. That was just one possible harness feature that wasn't very viable. And later, providers started relying on cross-turn data.