I feel like devs should have spent a bit more time preparing the model before even starting closed test, jev struggles to understand that the claim is false when it reads like an estabilished one.
Evaluationby thodoh2026-09-17
I feel like devs should have spent a bit more time preparing the model before even starting closed test, jev struggles to understand that the claim is false when it reads like an estabilished one. Struggles to pass a lot of my other benchmarks too - like comparing two numbers or counting how many letters in the word. Either devs didnt run it through the benchmarks or the model is just quite average.