Fair question. Short version: quality, not loyalty, and the software is young. Why mostly Anthropic so far We've
Adminby Mks2026-09-16
Fair question. Short version: quality, not loyalty, and the software is young.
Why mostly Anthropic so far
We've only really tested OpenAI and Anthropic. Most of the pipeline was built against Anthropic, and we haven't yet found a candidate that matches it on our content: French course material, structured outputs (JSON schemas for concepts, exams, verdicts), and long context on full lecture PDFs. When you're a two-person team, "doesn't break the schema" is worth more than price.
It's not Sonnet everywhere
The cheap tasks (chat, query rewriting) already run on Haiku, TTS is OpenAI, and SciSpace is coming in soon for essay writing and literature search.
The backend isn't tied to one provider
The per-task routing exists precisely so we can put any model on a single task and measure it. That's why Jev is interesting to us.
So if something cheaper holds up on French course material and typed outputs, it goes in the config. Happy to hear what you'd try.