What counts as one “task”?
A task is one job you want done — “triage support tickets” or “extract invoice fields.” You can run a bake-off on any task for free or under your $1 subscription. You only pay the extra $1 when you turn on continuous monitoring for that specific task.
Is the first bake-off really free?
Yes — one full bake-off with the complete report, no card and no sign-up, on your own API keys. The $1 subscription unlocks unlimited runs, saved evaluations, and the weekly leaderboard.
What does the $1-per-task monitoring actually do?
It keeps re-running your evaluation as new models launch, tracks cost and accuracy over time on a per-task dashboard, and alerts you (or auto-switches) the moment a cheaper model clears your bar. That ongoing watch is the part that keeps your bill dropping as the market moves.
Can I stop monitoring a task?
Any time. Turn the watch off and the $1 for that task stops at the end of the cycle — your bake-off history and dashboard stay readable on the $1 subscription.
Do I pay for the models’ API usage too?
Bake-offs run on your own API keys, so model usage is billed by the provider at cost. Grading is programmatic — no expensive LLM-judge — so re-testing during monitoring is close to free.
What about self-hosting and compliance?
Run fully self-hosted so data never leaves your environment, with residency and compliance filters pruning candidates before cost is considered. For >50 monitored tasks or on-prem deployments, talk to us about Enterprise.