02 / AI SYSTEMS · MULTI-OBJECTIVE SEARCH
Foundry.
Compare AI routing options against a quality target, latency limit, and budget.
Inspect a strategy
This policy meets your current limits. Nondominated means no other eligible policy improves one of the three objectives without worsening another.
The evidence behind the points
No named model gets an invented benchmark here. The starter table uses fictional profiles. Quality is an observed task pass rate when you supply your own rows.
| Task / profile | Pass | Mean | $/1k | n |
|---|---|---|---|---|
| triage / local | 91% | 180 ms | 0 | 100 |
| triage / balanced | 95% | 480 ms | 1.2 | 100 |
| triage / deep | 97% | 2100 ms | 5.8 | 100 |
| code / local | 65% | 420 ms | 0 | 100 |
| code / balanced | 87% | 1100 ms | 1.2 | 100 |
| code / deep | 96% | 4200 ms | 5.8 | 100 |
| review / local | 55% | 380 ms | 0 | 100 |
| review / balanced | 80% | 1400 ms | 1.2 | 100 |
| review / deep | 94% | 5200 ms | 5.8 | 100 |
| secrets / local | 84% | 210 ms | 0 | 100 |
| secrets / balanced | 92% | 600 ms | 1.2 | 100 |
| secrets / deep | 96% | 2400 ms | 5.8 | 100 |
Use your evaluation results
Paste the same schema with a row for every task/profile pair. Set source to user-supplied. Costs, latency and pass rates should come from comparable held-out evaluations.
Smarter allocation, with visible assumptions
The search enumerates all category-to-profile assignments, then compares cost, expected pass rate, and mean latency. It does not train a router, call paid models, predict tail latency, or establish statistical confidence. A zero API price for a local profile excludes hardware and electricity. Routing research such as RouteLLM motivates the tradeoff; the results shown here come only from this table. Re-evaluate on fresh tasks before using a policy in production.
The person behind the project
A note from Luis.
My coding-agent evaluation work makes me interested in what an AI system can demonstrate, including where it fails. This notebook exposes the quality, cost, and latency assumptions. Its starting model profiles are hypothetical, not a benchmark I have measured.
Who it helps
Engineers comparing the tradeoffs of AI routing policies.
Try this
Change the quality target or budget and inspect why a routing option becomes infeasible.
Inspect a reproducible system handoff