Which model leads for one task?
Compare registered models and their exact evaluated settings. This choice leaves your task priorities unchanged.
Resources
AI Model Selector
There are dozens of benchmark sites and even more model subscriptions. This page helps knowledge workers find the right models for their work and their budget.
Rate your work, set a budget and usage, then read the ranking.
Your work
Choose a profession to set task priorities, then adjust any task. Budget, usage and how you work keep their settings.
Based on Kepano's seven-point rating scale. Labels are selector wording.
Budget and use
Agent app: Grok Bot, Codex app and Claude desktop. IDE: VS Code, Cursor and Windsurf (now Devin Desktop). CLI: Claude Code, Codex CLI and Antigravity CLI. Cloud: hosted agents or a separately priced VPS.
Willing to run local models on this device?
Local model evidence and exclusions
Basic keeps these settings. Change them in Advanced. The ranking uses the values shown here.
Tools and data optional
Add APIs, MCP tools, work apps or a VPS. Their monthly costs use the same budget. Model scores stay unchanged. Choose one tier per product; use “Already paid or included elsewhere” only when another account or selected plan covers it.
Choose tools and read their terms
Does the pick change?
| USD/month | Light | Regular | Heavy | Multi-agent |
|---|
A cell shows the top-ranked set: sets that cover demand at the safety margin come first, then higher quality, and within 0.5 quality points the cheaper set. Red cells run out and amber cells cover with less than the margin. Parallel agents set the Multi-agent column. The plan matrix below shows every budget and usage level with the same rules.
Ranked combinations
Performance and monthly cost
Quality score × monthly price, for your current choices
- Efficient profile
- Other qualifying profile
- Measured frontier
Advanced adds
- The full 1–7 rating for each task, and Compare models beside each task
- How you work, Media, Talk to it, Dictate and Data location
- Local models on your own device, with device inputs
- Tools and data add-ons priced inside the budget
- Does the pick change? across budget and usage
- Matrix settings, a bundle tester and supply assumptions
- The full ranked table with plan contents and sessions per month
- Excluded plans and reasons
- How scores work, with a worked example
- Agent Overall evidence
Plan matrix
The best set of plans for each usage level and monthly budget, on your tasks and settings above. Each cell splits the work across up to four plans from different providers. It shows what the set costs, the quality of the work it delivers and how much headroom its usage pools leave. Select a cell to see every combination that was considered.
Matrix settings · quality standard, safety margin, plans per set, free plans, computer use, floors
These settings apply to every section of the page.
Each cell shows the plans, the total price and the quality: the average quality per unit of work, where Opus 5.5 on every task is about 129. Headroom is how many times the set's pools cover the demand at that cell's model mix, or the share served when it runs out. The grey line shows the result when work may drop to lower tiers to get done.
Cell detail
Select a cell in the matrix.
Test a bundle
Tick the plans to buy. The table shows how that set serves each usage level, and which plan and model does the work.
Supply assumptions
Supply is the API-equivalent USD a plan's usage pool holds each month, at the provider's own prices. A plan's work capacity is its supply divided by its model's cost per task relative to Opus. Edit any figure and every section recomputes.
Plan detail
Every solved set for your budget and usage, in ranking order, with the same fields as a matrix cell. Each row lists who does the work, how much of each pool it uses, the sessions a month each pool buys and the plans' contents.
Session costs in
| # | Combination | Monthly cost | Quality | Fit | ≈ Sessions/month | Register |
|---|
Full plan contents
Excluded plans and reasons
How scores work
Read the formula and a worked example for your current choices
Source scores and fixed anchors
Nine tasks use six Arena text boards and Code Arena. Finance and strategy share the business board; coding and systems share the software board. Writing, legal, medical and engineering use writing, legal, medicine and mathematical boards. Visual uses Code Arena (webdev). Duplicate names with no registered binding are removed only when each board retains its original central anchors. On each board, lower_anchor is the minimum central rating and upper_anchor is the maximum central rating across every published model setting. The formula is arena_score = 100 * (rating - lower_anchor) / (upper_anchor - lower_anchor). It preserves rating-gap ratios within that board. Equal ratings have equal scores. A zero-range board has no normalised score and leaves its tasks unmeasured. Its raw evidence remains available.
Each raw interval endpoint uses those same fixed central anchors, without clipping to 0–100. The transformed range is conditional on fixed central anchors. It has no joint confidence level and does not include uncertainty in the anchors. Task quality uses central scores. Raw endpoints, votes, publication date, exact model setting and the complete cohort remain recorded.
Artificial Analysis supplement
Artificial Analysis supplies measured Intelligence Index for orchestrator, AA-Briefcase presentation Elo for visual, Terminal-Bench 4.0 for coding and systems, and the published Finance & Accounting, Strategy & Ops, Legal, Healthcare & Medical and Engineering capability indexes for the matching tasks. Strategy uses Strategy & Ops alone. Estimated or unknown Intelligence Index status is excluded from its scores and cohorts. Each measure uses every eligible snapshot record and aa_score = 100 * (cohort_size - midrank) / (cohort_size - 1). Ties use their mean rank. Fewer than two eligible records leaves the measure unavailable. Explicit per-category bindings can use a different tested effort. AA intervals and benchmark publication dates are unavailable; observation dates remain visible. Conditional interval scenarios hold AA central scores fixed. MMMU-Pro remains visible as reading evidence only. It measures reading images, charts and slides. AA-Briefcase presentation judges produced deliverables in its Stirrup harness; Code Arena measures preference on generated web apps. These measures have limited coverage of native presentation tools. Medical and engineering can bind another published effort of the same model when the default effort lacks the index. The page names that effort. Surveying, geospatial and quantity work have no direct measure.
Arena raw-gap scores and AA percentiles use source-specific units, tasks and cohorts. Scaling both to 0–100 permits the chosen weighted index; it does not calibrate equal ability or accuracy. Raw rating gaps retain meaning within each board only. Changed anchors or cohorts can change the combined result.
Professions and task priorities
The seven profession presets populate the existing task ratings. The page marks these values From profession; an edited task shows Chosen. Unselected tasks mean Not needed (1). Medical and engineering start at 1. Choosing No profession restores the neutral public defaults. Budget, usage and interfaces keep their values. The URL stores the profession, effective task ratings and edit markers. Existing URLs retain their ratings.
Task quality
A model's quality on a task combines both sources and the user feedback: task_quality = mean(aa_score, arena_score) * (1 + praise_complaint). When one source has no measure, the other stands alone. praise_complaint is the model's Agent Arena praise-versus-complaint score on the Overall board, the only bucket captured. Opus 5.5 on most tasks is about 129. Classification and housekeeping have no measure: they leave quality but keep their share of demand.
Each usage level holds a quality floor. A model serves a task only when its quality is at least the floor share of the best model on that task: 0.5 for Light and Regular, 0.8 for Heavy and 0.9 for Multi-agent. "Drop to lower tiers" relaxes every floor to 0.5. Top tier is the share of work on models within 3% of the best on that task. Other work is the lower-tier share.
Demand
Demand is Opus-equivalent USD a month: demand = working_days * agents * ((1 - agentic_share) * chat_sessions * chat_cost + agentic_share * agent_sessions * agent_cost). Working days are 21.7. Light is 3 chat sessions of 2–9 turns and 2 agent sessions of 10–49 turns a day. Regular is 4 chat and 10 agent sessions of 10–49 turns. Heavy is 3 sessions of 50+ turns of each kind. Multi-agent is Heavy per agent, 3 agents by default, and demand scales with the agent count. Coding, systems and orchestration are agentic; other tasks use the chat sessions. A session costs Opus 5.5's measured OpenRouter Claude Code median for its length. The Multi-agent, 5+ agents column of the plan matrix uses a calibrated demand: 8% above the larger of Claude Max 20x and a measured set that ran out.
Supply
Claude Max 5x holds min(window_usd * windows * working_days, weekly_cap_usd * 52 / 12): the lower of USD 110 per 5-hour window times 2 windows a working day, and a USD 656 weekly cap times 52/12, which is USD 2,843 a month. Claude Pro holds one fifth of that and Max 20x four times it. Other plans use the register's sourced supply. A plan with no published size is estimated in work units: a paid plan scales its provider's measured plan by price, and a free plan gets 0.1 of the cheapest paid plan's work, or 20 units when the provider has no paid plan. A model's cost ratio is its Artificial Analysis cost per task divided by Opus 5.5's, and work_capacity = supply / cost_ratio(top_model).
Allocation
Every set of up to four plans from different providers within the budget is screened. Tools and data and dictation costs come out of the budget first. A plan joins a set only when it includes one of the selected ways of working, and every selected way must be included by a plan in the set. Subscription plans need recorded included access (R30); conditional, separately billed and unverified routes stay out. Images, video, two-way voice and computer use, when required, keep only sets with such a plan. Local models count no supply and stay evidence only.
For the best 150 sets, and for every set of one or two plans, a linear program first serves as much demand as the pools allow at the safety margin, then maximises quality. A plan that does no work is dropped from its set. A set's quality is the sum of its task contributions: task_contribution = share * served_quality / (demand * margin). share is the task's rating divided by the sum of measured ratings, served_quality is the sum of work times quality on that task, and demand is the task's share of the column's demand. Unserved work adds nothing.
Fit and ranking
Headroom is how many times the set's pools cover the demand at its model mix: 2× or more covers, 1–2× is tight, and under 1× runs out and shows the share served. Free tiers throttle sustained use, so a free-only set qualifies at Light only; above Light it shows red and does not count. A free plan can still join a paid set. Sets that cover demand at the safety margin rank first, then higher quality. Within 0.5 quality points the cheaper set ranks first. The recommendation is the top ranked set.
Value lens
The model comparison orders each task by quality, by Artificial Analysis cost per task, or by value: value = 0.8 * 100 * quality / best_quality + 0.2 * cost_score, where cost_score = 100 * (ln max_cost - ln cost) / (ln max_cost - ln min_cost) across the compared models with a cost per task.
Current worked example
Agent Overall evidence
Agent Arena signals. Values are percentage-point effects versus the reference. Praise versus complaint enters quality: a model's task quality is multiplied by one plus this score. Overall and steerability are shown for reference only. Each Agent setting is bound independently of the text and AA settings. Hosted records are proxies for local quantisations.
Compare Overall, praise versus complaint and steerability
| Registered model | Exact Agent setting | Overall agent/overall | Praise versus complaint agent_praise_complaint/overall | Steerability agent_steerability/overall |
|---|