Barnaby Robson

Resources

AI Model Selector

There are dozens of benchmark sites and even more model subscriptions. This page helps knowledge workers find the right models for their work and their budget.

Rate your work, set a budget and usage, then read the ranking.

01 / 05

Your work

What do you do?

Choose a profession to set task priorities, then adjust any task. Budget, usage and how you work keep their settings.

Tasks

Based on Kepano's seven-point rating scale. Labels are selector wording.

02 / 05

Budget and use

Budget / month

Usage

Set for you

Basic keeps these settings. Change them in Advanced. The ranking uses the values shown here.

03 / 05

Ranked combinations

    Performance and monthly cost

    Quality score × monthly price, for your current choices

    • Efficient profile
    • Other qualifying profile
    • Measured frontier

    Advanced adds

    • The full 1–7 rating for each task, and Compare models beside each task
    • How you work, Media, Talk to it, Dictate and Data location
    • Local models on your own device, with device inputs
    • Tools and data add-ons priced inside the budget
    • Does the pick change? across budget and usage
    • Matrix settings, a bundle tester and supply assumptions
    • The full ranked table with plan contents and sessions per month
    • Excluded plans and reasons
    • How scores work, with a worked example
    • Agent Overall evidence
    04 / 05

    Plan matrix

    The best set of plans for each usage level and monthly budget, on your tasks and settings above. Each cell splits the work across up to four plans from different providers. It shows what the set costs, the quality of the work it delivers and how much headroom its usage pools leave. Select a cell to see every combination that was considered.

    Runs out
    Covers demand, but with less headroom than the safety margin
    Covers with the safety margin
    2.0× or more covers comfortably
    1.0–2.0× tight: covers, but heavy weeks may run out
    serves 79% runs out
    Top tier share of work on models within 3% of the best on that task
    Lower tier: Sonnet 9% work moved to a cheaper model to fit the pools

    Each cell shows the plans, the total price and the quality: the average quality per unit of work, where Opus 5.5 on every task is about 129. Headroom is how many times the set's pools cover the demand at that cell's model mix, or the share served when it runs out. The grey line shows the result when work may drop to lower tiers to get done.

    Cell detail

    Select a cell in the matrix.

    Provider logo sources and licences