IT Pro Expert
Search
IT · 11 Sep 2026 · 67 min read

Which AI model to use? (Sept 2026)

Confused about which AI model to choose in September 2026? Compare strengths, use cases and performance to pick the right fit for your business.

Which ai for what task sept 2026

Every AI lab says its newest model is the best one. Most of them are right about something. This page sets out, category by category, which of the September 2026 models actually leads on medical, coding, legal, cybersecurity, reasoning, images, video and more — with the published scores, the price, and the caveats the launch posts leave out.

The short answer

Best all-round

Claude Opus 5.5, released 22 September, scores 58 on the Artificial Analysis Intelligence Index v4.3 — five points clear of the next models, the widest lead at the top in months. It leads six of the ten tests in the index and costs $4 in / $20 out per million tokens, 20% less than Opus 5.

Close behind

Claude Fable 5.1 and GPT-6 Astra, tied at 53. Astra still leads maths and uses the fewest tokens of any frontier model; Fable 5.1 costs twice Opus 5.5's per-token price. Sonnet 5, now permanently $2 / $10, covers routine work.

Fastest and cheapest

Gemini 3.8 Flash at around 262 tokens per second and $0.75 / $3.75 per million until 31 December, when the price doubles. DeepSeek V4.1 Flash undercuts it at $0.15 / $0.60 off-peak.

Best to self-host

Xiaomi's MiMo-V2.6-Pro (46 on the index, MIT licence, released 21 September) is now the strongest open model, though at 1.02 trillion parameters it needs datacentre-class hardware such as a GB300 station. DeepSeek V4.1 Flash remains the practical pick for code and agents, and GLM-5.3 (45) for general work. Your data never leaves your own servers.

No single model wins everything. The right answer depends on the job.

The models compared

These are the current generally available flagships from each lab, plus the leading open-weight models. Claude Opus 5.5 is the first model in Anthropic's Claude 5.5 family; Anthropic says Sonnet 5.5 and Haiku 5.5 will follow in the coming weeks. Claude Mythos 5.1 is deliberately absent: it is the same underlying model as Fable 5.1 with some safeguards lifted, and it is available only to vetted organisations in Anthropic's trusted-access programmes, not on any public plan or the standard API.

ModelMakerReleasedContextAPI price ($ per million tokens, in / out)
Claude Opus 5.5Anthropic22 Sep 20261M$4 / $20 (cache reads $0.20)
Claude Fable 5.1Anthropic1 Sep 20261M$10 / $50
Claude Opus 5Anthropic24 Jul 20261M$5 / $25
Claude Sonnet 5Anthropic30 Jun 20261M$2 / $10 (planned rise to $3 / $15 cancelled)
GPT-6 AstraOpenAI3 Sep 20261.05M$10 / $50
GPT-5.6 SolOpenAI9 Jul 20261.05M$4 / $20 promotional to at least 21 Nov (list $5 / $30)
GPT-5.6 TerraOpenAI9 Jul 20261.05M$2 / $12
GPT-5.6 LunaOpenAI9 Jul 20261.05M$0.20 / $1.20
Gemini 3.8 FlashGoogle2 Sep 20261M$0.75 / $3.75 (doubles to $1.50 / $7.50 on 1 Jan 2027)
Gemini 3.1 ProGoogleFeb 20261M$2 / $12 (prompts up to 200K)
Grok 4.7SpaceXAI (xAI)21 Sep 2026500K$2 / $6
Grok 4.6SpaceXAI (xAI)12 Aug 2026500K$2 / $6
Muse Spark 1.3Meta2 Sep 20261M$1.25 / $4.25 ($0.10 / $0.20 if Meta may train on your data)
MiMo-V2.6-ProXiaomi21 Sep 20261M$0.435 / $0.87 (open weights, MIT)
Kimi K3Moonshot AI16 Jul 20261M$3 / $15 (open weights, custom licence)
Qwen3.8-MaxAlibaba3 Aug 2026~1M$2 / $6 (open weights, custom licence)
GLM-5.3Z.aiAug 20261MLow; open weights
Step 5 PreviewStepFun20 Sep 2026 (API)1M$1.00 / $2.70 (weights due 15 Oct)
DeepSeek V4.1 FlashDeepSeek10 Sep 20261M$0.30 / $1.20 peak, half off-peak (open weights, MIT)
DeepSeek V4 ProDeepSeekAug 20261MOpen weights (MIT); peak / off-peak API pricing
Hy4 previewTencent28 Aug 20261M$0.83 / $2.50 (open weights, Apache 2.0)
Mistral Medium 3.5Mistral (France)Apr 2026256KLow; not fully published

Category rankings

Category leader Second and third The rest I independent   V vendor-reported

Reasoning and overall intelligence

The Artificial Analysis Intelligence Index v4.3 combines ten tests, including long-horizon knowledge work, business workflow automation, terminal tasks, document reasoning and factual accuracy, with 45% of the weighting on private test sets so vendors cannot train on them. The current ceiling is 58. Bars here are scaled to the leader.

1Claude Opus 5.558I
2Claude Fable 5.153I
2GPT-6 Astra53I
4Claude Opus 551I
5Claude Fable 550I
6Muse Spark 1.348I
7GPT-5.6 Sol47I
8Grok 4.746I
8MiMo-V2.6-Pro open weights46I
10GLM-5.345I
10Qwen3.8-Max API45I
12Grok 4.644I
12Kimi K344I
12Step 5 Preview44I
15GLM-5.3-Flash42I
15GPT-5.6 Terra42I
17Gemini 3.8 Flash41I
18Qwen3.8 2.4T open weights40I
19DeepSeek V4.1 Flash39I
20GPT-5.6 Luna37I
21DeepSeek V4 Pro36I

If you have seen index scores of 57, 61 or 66 quoted for these models, those are older scales from earlier in September — Kimi K3, for example, launched at 57 on version 4.1, and Opus 5 at 61. Only v4.3 numbers are comparable with each other; the figures here are from the current v4.3.2 revision. Opus 5.5 was run at max effort with Anthropic's fallback models switched on, and trails on only three of the ten tests: CritPt, AA-LCR and GDP.pdf. Grok 4.7 was benchmarked on 21 September, the day it launched.

Science knowledge

GPQA Diamond is a set of graduate-level physics, chemistry and biology questions written so that experts outside the field cannot answer them with Google. The frontier is now bunched within six points of the ceiling.

1GPT-6 Astra96.0%V
2Gemini 3.8 Flash95.3%V
3GPT-5.6 Sol94.6%V
4Gemini 3.1 Pro94.3%I
5Claude Fable 5.193.7%V
6Kimi K393.5%I
7Claude Opus 593.2%V
8Qwen3.8-Max92.6%I
9GPT-5.6 Luna92.3%I
10DeepSeek V4 Pro90.1%I

Anthropic did not publish a GPQA Diamond score for Claude Opus 5.5 at launch.

Maths

FrontierMath Tier 4 is research-level mathematics. Note that these are OpenAI-reported figures and that OpenAI funded FrontierMath and had sight of many of its problems, so treat the gap as directional rather than settled.

1GPT-6 Astra97.6%V
2Claude Fable 5.187.8%V
2Claude Fable 587.8%V
4GPT-5.6 Sol~83%V
5Claude Opus 573.2%V

Coding

The fairest single answer to "can this thing do engineering work" is now Artificial Analysis's Coding Agent Index. Each model runs inside its own maker's coding agent — Claude Code, Codex, Grok Build, Muse Code — on the same three test sets (DeepSWE, Terminal-Bench 4.0 and SWE-Atlas-QnA), scored independently.

1Claude Fable 5.1 Claude Code62I
1GPT-6 Astra Codex62I
3Claude Opus 5 Claude Code60I
4Grok 4.7 Grok Build56I
5GPT-5.6 Sol Codex55I
6Muse Spark 1.3 Muse Code54I

Astra ties Fable 5.1 at roughly 60% of the cost per task, because it uses fewer tokens than any other agent on the index. Grok 4.7 gained nine points on Grok 4.6 to take fourth place, overtaking GPT-5.6 Sol. Claude Opus 5.5 had not yet been run through the Coding Agent Index at the time of writing. On Anthropic's own launch figures (vendor-reported), it scores 57.8% on CursorBench 4.0 against 51.8% for Fable 5.1 and 41.7% for GPT-5.6 Sol, and 54.4% on FrontierCode against 53.3% for GPT-6 Astra.

The older shell test is now saturated

Terminal-Bench 2.1, run by Artificial Analysis on the same harness for every model, was the index's shell test until 7 September. Ten models now score 84% or more, so it no longer separates the leaders, but it is still the widest independent comparison that includes the cheaper models.

1Claude Fable 5.191.4%I
2GPT-6 Astra89.9%I
3GPT-5.6 Sol89.5%I
4Claude Opus 589.1%I
5Grok 4.688.4%I
6GPT-5.6 Terra88.0%I
7Gemini 3.8 Flash87.6%I
8Kimi K385.0%I
9GPT-5.6 Luna84.7%V
10GLM-5.3-Flash84.3%I
11DeepSeek V4 Pro78.7%I

The harder tests tell a different story

DeepSWE 1.1 puts a model inside a real repository and asks it to fix real issues. Terminal-Bench 4.0 is the much harder shell test that replaced 2.1 in the index. On these the order changes and the field spreads out.

ModelDeepSWE 1.1Terminal-Bench 4.0 (vendor)Terminal-Bench 4.0 (independent)
Claude Opus 5.5—66.4% V59.6% I
Muse Spark 1.375.4% V——
GPT-6 Astra74.1% V57.7% V59.1% I
Gemini 3.8 Flash73.8% V19.1% V—
Claude Opus 573.6% V—49.0% I
GPT-5.6 Sol72.7% V37.3% V39.9% I
Claude Fable 569.9% V——
Claude Fable 5.167.4% V55.8% V52.0% I

Opus 5.5's vendor figure is at its xhigh effort setting, with a stated margin of error of about 2.6 points; on Artificial Analysis's independent run it is level with GPT-6 Astra and 11 points ahead of Opus 5. DeepSeek V4 Pro's 80.6% on the older SWE-bench Verified remains the best published open-weight result on that test. Grok 4.7 scored 73% on DeepSWE and 33% on Terminal-Bench 4.0 when Artificial Analysis ran it inside its own Grok Build agent on launch day.

What about PHP, C# and SQL?

Every headline coding score above is Python-biased: SWE-bench Verified is 500 tasks drawn from twelve open-source Python repositories, and DeepSWE and Terminal-Bench lean the same way. Nobody publishes a PHP or C# leaderboard. Two datasets get close, and both point the same way.

TestWhy it matters for PHP / C#LeaderScoreNext best
SWE-bench MultilingualReal GitHub issues across nine languages including PHP (C# is not among them). Anthropic is the only frontier lab that publishes a result.Claude Opus 4.677.8% VNo published GPT, Gemini or Grok figure. Open-weight Qwen3.6-35B: 67.2%
Scale SEAL private commercial repositoriesUnseen business codebases that look nothing like popular open-source projects — the closest proxy to a real PHP or .NET application.Claude Opus 4.6 (thinking)47.1% IMuse Spark 44.7%, GPT-5.4 xHigh 43.4%, Gemini 3.1 Pro 32.2%

Opus 4.6 dropped only three points moving from Python-only Verified (80.8%) to Multilingual (77.8%), and on the private commercial set Claude degrades less than its competitors when moving from public to unseen repositories. Both figures are from an earlier Claude generation, because the current ones have not been re-run on these tests; the pattern is the point.

Desktop and computer use

OSWorld 2.0 gives a model a real desktop and 108 long workflows that take a human around 1.6 hours each. The headline numbers are partial-credit; the strict "finished the whole job" rate is far lower and shown alongside.

1Claude Opus 5.581.8%V
2Claude Fable 5.177.9%V
3Claude Opus 575.4%V
4GPT-6 Astra72.6%V
5GPT-5.6 Sol65.7%V
6Gemini 3.8 Flash59.0%V

Strict end-to-end completion: Fable 5.1 41.7%, Opus 5 39.6%; no strict figure has been published for Opus 5.5. Anthropic's launch table, run on its current setup, puts Opus 5.5 at 81.8% against 80.7% for Fable 5.1 and 74.0% for Opus 5, so read its lead over Fable as about a point. Fable scores zero on any task where its safety classifiers intervene, which drags its average down on security-adjacent work. Astra completes tasks around 47% faster than GPT-5.6 Sol, so it often wins on wall-clock time despite the lower score.

Global knowledge

Humanity's Last Exam is 2,500 expert-written questions across every academic field. With tools enabled, Claude leads and this is the one academic test GPT-6 Astra loses.

1Claude Opus 5.567.7%V
2Claude Fable 5.165.0%V
3Claude Fable 563.8%V
4Claude Opus 563.6%V
5GPT-6 Astra57.2%V
6Gemini 3.8 Flash54.9%V

Artificial Analysis's independent run, on its own setup, puts Opus 5.5 at 61.4% against a previous best of 59.1% from Fable 5.1, so the ordering holds. Gemini's figure is on the HLE-Verified variant, so treat it as approximate. Factual accuracy specifically is covered in the hallucination section below.

Biology and life sciences

This is the thinnest category. Only OpenAI publishes head-to-head life-science scores, Anthropic routes biology questions on Fable to Opus 5, and the strongest biology model on either side (Claude Mythos 5.1) is not on public sale: access is limited to vetted organisations in Anthropic's trusted-access programmes. Anthropic says the new Opus 5.5 matches or beats Mythos 5.1 across many areas of biology, which is why it ships with the same biology safeguards as Fable 5.1.

ModelGeneBench ProLifeSciBenchNote
GPT-6 Astra37.8% V60.3% VLeads both published tests
GPT-5.6 Sol28.7% V59.9% VNear-level on LifeSciBench
Claude Opus 5.5not publishednot publishedFable 5.1's biology safeguards; vetted labs and companies can apply to Anthropic's new Life Sciences Verification Program for fuller access
Claude Fable 5.1 / Opus 5not publishednot publishedBiology requests on Fable are routed to Opus 5
Grok 4.6not publishednot publishedWon a biosecurity benchmark no other model matched; no percentage released
Claude Mythos 5.1restrictedrestrictedVirology evaluations of 0.81–0.87 reported; vetted organisations only

Legal (professional)

Legal benchmarks are the most fragmented, and each one is led by a different model. The consistent finding across all of them: models pass around 90% of the individual criteria on a legal task but complete very few whole tasks end to end.

BenchmarkWhat it measuresLeaderScore
BigLaw BenchDrafting and analysis to law-firm standardClaude (Opus 4.8 reference)91.1% I
Harvey Bench / GDPvalProfessional legal work productGrok 4.6rank only, no % published
Legal Research Bench (Vals)Case-law research accuracyGPT-5.6 Sol48.1% I
HAQQLegal question answeringDeepSeek V4 Pro36.8 / 50 I
Harvey LABEnd-to-end agentic legal tasksMuse Spark 1.120.0% I

Reasonable working rule: Claude for drafting and precision, GPT for research, and a human reads everything before it goes out. None of these models should be trusted to complete legal work unsupervised.

Cybersecurity

Two tests matter. ExploitBench measures whether a model can turn a known vulnerability into a working exploit; CyberGym measures whether it can find vulnerabilities in real codebases. A model can be excellent at one and mediocre at the other.

1GPT-6 Astra100%V
2Claude Fable 5~78%I
3GPT-5.6 Sol76.5%I
4Claude Opus 570%V
5GLM-5.354.4%V

ExploitBench above. On CyberGym (vulnerability discovery) the order flips: DeepSeek V4.1 Flash 88.1%, Gemini 3.8 Flash cyber variant 86.2%, GLM-5.3 84.5%; Xiaomi reports 94.0% for MiMo-V2.6-Pro on its own setup, but only 47.9% on ExploitBench. Claude Fable 5.1 has no published cyber score because penetration-testing, exploit and binary-scanning requests are routed to Opus 4.8. Claude Opus 5.5, which Anthropic rates as comparable to Mythos 5.1 in cybersecurity, has no published score for the same reason: it is the first Opus to ship with Fable-class cyber safeguards. Gemini's cyber variant is only available inside Google's Fairwind programme, and Astra's public build declines advanced offensive work unless your organisation is enrolled in OpenAI's Daybreak scheme.

Tool use and agents

AutomationBench-AA is Artificial Analysis's version of Zapier's business-workflow test: multi-step tasks across real apps. The headline score gives partial credit; the strict figure counts only workflows finished without breaking a single rule.

1GPT-6 Astra68.5%I
2Grok 4.666.7%I
3GLM-5.362.2%I
4GPT-5.6 Sol60%I

Artificial Analysis says Claude Opus 5.5 is level with GPT-6 Astra on AutomationBench-AA, but its headline score was not in the launch summary, so it is not charted here. Strict, no-rule-broken completion: Astra 41.6%, Fable 5.1 32.1%, Opus 5 28.3%; Opus 5.5 scored 40.0% in Zapier's own strict run during early access, as published by Anthropic. On AA-Briefcase, the multi-week knowledge-work test, the order reverses: Opus 5.5 now leads at 1,822 Elo, 143 points ahead of Fable 5.1 and the first Anthropic model to beat GPT-5.6 Sol on presentation quality, with Fable 5.1 and Opus 5 behind it and Grok 4.7 joining the frontier just behind them, even though Grok 4.7 slipped 1.1 points against Grok 4.6 on AutomationBench. Grok 4.6 remains the value pick for automation: near the top at a fraction of Astra's price.

Medical and health

Three different questions get asked under "medical", and they have three different answers.

Professional use: answering to a physician's standard

HealthBench Professional is built from 5,000 conversations and physician-written rubrics from 262 doctors in 60 countries, scored on the length-adjusted version so that longer answers do not automatically win. It is OpenAI's benchmark, and all figures are OpenAI-run.

1GPT-6 Astra63.4%V
2Claude Fable 560.9%V
3GPT-5.6 Sol60.5%V
4Claude Opus 557.5%V
5Claude Fable 5.156.6%V

Fable 5.1 scoring below Fable 5 is a safety-routing artefact, not a capability regression: questions its classifiers flag are answered by a smaller model and marked down accordingly. No Gemini, Grok, open-weight or Claude Opus 5.5 score on HealthBench Professional has been published yet.

General medical information for the public

No consumer-grade health benchmark has been published for this generation of models, so the professional ranking above is the best available guide, read together with the hallucination table further down. The practical change this month is on the refusal side: Anthropic reports that Fable 5.1's biology safeguards fire 85% less often on benign elementary biology and medical questions than Fable 5's did, which means fewer "I can't help with that" answers to ordinary health queries.

Scan and imaging interpretation

The imaging benchmarks exist, but they live in academic papers rather than on vendor leaderboards: Radiology's Last Exam (RadLE) tests chatbots against board-certified radiologists on hard spot-diagnosis cases, ReXVQA covers chest X-ray question answering, CXR-LT is a multi-centre chest X-ray challenge with over 145,000 images, and newer sets such as NeuroQA cover 3D brain MRI. No lab publishes scores on any of them, so the only figures come from independent studies — which run a generation behind the current models.

Modality (RadLE)RadiologistsBest chatbotOthers
MRI98%GPT-5, 45%Gemini 2.5 Pro 35%, o3 33%, Grok-4 23%, Claude Opus 4.1 0%
X-ray89%GPT-5, 31%Gemini 2.5 Pro 22%, o3 22%, Grok-4 8%, Claude Opus 4.1 3%
CT79%Gemini 2.5 Pro, 29%GPT-5 22%, o3 19%, Grok-4 8%, Claude Opus 4.1 1%

Other studies agree on the ordering. On pneumothorax from chest radiographs, ChatGPT-4o was most accurate at 69.6%, then Claude 3.5 at 64.9% and Gemini 2.0 at 57.4% — but every model fell to between 12% and 21% on children, and ChatGPT dropped from 81.6% on large pneumothoraces to 42.2% on small ones. On measuring liver metastases against a radiologist's reference, Gemini reached an agreement score of 0.81, GPT-o3 0.52 and Claude 4 Opus 0.07.

Claude's pattern is consistent: it is the weakest of the majors at reading pixels and the strongest at reading radiology text. On structuring 3,949 head CT reports, Claude was significantly more accurate than GPT and Gemini for intracranial haemorrhage, and on 56 JAMA neuroradiology cases Claude 3.5 achieved the highest accuracy (80.4%) when given the image and the report together.

Explaining a scan

GPT first, Gemini second. Fine for describing findings in plain English to a patient. Not close to a radiologist: the best result on any modality was 45% against 98%.

Working with reports

Claude. Structuring, summarising and extracting findings from radiology text is where it leads, even though it trails badly on the images themselves.

Actually analysing scans

A specialist model, not a chatbot. Google's MedGemma is purpose-built for medicine, reads X-ray, CT and MRI volumes natively, runs locally, and in the ReXVQA reader study scored 83.84% — above every human reader in the panel.

Clinical use

The regulated products deployed in NHS radiology — chest X-ray triage, stroke CT, fracture detection — are UKCA/CE-marked medical devices with published sensitivity per condition. No general chatbot is one.

Two caveats. Every study above tested general models one generation or more behind the current ones — GPT-6 Astra, Claude Opus 5.5, Claude Fable 5.1 and Gemini 3.8 Flash have not yet been run through RadLE. And the RadLE cases are deliberately hard: on routine chest X-rays the chatbots do markedly better, as the pneumothorax study shows, and MedGemma does better still.

Speed

Two different things get called speed. Tokens per second is how fast text appears on screen; time per task is how long an agent takes to finish a job, which depends as much on how many tokens a model needs as on how fast it produces them.

ModelOutput speed (tokens/sec)Time to first tokenToken efficiency
Gemini 3.8 Flash262–2653.0s medianFast but verbose
DeepSeek V4.1 Flash207.51.15sFast but very verbose: 250M tokens to run the full index, against a 140M median
Grok 4.7~188—About 81k output tokens and seven minutes per index task, three times Astra's token count
MiMo-V2.6-Pro~134—$0.13 per index task, one of the cheapest results at its level; an UltraSpeed variant costs ten times as much
GPT-5.6 Terra~103lowFastest GPT-5.6 tier
GPT-6 Astra71.410s+Most efficient at the frontier: about 27k output tokens per index task
Claude Fable 5.169.710.8s (p95)Token-hungry: 1.7× the output of Fable 5
Claude Opus 5.5not yet measured—Anthropic reports output over 30% faster than Opus 5. About 119k output tokens per index task at max, against 73k for Opus 5, but lower prices keep its cost per task level with Opus 5
Muse Spark 1.3—5.9s (p95)About 60k output tokens per index task, roughly three times version 1.2
Kimi K3——Averages nearly an hour per task on AA-Briefcase

If you are running agents, token efficiency is the number that sets your bill. Astra is not the fastest per token, but it needs about a third of the output tokens Grok 4.7 uses per task, and under a quarter of Opus 5.5's at max effort, which is why it has the lowest cost per task at its level on the index. Opus 5.5 narrows that gap at lower effort settings: Artificial Analysis finds four of its five effort levels on the best-value frontier among models scoring 50 or more.

Hallucination and accuracy

AA-Omniscience asks 6,000 hard factual questions and tracks two things separately: how often the model is right, and how often it makes something up when it does not know. A model can post a low hallucination rate simply by refusing to answer, so read the two columns together.

ModelAccuracy (higher is better)Hallucination rate (lower is better)Read as
Claude Fable 5.167%not in the published snapshotHighest accuracy in the published snapshot; answers more, so hallucination will not be low
Claude Fable 565%63.6%Very accurate, confidently wrong when it misses
Claude Opus 5~61%60.8%First Opus that answers rather than declines
GPT-6 Astra~55%51%Halved Sol's hallucination rate without refusing more
Gemini 3.1 Pro54.9%50.9%Balanced
Grok 4.747%29%Lowest measured rate of this month's frontier releases, down from 34% on Grok 4.6
Gemini 3.8 Flash—55.2%—
Kimi K346–48%~53%Regressed from K2.6
Claude Opus 4.846.6%35.9%Best-calibrated previous-generation flagship: knows what it doesn't know
Grok 4.3 (medium)—16%Lowest rate of any flagship-class model
GPT-5.6 Sol~51%92.2%Almost never admits ignorance
DeepSeek V4 Flash (July build)37%84%Strong coder, still overconfident on facts, though 11 points better than the April build

On the combined AA-Omniscience Index, which rewards right answers and penalises invented ones, GPT-6 Astra (44, at high effort) sat one point ahead of Fable 5.1 (43) before Opus 5.5 arrived. Artificial Analysis now lists Claude Opus 5.5 as the leader on AA-Omniscience, but its separate accuracy and hallucination figures were not in the launch summary. The very lowest hallucination rates on the board belong to small models that decline most questions. That is not a model you want; it is a model that has learned to say "I don't know". Also worth knowing: reasoning modes hallucinate two to three times more than plain modes on summarisation tests. If the facts matter, turn thinking off or give the model a document to work from.

Refusals and over-blocking

This is the category businesses ask about most and the one with the least data. No leaderboard measures "refused a perfectly reasonable request" across vendors, so the table below collects what has actually been published.

ModelWhat is knownFigure
Claude Opus 5.5The first Opus to launch with Fable 5.1-class safeguards on cybersecurity, biology and distillation, each falling back to another model transparently: most cyber tasks go to Opus 4.8. Fixing bugs in your own code is still allowed, and vetted organisations can apply for fuller access through Anthropic's cyber and life-sciences verification programmes.qualitative
Claude Fable 5.1Anthropic says its new safeguards block 60% fewer false positives than Fable 5, with around 60% fewer cyber interventions per session, and biology safeguards firing 85% less often on benign medical questions. Pentest, exploit and binary-scan requests still go to Opus 4.8.−60% / −85% vs Fable 5
Claude Fable 5Fell back to Opus on roughly 18% of AutomationBench-AA tasks and 9% of AA-Omniscience questions — the only absolute block rate any vendor's routing has produced.9–18% of tasks
Older Claude (4.6)OR-Bench over-refusal on a benign health-robotics set: Opus 4.6 33.5%, Sonnet 4.6 25.9%. Claude has historically been the safest and the most over-cautious.26–34%
GPT-6 AstraPublic build declines advanced offensive cyber work; wider access through the Daybreak programme for vetted organisations. Alignment testing: 0% unauthorised task completion.qualitative
GeminiCyber variant restricted to the Fairwind programme. Older Gemini 2.5 Flash showed 46.3% over-refusal on the same health set.46% (older gen)
MistralAccepts most prompts on OR-Bench; consistently the lowest over-refusal of the majors.low
Open-weight modelsIndependent bio-research testing found benign-tier over-refusal of 91.5% for Kimi K2.6, 76.6% for Claude Opus 4.7 and 57.9% for GPT-5.5 — and that a model's overall refusal rate is a poor predictor of how sensibly it refuses.58–92% (bio prompts)

A model that refuses to look at malware is no use to the people whose job is looking at malware.

Image generation

Image models are ranked by blind human votes in the Artificial Analysis arena: two images from the same prompt, pick the better one, thousands of times. Elo is the native unit; higher is better and 30 points is a meaningful gap. OpenAI's September release, GPT Image 2.5, now holds the top two places both for creating images and for editing them.

ModelMakerText-to-image EloImage editingBest atPrice per 1,000 images
GPT Image 2.5 FlareOpenAI1188 (#1)#2The quicker 2.5 tier; photorealism and text in images$211
GPT Image 2.5 SunburstOpenAI1182 (#2)#1The slower, more detailed tier; the best editor on the board$211
GPT Image 2OpenAI1171 (#3)#5Plans the layout before drawing; strong typography$211
Grok Imagine Image 2.0SpaceXAI1154 (#4)outside top 5New in August; top-five quality at under a third of OpenAI's price$60
MAI-Image-2.6Microsoft1147 (#5)#3Editing, product and branded design$39
Reve 2.1Reve1129 (#6)outside top 5Previously led image editing; now overtaken$200
Nano Banana 2 (Gemini 3.1 Flash Image)Google1122 (#7)outside top 5Fast, low-cost everyday image model$67
Muse ImageMeta1111 (#8)—Cheapest model in the top ten$10
Nano Banana Pro (Gemini 3 Pro Image)Google1100 (#11)outside top 5True 4K output, multilingual text, identity lock across edits$134
Ideogram 4.0 (Quality)Ideogram1012 (top open-weight)—Typography and text-heavy graphics$100
FLUX.2 [dev]Black Forest Labs1000—Runs locally; best open-weight option for photorealism$12
Midjourney V8.1Midjourneynot on arena—Aesthetic and creative ideationsubscription

Prices are Artificial Analysis's figure for 1,000 images at 1024×1024 on each maker's API. For diagrams, technical illustrations and anything with words in it, no benchmark measures "diagram quality" directly; text-rendering ability is the proxy, and GPT Image 2 and 2.5, Ideogram 4.0 and Nano Banana Pro lead it. For architecture diagrams you will usually get a better result asking a language model for SVG or Mermaid than asking an image model to draw one.

Video generation

Video models are ranked by blind votes in the Artificial Analysis Video Arena, and the board has turned over since the start of September. Google's Gemini Omni Flash now leads for clips with sound; the rest of the top five are built on Chinese models from Alibaba, MiniMax and ByteDance. Kling, which topped several leaderboards earlier in the year, has slipped to twelfth.

ModelMakerText-to-video Elo (with audio)Image-to-video Elo (with audio)Output and strengthsAPI price per minute
Gemini Omni FlashGoogle1233 (#1)1177 (#3)Conversational editing and clip extension; 1080p and 4K by upscaling$6.00
Wan 3.0Alibaba1229 (#2)1164Up to 30-second clips; also leads text-to-video without audio$12.00
MiniMax H3 Maxfal, built on MiniMax H31227 (#3)1195 (#1)Top-three quality at the lowest price in the top ten$2.40
MiniMax H3MiniMax1220 (#4)1181 (#2)Best open-weight video model$7.80
Seedance 2.0ByteDance1210 (#5)1174 (#5)Native audio, multi-shot, up to 12 reference files$9.07
Seedance 2.5ByteDancenot yet rankednot yet ranked30 seconds per pass, extendable; up to 50 references; timestamp and region editsHiggsfield: ~$3.50 per 10s at 720p
HappyHorse 1.1Alibaba114711067-language lip-sync$9.90
Grok Imagine Video 1.5SpaceXAInot ranked1099Tuned for animating stills$8.40
Kling 3.0 Pro (1080p)Kling AI (Kuaishou)1095 (#12)1055Native 4K, 15-second multi-shot scenes, multilingual lip-sync, motion control$20.16
Veo 3.1Google10881082Up to 4K with synchronised dialogue; cheaper Fast ($9.00) and Lite ($4.80) tiers$24.00
Runway Gen-4.5Runwaynot on current board—4K; best motion brushes and scene control for film worksubscription / API
Sora 2OpenAIwithdrawnwithdrawnApp closed 26 April; API ends 24 September 2026—

Elo scores are for clips with audio, as of 22 September. Without audio the order changes: Wan 3.0 leads text-to-video (1336) and Gemini Omni Flash leads image-to-video (1369). Prices are Artificial Analysis's figure for one minute of video on the maker's own API, except Seedance 2.5, which uses Higgsfield's credit price.

Seedance 2.5 and Higgsfield

ByteDance released Seedance 2.5 on 31 July. It generates up to 30 seconds with sound in a single pass, double Seedance 2.0, and can extend a clip in further rounds while keeping the same characters and setting. One generation can draw on up to 30 images, 10 video clips and 10 audio clips as references, and edits can target a single moment or region instead of re-rendering the whole clip.

It launched in China on Jimeng and Doubao, and outside China it is sold through creative platforms including Higgsfield, Dreamina, Runway, OpenArt and Magnific. Higgsfield, a San Francisco company that bundles more than 50 image, video and audio models under one subscription, is the version most people are talking about, because it wraps the model in direct controls for era, lens, lighting, physics and camera path. It charges 70 credits (about $3.50) for a 10-second clip at 720p and 120 credits (about $6) at 1080p; the model generates up to 1080p, with 4K by upscaling.

There is no blind-vote score yet, and ByteDance itself says complex motion and scenes with several interacting people still need work. Seedance 2.0 also drew legal complaints from Disney and Paramount over famous characters, so keep real people, brands and film characters out of your prompts.

Kling: capable, but no longer on top

Kling 3.0, launched in February, is still Kuaishou's flagship, with a faster Turbo version and an upgraded Omni model added on 17 June. Its strengths are unchanged: native 4K, up to 15 seconds across several shots, multilingual lip-sync and precise motion control. On the blind vote, though, the 1080p Pro model now sits twelfth for text-to-video with sound, and at $20.16 a minute on Kling's own API it costs more than all but one of the models ranked above it.

The business is growing fast. Kling AI was spun out of Kuaishou in July with a $3 billion funding round at an $18 billion valuation, backed by investors including Tencent, Alibaba Cloud and Baidu, and around three-quarters of its revenue comes from outside China. A Kling 4.0 is widely expected but has not been announced, so treat any site already selling "Kling 4" access with suspicion.

Seedance and Kling both come from Chinese companies, so prompts and uploads may fall under Chinese law, just as material sent to a US provider falls under US law. For client footage or unreleased products, check which company processes the data, and where, before you upload anything.

What people actually run

Everything above ranks models by how they score. This ranks them by how much work they are actually given. OpenRouter is a neutral exchange that routes requests across several hundred models from every major lab and publishes its own token volumes, so it is the closest thing to a usage leaderboard that exists. Here is where the tokens went in the seven days to 20 September 2026.

Most-used model Second and third The rest I measured by an independent third party, not a score
1DeepSeek V4.1 Flash DeepSeek15.8TI
2GLM 5.3 Flash Z.ai14.1TI
3Hy4 preview Tencent12.5TI
4GPT-5.6 Luna OpenAI9.72TI
5DeepSeek V4 Flash DeepSeek9.44TI
6MiMo-V2.5 Xiaomi7.07TI
7Hy3 Tencent4.78TI
8Nemotron 3 Ultra (free) Nvidia4.49TI
9DeepSeek V4 Flash, second listing DeepSeek3.77TI
10GLM 5.3 Z.ai3.0TI

Volumes are prompt plus completion tokens for the seven days to 20 September 2026, so they predate the MiMo-V2.6 and Claude Opus 5.5 launches. OpenRouter lists variants of the same model separately, which is why DeepSeek V4 Flash appears twice. Source: OpenRouter (openrouter.ai/rankings), licensed under CC BY 4.0.

Read that against every ranking above and the mismatch is close to total. Only one model this page names as a category leader appears: DeepSeek V4.1 Flash, top of the CyberGym vulnerability-discovery test, which went from seventh place to first within ten days of its release. Claude Fable 5.1, GPT-6 Astra and Claude Opus 5 are nowhere near the top ten. Eight of the ten entries are Chinese; the two American ones are OpenAI's budget tier and a free Nvidia model. Counted by requests rather than tokens, the week beginning 14 September split DeepSeek 25.4%, Google 18.6%, OpenAI 17.0% and Z.ai 9.4%, with Anthropic on 2.7%.

Benchmarks measure the ceiling. Usage measures the floor, and almost all the work happens on the floor.

The pattern behind it is straightforward once stated. Almost every model in the top ten is open-weight, free or priced near the floor, and an exchange like this accumulates whatever is cheapest to serve. The frontier labs' pricing keeps their best models off it by design. That is worth knowing when you are choosing a stack: the models that win the tables above are bought for the hardest tenth of the work, and something like this list is what handles the rest.

What the work actually is

OpenRouter also classifies requests by task and reports the share of spend each takes. In the week to 12 September the split was a useful corrective to a page organised around medical, legal and scientific benchmarks.

CategoryShare of spendWhat sits inside it
General31.8%Classification, question answering, content writing, conversation, customer support, summarising
Code30.8%Code generation, debugging, file operations, shell execution, code review, frontend and UI
Agent28.2%Workflow execution, multi-step planning, tool dispatch
Data9.2%Data extraction and transformation

Agent work is already more than a quarter of what people pay for, on a par with writing code. None of the professional benchmarks on this page measure it, which is why the tool-use and computer-use sections matter more than their position in the page suggests. The categories this page covers in most depth, medical and legal, do not appear in the spend breakdown at all.

One reason any of this can be wrong

For two weeks in August 2026 an anonymous listing called Ox Alpha carried 27 trillion tokens through the exchange, 12% to 14% of all traffic, with no lab attached to it. Labs routinely run unreleased models under codenames for a week or two before launch; Z.ai later confirmed that Ox Alpha was GLM-5.3-Flash and published the weights. So a usage table like the one above can be materially wrong about who is winning, for weeks at a time, while being exactly right about the totals.

Why everyone is talking about Kimi

Kimi is the assistant and model family from Moonshot AI, a Beijing start-up founded in 2023. Three things have put it in the headlines since July: a model that rivals the US frontier and can be downloaded, a business growing unusually fast, and a public accusation from a rival.

2.8Tparameters in Kimi K3, the largest open-weight model yet released
44on the Artificial Analysis index, two behind the top open model, Xiaomi's MiMo-V2.6-Pro
300bntokens a day from K3 models on OpenRouter
$2bnannualised revenue Moonshot is targeting by the end of 2026

A frontier-class model you can download

Moonshot released Kimi K3 on 16 July and published the full weights on 27 July: 2.8 trillion parameters with 104 billion active per token, a one-million-token context window and native vision. At launch it took first place in the Frontend Code Arena, ahead of Claude Fable 5 and GPT-5.6 Sol, although Moonshot itself conceded that K3 still trailed those two models overall. The release landed in Washington as a question about how quickly a Chinese start-up had closed the gap. A US congressional committee had been pressing American companies over their use of Kimi since April, and in March Cursor admitted that its Composer 2 coding model was built on an earlier Kimi release.

It is not a bargain model. Moonshot charges $3 per million input tokens and $15 per million output, five times the input price of its predecessor, K2.6, although K3 is free to use in the Kimi app. The weights come under Moonshot's own licence rather than MIT: free for most businesses, but model-hosting companies with more than $20 million in annual revenue need a separate agreement.

A fast-growing business

Bloomberg reports that Moonshot is aiming for $2 billion in annualised revenue by the end of the year, double its August run rate, and OpenRouter shows K3 models generating up to 300 billion tokens a day. On 17 September Moonshot launched Kimi for financial services, wired into data from S&P Global Market Intelligence, Crunchbase, the US SEC's EDGAR filings, the IMF, the World Bank and the Federal Reserve's FRED database. It has reportedly filed confidentially for a Hong Kong listing.

The allegation

On 10 September Anthropic, which makes Claude, published a threat-intelligence report accusing Moonshot of more than copying. It alleged that in one ten-day window Moonshot relayed nearly 300,000 Kimi customer requests to Claude, mostly to Claude Opus, through 5,380 fraudulent accounts that appeared to be in Singapore and Japan; showed Claude's answers to users as if Kimi had written them; and kept the exchanges to train its own models, as part of a campaign in which Anthropic says more than 23 million responses were collected. The same report made distillation allegations against DeepSeek and MiniMax.

What it means if you are weighing Kimi up

  • Know which model answers you. Whatever the truth of this case, a model name on screen is a claim, not a guarantee. Ask any AI vendor or reseller which model and which country process your data, and get the answer in the contract.
  • Hosted Kimi is a Chinese service. Prompts sent to kimi.com or Moonshot's API are handled by a Chinese company. Keep client data off it unless your data-protection impact assessment clears it.
  • The weights change the picture. Because K3 can be downloaded, it can run on infrastructure you control, and US hosts such as Together AI have committed to serving it from US infrastructure. At 2.8 trillion parameters, though, self-hosting is a job for a GPU cluster, not a single server; DeepSeek V4.1 Flash, at 552 billion parameters, is far easier to run.
  • Judge it on price, not hype. At $3 / $15, K3 costs more per token than Grok 4.7 or Muse Spark 1.3 and scores lower than both. Xiaomi's MiMo-V2.6-Pro, also open-weight and under the more permissive MIT licence, now scores higher at $0.435 / $0.87. On score and price alone, K3 is hard to justify.

What it costs to use them yourself

API prices are in the model table above. For individual chat subscriptions, the whole market has settled on a $20-a-month standard tier, with budget tiers below it and power-user tiers of $100 to $300 above it. Prices are the vendors' US list prices; UK checkouts are billed in sterling with VAT added, so check the vendor page for the exact figure on the day.

ServiceFree tierBudgetStandardPower userWhat the standard tier gets you
ChatGPTYes, GPT-5.6 Luna, unlimited text; ads in some marketsGo $8Plus $20Pro $100 / $200GPT-5.6 Sol in chat; GPT-6 Astra only inside ChatGPT Work and Codex (Astra in ordinary chat needs Pro)
ClaudeYes, daily caps—Pro $20 ($17 annual)Max $100 (5×) / $200 (20×)Opus 5.5 included, with five-hour usage limits raised at its launch; Fable 5.1 only on pay-as-you-go usage credits (Max includes it for up to half the weekly allowance)
GeminiYes, 3.6 Flash plus limited 3.1 ProAI Plus $4.99AI Pro $19.99AI Ultra $99.99 / $199.993.1 Pro with 1M context, plus 3.8 Flash in the app; Ultra was cut from $249.99 at I/O 2026
GrokYes, Grok 4.6 with tight limitsX Premium $8 / SuperGrok Lite $10SuperGrok $30SuperGrok Plus $100 / Heavy $300Grok 4.6 and Grok Bot; Grok 4.7 launched on the API, Cursor and Grok Build first, with no app date yet
KimiYes, K3 is free in the app—from about $19up to about $199Kimi K3, with context length tiered by plan
DeepSeekYes, web chat is free—no monthly plan—Everything else is per-token API ($0.15 / $0.60 off-peak for V4.1 Flash), or self-hosted

Two things to check before relying on a consumer plan for work. Training terms differ: xAI's individual Grok plans may use your conversations for training while Grok Business ($30 per seat) does not, and Meta's cheapest Muse Spark API tier is cheap precisely because Meta may train on what you send. Availability differs too: some Gemini app features have excluded the UK and EEA at launch, and Claude Mythos 5.1 is not sold on any public plan. Business and Enterprise tiers from every vendor add admin controls, data-processing terms and, in most cases, EU or UK data residency.

Running an AI agent on your own PC

Everything above runs in someone else's data centre. A local agent keeps prompts and files on your own machine, in exchange for a much smaller model. It takes two pieces: a model runner such as Ollama or LM Studio, and an agent layer that plans and uses tools around the model. OpenClaw suits a persistent personal assistant, OpenHands and Cline suit software engineering, and goose suits general local automation. Hermes Agent, from Nous Research, learns reusable skills from each completed task and ships tool-call parsers for local models, so a small Qwen behaves predictably.

Memory decides almost everything. A 4-bit (Q4) model needs roughly 0.6 GB of video memory per billion parameters, before you add room for the conversation. The newest mixture-of-experts models help here: they hold many parameters in memory but only work a few billion of them per word, so they run far faster than their size suggests.

What tokens per second actually feels like

Models write in tokens, not words. In English prose a token is about three-quarters of a word, so 100 tokens is roughly 75 words; code uses more tokens per line because of symbols and indentation. An average adult reads silently at about 238 words a minute, which is only around 5 tokens per second. Anything above 10 or 15 tokens per second therefore outruns your eyes on a normal answer. Speed starts to matter when the model writes far more than you read: reasoning models think privately before answering, and coding agents write whole files and retry until the tests pass. Drag the slider or pick a machine to see the difference.

A token is a chunk of text, usually part of a word. This paragraph is being written at the speed you picked, so you can see how it feels in practice. Above about fifteen tokens per second it outruns most readers, and a short answer like this one feels instant for a short answer. It feels very different when a coding agent has to write a two-hundred-line file, run the tests, read the errors and write the file again, because every one of those steps is paid for in tokens. The faster the machine, the less time you spend waiting for work you will never read.

What the model writesSizeTime at this speed
A short answer to a question~150 words, ~200 tokens7 sec
A detailed, one-page answer~600 words, ~800 tokens27 sec
A 150-line code file~1,500 tokens*50 sec
A reasoning model's hidden thinking, then a short answer~2,200 tokens1 min 13 sec
One agent task, at GPT-6 Astra's average output~27,000 tokens15 min

Times assume the speed above holds for the whole answer and ignore the wait before the first word appears. *Assumes about 10 tokens per line of code. The reasoning example uses 2,000 thinking tokens before a 200-token answer; the agent example is the ~27k output tokens Artificial Analysis measured for GPT-6 Astra per index task, against ~119k for Claude Opus 5.5 at max effort.

With that in mind, here is what each build can realistically run. Every PC row assumes the same base machine: an Intel Core Ultra 7 with 64 GB of RAM. The coding level column uses our own four-step scale, explained after the tables.

The Coding vs Opus 5 column shows how close the best coding model each machine runs at a usable speed comes to Claude Opus 5, on one shared yardstick: the Artificial Analysis Coding Index, where Opus 5 scores 78.0. The best coder that fits on a single 24 GB graphics card, Qwen 3.8 27B, scores 68.1, or 87% of Opus 5. The best that any machine on this page holds, Qwen3.8-Flash-Next at about 123 GB, scores 73.0, or 94%. Claude Opus 5.5 was released today and is not on the index yet; it beats Opus 5 on every coding test Anthropic published, and by 11 points on the independent Terminal-Bench 4.0 run, so the gap to Opus 5.5 is wider than these bars show.

BuildBest all-round modelBest for agents and codingWhat to expectCoding levelCoding vs Opus 5
PC, no graphics cardGemma 4 26B-A4B or Qwen 3.6 35B-A3B, both mixture-of-experts with 3–4B parameters active per wordQwen 3.6 35B-A3B for coding; gpt-oss 20B for the cleanest tool callsWorkable for private drafting and summarising, too slow for coding agents. A 64 GB machine with no GPU runs a dense 14B model at about 10–15 tokens per second in published tests. gpt-oss 120B will not fit: its weights alone are 63.5 GB.154%
PC with an RTX 4070 (12 GB)Qwen 3.5 9B at Q6 (about 9 GB, around 30 tokens per second) or Gemma 4 12B at Q4 (about 7 GB)Gemma 4 12B for coding; gpt-oss 20B loads at Q4 but leaves almost no room for contextSnappy chat and light tool use. Anything 14B or larger spills into system RAM and drops to around 7 tokens per second.1–240%
PC with an RTX 4090 (24 GB)Gemma 4 26B-A4B at Q4: about 16 GB, around 85 tokens per second, 256K contextQwen 3.8 27B (about 17 GB at Q4), the strongest coder that fits one card. Qwen 3.6 27B is faster per token, and Poolside's Laguna XS 2.1 is a lighter agentic optionThe sweet spot, and the first tier where a local coding agent is genuinely useful. Skip 70B dense models: they do not fit at Q4.387%
PC with 2× RTX 3090 (48 GB in total)Gemma 4 31B or Qwen 3.6 27B at higher precision with a longer contextQwen 3.8 27B for coding on one card, with an assistant model on the otherEach 3090 runs Qwen 3.6 27B at about 35 tokens per second. 70B models fit at Q4, but the newer 27–31B models outscore them. Use vLLM for NVLink-aware splitting of one model across both cards. Still short of the 60 GB gpt-oss 120B wants.387%
PC with 2× RTX 5090 (64 GB in total)Qwen 3.6 35B-A3B at Q6 on one card (about 28 GB, around 80 tokens per second) with Gemma 4 31B on the otherQwen 3.8 27B at high precision; Mistral Small 4 (119B) fits across both cards but codes worseThe fastest consumer build, but more memory is not more intelligence: Mistral Small 4 scores 43 on BenchLM's leaderboard against 55 for Gemma 4 31B, which fits on one card.387%

Apple Mac Studio and MacBook Pro

Macs work differently. The processor and graphics share one pool of unified memory, so a Mac can load models far larger than any consumer graphics card holds. The trade-off is speed: writing each token is limited by memory bandwidth, where Nvidia cards are faster, while reading your prompt is limited by raw compute, where the M5 generation made its biggest jump. In Apple's own measurements, M5 produced the first token 3.3 to 4 times faster than M4 but subsequent tokens only 19–27% faster. For coding agents, which send long prompts full of code, that first-token speed is the one you feel.

Two supply notes. There has never been an M3 Max Mac Studio: that chip shipped only in the MacBook Pro, so it is listed here as a laptop. And the memory shortage has hit Macs hard: Apple removed the 512 GB M3 Ultra option in March 2026, and by May the M3 Ultra Mac Studio was sold only with 96 GB. The M5 Max and M5 Ultra Mac Studio, announced on 26 August, restore the 512 GB tier.

MacMemory and bandwidthWhat it can runSpeedClosest cloud equivalentCoding levelCoding vs Opus 5
M3 Max (MacBook Pro only)up to 128 GB, 400 GB/sQwen 3.6 27B and Gemma 4 31B comfortably; gpt-oss 120B squeezes in with 128 GBSlow on dense models: the M1 Max, which has the same 400 GB/s, measured 15.5 tokens per second on Qwen 3.6 27BClose to early-2026 flagships on standard coding tests; well behind today's frontier on hard tasks2–387%
M4 Max (Mac Studio 2025)up to 128 GB, 546 GB/sgpt-oss 120B (about 65 GB) and 122B-class mixture-of-experts models16.6 tokens per second in a published Q4 run of Qwen 3.6 27B; far faster on mixture-of-experts modelsAs above; the larger MoE models add knowledge, not frontier reasoning387%
M3 Ultra (Mac Studio 2025)96 GB (256 and 512 GB withdrawn), 819 GB/sAt 96 GB, slightly less room than a 128 GB M4 Max; used 256–512 GB units can hold 235B-class models28.6 tokens per second measured on Qwen 3.6 27B via OllamaAs above, with the fastest dense-model speeds of the 2025 Macs387%
M5 Max (Mac Studio 2026)up to 128 GB, 614 GB/sgpt-oss 120B; DeepSeek V4 Flash (284B) at 2-bit76–88 tokens per second measured on gpt-oss 120B in MLX (on an M5 Max MacBook Pro); about 39 on DeepSeek V4 Flash at 2-bit (community figure)As the M4 Max, but noticeably faster: the most practical big-model Mac387%
M5 Ultra (Mac Studio 2026)96–512 GB, 1.2 TB/sAt 512 GB (about 435 GB usable): Llama 4 Maverick 400B, several 120B models at once, and by our sizing arithmetic DeepSeek V4.1 Flash at Q4Estimates: 20–25 tokens per second on a 70B dense model, 70–100 on a mixture-of-experts model with 20B active parametersWith DeepSeek V4.1 Flash loaded: roughly level with GPT-5.6 Luna and Gemini 3.8 Flash on the intelligence index394%

M5 Ultra prices start at $5,499 with 96 GB and run to $18,299 fully loaded; the 96 and 256 GB models ship from 22 September and the 512 GB model in late October. M5 Ultra speeds are published estimates, not measurements, until shipped units are benchmarked. Measured figures depend on the runtime (MLX, Ollama, llama.cpp), quantisation and context length, so treat them as a guide.

Dedicated AI computers

A new class of machine sits between the gaming PC and the datacentre. Some copy Apple's approach of one large pool of shared memory, but use AI chips instead of desktop processors. Others stack professional graphics cards with 96 GB each. All of them run the same open-weight models as the builds above. What changes is how big a model fits, how fast it runs, and how many people it can serve at once.

MachineMemory and bandwidthWhat it can runCoding levelCoding vs Opus 5Estimated cost (US)Status
NVIDIA DGX Spark
deskside appliance
128 GB unified, 273 GB/s. GB10 Grace Blackwell chip with a 20-core Arm CPUNVIDIA quotes models up to about 200B parameters, or about 405B with two units linked. Its sweet spot is mixture-of-experts models such as gpt-oss 120B; dense 70B models run slowly on this bandwidth387%$4,699 (launched at $3,999)On sale. Runs NVIDIA's Linux-based DGX OS, not Windows
HP ZGX Nano G1n
deskside appliance
128 GB unified. Same GB10 chip as the DGX SparkThe same as the DGX Spark387%$6,499–7,399On sale; the dearest GB10 box reviewed so far
HP Z2 Mini G1a
x86 mini workstation
128 GB unified, 273 GB/s. AMD Ryzen AI Max+ PRO 395gpt-oss 120B on its integrated Radeon graphics, no discrete card. The same chip runs gpt-oss 120B at about 30 tokens per second. Runs normal Windows 11 Pro387%About $3,300–4,900; prices have climbed with memory costsOn sale. AMD's next chip, the Ryzen AI Max Pro 400, raises the ceiling to 192 GB
HP OmniBook Ultra 16 and X 14
NVIDIA RTX Spark laptops
Up to 128 GB unified memory, up to 1 petaflop of FP4 AI computeThe same model sizes as a DGX Spark, on battery, in a Windows laptop3 (expected)87%Not announcedAnnounced 4 September; due this autumn. An always-on HP OmniDesk desktop follows
Dell Precision 7875
multi-GPU tower
2× RTX PRO 6000 Blackwell Max-Q: 192 GB of graphics memory at 1.8 TB/s per card120B-class models at high precision with fast replies; DeepSeek V4 Flash (284B) at Q4 by our sizing arithmetic394%About $40,000–55,000 (our estimate: each card now costs $14,000–16,000)On sale with Threadripper PRO 9000 processors
HP Z8 Fury G6i
multi-GPU tower
Up to 4× RTX PRO 6000 Max-Q: 384 GB of graphics memoryDeepSeek V4.1 Flash at Q4 held entirely in fast graphics memory (by our arithmetic), fast enough to serve a team394%About $70,000–90,000 fully loaded (our estimate; HP publishes no price)On sale
MSI XpertStation WS300 and HP ZGX Fury
GB300 deskside station
748 GB coherent: 252 GB of HBM3e at 7.1 TB/s plus 496 GB of LPDDR5XTrillion-parameter models at FP4, so the top open model, MiMo-V2.6-Pro, fits. The only machines on this page that can run it394%$85,000–99,999 for the MSI; HP has not published a priceMSI shipping; HP orderable
AMD Threadripper Halo Station
x86 accelerator workstation
96-core Threadripper PRO, 2 TB DDR5, 2× Instinct MI350P, each with 144 GB of HBM3E at 4 TB/s (288 GB in total; 576 GB with four)AMD claims trillion-parameter models. That is a capacity claim; expect it to be usable for experiments rather than fast serving394%No price. Parts alone exceed $100,000; a finished system could pass $150,000Prototype shown at IFA on 4 September; no date

Every machine here tops out at coding level 3. Even the $85,000 GB300 stations run the best open model, which scores 46 on the intelligence index, against 58 for Claude Opus 5.5 in the cloud. What the expensive machines buy is capacity and speed for a whole team, with nothing leaving the building, rather than frontier intelligence. Bandwidth decides single-user speed: the GB10 and Ryzen AI Max boxes share the same 273 GB/s, about a third of an M3 Ultra, so they favour mixture-of-experts models over dense ones.

What each option costs

Prices are US list or street prices seen in August and September 2026, or our own estimates where no price has been published. Memory and graphics card shortages have pushed almost every figure up this year, so check the day's price before you buy. UK checkouts are in sterling with VAT added. Electricity, software and support are not included. The two new columns show the ceiling of each option: the strongest model it can hold, with its Artificial Analysis Intelligence Index score where one exists, the coding level from our four-step scale, and the fastest published speed. Speeds are for different models, so compare them with care: a fast figure on a small model is not the same as a fast figure on a large one. All PC rows share the Core Ultra 7, 64 GB base. The Coding vs Opus 5 column is explained above the first local table.

OptionEstimated costMax intelligence (best model it can hold)Coding vs Opus 5Max speed (fastest published figure)Basis
PC, no graphics cardAbout $1,800–2,500Gemma 4 26B-A4B or Qwen 3.6 35B-A3B. Coding level 154%10–15 tok/s (dense 14B)Our estimate. A 64 GB DDR5 kit alone now costs $769–929
PC with an RTX 4070 (12 GB)About $2,300–3,200Qwen 3.5 9B or Gemma 4 12B. Level 1–240%~30 tok/s (Qwen 3.5 9B)Our estimate. The card is discontinued, so remaining or used stock
PC with an RTX 4090 (24 GB)About $3,200–6,500Qwen 3.8 27B: index 34. Level 387%~85 tok/s (Gemma 4 26B-A4B)A used 4090 costs $1,400–2,000, a new one $3,700–4,300
PC with 2× RTX 3090 (48 GB)About $4,000–5,500Qwen 3.8 27B at higher precision: index 34. Level 387%~35 tok/s (Qwen 3.6 27B)Used 3090s cost $1,000–1,300 each; allow for a larger power supply
PC with 2× RTX 5090 (64 GB)About $9,000–10,500Qwen 3.8 27B at high precision: index 34. Level 387%~80 tok/s (Qwen 3.6 35B-A3B)About $3,500 per card, plus a workstation-class power supply
MacBook Pro, M3 Max, 128 GBResale onlyQwen 3.8 27B; gpt-oss 120B fits. Level 2–387%~15 tok/s (Qwen 3.6 27B, est.)Discontinued; no reliable current price
Mac Studio, M4 Max, 128 GBAbout $4,000–7,000Qwen 3.8 27B; gpt-oss 120B and 122B-class models. Level 387%16.6 tok/s (Qwen 3.6 27B); faster on MoEReplaced by M5; €4,399 at a European retailer, $4,000–7,000 on US resale during the shortage
Mac Studio, M3 Ultra, 96 GB$3,999 listQwen 3.8 27B; gpt-oss 120B. Level 387%28.6 tok/s (Qwen 3.6 27B)Replaced by M5; larger memory options were withdrawn
Mac Studio, M5 MaxFrom $2,499 (£2,499)Qwen 3.8 27B; DeepSeek V4 Flash at 2-bit (128 GB). Level 387%76–88 tok/s (gpt-oss 120B)Apple list price for 36 GB; the 128 GB model costs more
Mac Studio, M5 Ultra$5,499–18,299DeepSeek V4.1 Flash at 512 GB: index 39. Qwen3.8-Flash-Next for coding. Level 394%70–100 tok/s (MoE, 20B active, est.)Apple list: 96 GB to fully loaded 512 GB
HP Z2 Mini G1a, 128 GBAbout $3,300–4,900Qwen 3.8 27B; gpt-oss 120B. Level 387%~30 tok/s (gpt-oss 120B)Retail listings
NVIDIA DGX Spark$4,699Qwen 3.8 27B; models up to ~200B. Level 387%~30 tok/s (gpt-oss 120B, est. from matching bandwidth)NVIDIA US store, August 2026
HP ZGX Nano G1n$6,499–7,399As DGX Spark. Level 387%As DGX SparkHP direct, 2 TB to 4 TB
HP OmniBook RTX Spark laptopsNot announcedAs DGX Spark (up to 128 GB). Level 3 expected87%Not publishedDue this autumn
Dell Precision 7875, 2× RTX PRO 6000About $40,000–55,000Qwen3.8-Flash-Next (~123 GB) for coding. Level 394%Not published; 1.8 TB/s per cardOur estimate from card prices
HP Z8 Fury G6i, 4× RTX PRO 6000About $70,000–90,000DeepSeek V4.1 Flash: index 39. Qwen3.8-Flash-Next for coding. Level 394%Not published; 1.8 TB/s per cardOur estimate from card prices
MSI XpertStation WS300 (GB300)$85,000–99,999MiMo-V2.6-Pro: index 46, the best open model. Level 394%Not published; 7.1 TB/s HBM3eMSRP and retail listings
AMD Threadripper Halo Station$100,000–150,000+Trillion-parameter models (AMD claim). Level 394%Not published; 4 TB/s per acceleratorPress estimates from component prices; AMD has set no price
For comparison: a cloud plan$20 a monthClaude Opus 5.5: index 58. Level 4100%+~262 tok/s (Gemini 3.8 Flash API); 71.4 (GPT-6 Astra)List price; Artificial Analysis measurements

For scale: $4,699, the price of a DGX Spark, would pay for a $20-a-month Claude or ChatGPT plan for about 19½ years, or buy about 235 million Claude Opus 5.5 output tokens at $20 per million. Local hardware pays back through privacy, predictable cost and very high volume, not by being cheaper for occasional use.

The best setup for your budget

Pulling the tables together, here is what we would buy at each price point today, and what is worth waiting for. Prices are US prices in September 2026. "A PC you already own" assumes a desktop with a free graphics slot and a power supply big enough for the card. The Coding vs Opus 5 bar is the same yardstick used in the tables above.

BudgetBest pickWhat it gets youCoding vs Opus 5AlternativesComing later this year
Under $1,000Add an RTX 5060 Ti 16 GB to a PC you already own: about $670–805 now (it launched at $429)Runs Gemma 4 12B and gpt-oss 20B comfortably, and a 27B model squeezed to 3-bit if you accept some quality loss. Coding level 1–240%No PC? A Mac mini M6 with 16 GB ($899) runs the same class of model, quietly, but on slower memory (153 GB/s)Card prices are rising, not falling, so buy when you need it
$1,000–2,000Add a used RTX 3090 (24 GB) to a PC you already own: $1,000–1,300The cheapest route to Qwen 3.8 27B, the strongest coder that fits on one card, at around 35 tokens per second on 27B models. Level 387%Starting from scratch: a Mac mini M6 with 32 GB ($1,299) runs the same model, but slowly (170 GB/s). A used RTX 4090 ($1,400–2,000) is about 40% faster than a 3090The Mac mini M5 Pro starts at $1,699, but its 24 GB base is tight; 48 GB costs $2,299
$2,000–4,000A complete PC with a used RTX 4090: about $3,200–4,500The fastest single-card setup: Gemma 4 26B at around 85 tokens per second, and Qwen 3.8 27B for coding. Upgradable later. Level 387%Mac Studio M5 Max with 48 GB ($3,099) or 64 GB ($3,499): silent, 614 GB/s. HP Z2 Mini G1a with 128 GB ($3,300–4,900) if you need 120B-class modelsHP's OmniBook RTX Spark laptops, with up to 128 GB, are due this autumn; no price yet
$4,000–6,000Mac Studio M5 Max with 128 GB: $5,099Holds gpt-oss 120B, measured at 76–88 tokens per second on this chip, plus DeepSeek V4 Flash at 2-bit, with Qwen 3.8 27B for coding. Level 387%M5 Ultra with 96 GB ($5,499) for the fastest speeds on 27B-class models. NVIDIA DGX Spark ($4,699) if you need NVIDIA's CUDA software. A PC with 2× used RTX 3090 ($4,000–5,500)—
$6,000–10,000A PC with 2× RTX 5090: about $9,000–10,500, at the top of the bandThe fastest consumer setup: around 80 tokens per second on Qwen 3.6 35B-A3B, with Qwen 3.8 27B at high precision on one card and an assistant on the other. Level 387%Two linked DGX Sparks (about $9,400) share 256 GB, which NVIDIA says handles models of about 405B. That is enough for Qwen3.8-Flash-Next (94%), at modest speed. An M5 Ultra with the 36-core chip and 96 GB costs $6,800The 256 GB M5 Ultra sits just above this band
$10,000–15,000Mac Studio M5 Ultra with 256 GB: about $10,800, or $11,299 with 2 TBHolds Qwen3.8-Flash-Next (about 123 GB), the best local coder on this page, with 1.2 TB/s of bandwidth and room for a second model. Level 394%A single RTX PRO 6000 (96 GB) costs $14,000–16,000 for the card alone, so it does not fit this band as a full systemThe 512 GB M5 Ultra arrives in late October, unpriced; it adds DeepSeek V4.1 Flash (index 39)

The curve flattens fast. $1,000–2,000 already reaches 87% of Opus 5 on the Coding Index, and a further $9,000 adds only another seven points. What bigger budgets mainly buy is speed, room for larger general-knowledge models and capacity for a team. No budget here reaches Claude Opus 5.5, which a $20-a-month plan includes.

Local against the cloud: speed

The fastest cloud models write several times faster than any local machine, and they do it while running much larger models. The local figures below are for the best model each machine runs comfortably, so they are not like-for-like: a local 27B model is smaller than anything in the cloud rows.

A cloud API, independently measured L local, measured E local, estimated
1Gemini 3.8 Flash cloud~262A
2DeepSeek V4.1 Flash cloud207.5A
3M5 Max, gpt-oss 120B Mac76–88L
4RTX 4090, Gemma 4 26B-A4B PC~85E
5RTX 5090, Qwen 3.6 35B-A3B PC~80E
6GPT-6 Astra cloud71.4A
7Claude Fable 5.1 cloud69.7A
8RTX 3090, Qwen 3.6 27B PC~35E
9Ryzen AI Max+ 395, gpt-oss 120B HP Z2 Mini G1a class~30L
10M3 Ultra, Qwen 3.6 27B Mac28.6L
11M5 Ultra, 70B dense model Mac20–25E
12M4 Max, Qwen 3.6 27B Mac16.6L
13Core Ultra 7, 14B dense model PC10–15E

Tokens per second, single user. Bars are scaled to Gemini 3.8 Flash. Cloud figures are Artificial Analysis output-speed measurements; Claude Opus 5.5 has not yet been measured independently. A local machine serves one person at full speed; a cloud API serves you at this speed however many of your staff use it at once.

Local against the cloud: intelligence and coding

This is where the gap is widest. On the Artificial Analysis Intelligence Index, the best model a machine under $20,000 can hold is DeepSeek V4.1 Flash, on a 512 GB M5 Ultra. It scores 39, level with the budget cloud tier. Only the $85,000-plus GB300 stations can hold the strongest open model, MiMo-V2.6-Pro, and no local machine can run the frontier models at all.

1Claude Opus 5.5 cloud only58I
2GPT-6 Astra / Claude Fable 5.1 cloud only53I
3MiMo-V2.6-Pro open; fits a GB300 station46I
4GLM-5.3 open, needs datacentre GPUs45I
5Gemini 3.8 Flash cloud budget tier41I
6DeepSeek V4.1 Flash fits a 512 GB M5 Ultra or 4× RTX PRO 600039I
7GPT-5.6 Luna cloud budget tier37I
8Qwen 3.8 27B fits one 24 GB graphics card34I
9Qwen 3.6 27B fits one 24 GB graphics card21I

Artificial Analysis Intelligence Index v4.3. Qwen 3.8 27B, the strongest model that fits on a single 24 GB graphics card, scores 34 at its highest effort setting; its predecessor Qwen 3.6 27B scores 21. On coding the picture is more mixed. On the Artificial Analysis Coding Index, Qwen 3.8 27B scores 68.1 against 78.0 for Claude Opus 5, and on the older SWE-bench Verified test Qwen 3.6 27B's 77.2% is within four points of Claude Opus 4.6's 80.8%. On the harder Terminal-Bench 4.0, the best open model, MiMo-V2.6-Pro, reports 34.9% against 59.6% for Claude Opus 5.5.

Level 1: helper

Explains code, writes snippets and small scripts that you copy in and run yourself. Any machine on this page manages it.

Level 2: single-file

Writes or edits one file or a short program reliably, but agent loops that run tests and retry are slow enough to be frustrating.

Level 3: repository agent

Plans and edits across a real codebase and runs the tests, using Qwen 3.6 27B-class models at a usable speed. The ceiling for local hardware today.

Level 4: frontier

Long, unattended work on large codebases, such as migrations and audits that run for hours. Cloud only: Claude Opus 5.5, GPT-6 Astra and Claude Fable 5.1.

Desktop coworkers compared: Claude, ChatGPT and Copilot

All three vendors now sell an assistant that works across your files, apps and desktop rather than just answering questions. Anthropic merged Claude's chat and Cowork modes into one window on 16 September. OpenAI rebuilt its desktop app around ChatGPT Work on 9 July, combining Chat, Work and Codex. Microsoft's Copilot Actions carries out tasks in a separate agent workspace on Windows, but as of September 2026 it has not reached general availability and remains an experimental preview.

1. Claude (Cowork in Claude Desktop)

The most capable and the most careful. Claude Opus 5.5 leads the desktop-task benchmark, the connector directory is deep, and scheduled tasks keep running in the cloud with your PC off. It is slower than ChatGPT and explains more of its working.

2. ChatGPT desktop app (Work)

The fastest and the widest. It has the largest app directory and the only event triggers of the three, so it can react to a new email rather than waiting for a timer. It matched Claude's accuracy in independent tests while finishing sooner.

3. Microsoft Copilot Actions

The most cautious design: each agent gets its own Windows account and workspace. But it is an experimental preview with no scheduler and folder access limited by default. For Microsoft 365 automation today, Copilot Studio is the more mature route.

CapabilityClaude (Cowork / Claude Desktop)ChatGPT desktop appMicrosoft Copilot Actions
Connector directory (MCP)1,253 connectors listed in August 2026, 474 of them partner-verified2,289 apps listed in August 2026Agent connectors registered in the Windows On-Device Registry (preview). Copilot Studio separately offers 1,400+ connectors plus MCP servers
Your own MCP serversYes: custom connectors, plus local servers in Claude DesktopRemote servers only, via Developer mode, set up from chatgpt.com in a browserVia Copilot Studio (Add a tool, then Model Context Protocol), not in the Copilot app itself
Full desktop control without MCPYes, in beta for Pro and Max on macOS and Windows. Uses a connector first, then the browser, then the screen. On macOS 15 or later it works in background windows without taking over your pointerYes. Computer use in the desktop app on Mac and Windows, demonstrated organising Apple Notes, plus voice control of the computer since 23 JulyYes, but inside a separate agent workspace with its own Windows account. By default it can reach only Documents, Downloads, Desktop, Music, Pictures and Videos
Desktop task score (OSWorld 2.0, partial credit)81.8% with Opus 5.5 (vendor)72.6% with GPT-6 Astra (vendor)Not published
How fastSlower. In a 22 September head-to-head it took 1:44 against 1:17 on a research task and about 4 minutes against 28 seconds on a Drive-to-spreadsheet job, with identical accuracy. Screen control is its slowest route by designFastest on all three tests in the same comparison. GPT-6 Astra completes desktop tasks around 47% faster than GPT-5.6 SolNo published timings. It runs in a parallel session, so it does not block you while it works
Runs without re-prompting?Yes, on a timer. Scheduled tasks run hourly, daily, weekly or on weekdays, remotely, even with the PC asleep or the app closed. Tasks that need local files or apps run on your machine, which must be on. No event triggersYes, on a timer or an event. Scheduled tasks plus triggers from new Gmail, Slack or GitHub activity on Plus plans and above. A task can pause if it needs your action, and sending messages may need approvalNo. Each task starts from a prompt and runs while its conversation is open; Windows will not sleep until you close it. No scheduler
Paid from your monthly plan or API credits?Your plan first. Cowork shares one five-hour and weekly allowance with chat and Claude Code on Pro, Max, Team and seat-based Enterprise. Past the limit you can wait, upgrade, or switch on optional usage credits billed at standard API rates. Fable 5.1 runs only on usage credits outside Max. Separate Console API credits are used only if you choose themYour plan first. On Plus and Pro, Work, Codex and ChatGPT for Excel draw on one shared agentic allowance. Past it you can buy optional flexible credits, with automatic reload if you want it; OpenAI states these are not API credits. Business and Enterprise use a pooled workspace credit balanceNo separate charge has been published for the Windows preview. For business agents, Microsoft 365 Copilot is a per-user licence from $30, and internal agents used by licensed staff are largely zero-rated. Copilot Studio agents, and Microsoft 365 Copilot's own Cowork tasks, consume Copilot Credits at $0.01 each pay-as-you-go or $200 per 25,000 a month

Connector counts come from Node8's August 2026 capture of both public directories and change weekly. OSWorld scores are each vendor's own and use different setups. Timings are from The New Stack's three-task comparison published on 22 September.

Sixty common business apps and where they connect

Most business software now publishes an MCP server, the plug that lets an AI assistant read and act inside it. What differs is whether the app is a one-click listing or something your IT team has to wire in. These are sixty of the most widely used apps with MCP servers, not a ranking by usage.

✓ one-click listing in the product's directory ◐ add it yourself as a custom MCP server Native built into the product
Productivity, files and meetingsClaudeChatGPTCopilot
Gmail✓✓◐
Google Drive✓✓◐
Google Calendar✓✓◐
Microsoft 365 (Outlook, Teams, SharePoint)✓✓Native
Slack✓✓◐
Zoom✓✓◐
Notion✓✓◐
Atlassian (Jira, Confluence)✓✓◐
Asana✓✓◐
monday.com✓✓◐
ClickUp✓✓◐
Trello✓✓◐
Wrike✓✓◐
Teamwork.com✓✓◐
Todoist✓✓◐
Linear✓✓◐
Smartsheet✓✓◐
Airtable✓✓◐
Dropbox✓✓◐
Box✓✓◐
Egnyte✓✓◐
Docusign✓✓◐
Jotform✓✓◐
Calendly✓✓◐
Fireflies✓✓◐
Otter.ai✓✓◐
Fathom✓✓◐
Granola✓✓◐
Gamma✓✓◐
Whimsical✓✓◐
Sales, finance, IT and developmentClaudeChatGPTCopilot
HubSpot✓✓◐
Zoho CRM✓✓◐
Pipedrive◐✓◐
Intercom✓✓◐
Mailchimp✓✓◐
Klaviyo✓✓◐
Semrush✓✓◐
Ahrefs✓✓◐
Stripe✓✓◐
PayPal✓✓◐
QuickBooks✓✓◐
Xero✓◐◐
Shopify✓✓◐
Canva✓✓◐
Figma✓✓◐
Miro✓✓◐
WordPress.com✓✓◐
Wix✓✓◐
Zapier✓◐◐
ServiceNow✓◐◐
Freshservice✓◐◐
PagerDuty✓◐◐
Malwarebytes✓✓◐
GitHub◐✓◐
Postman✓◐◐
Sentry✓✓◐
Supabase✓✓◐
Vercel✓✓◐
Cloudflare✓✓◐
Datadog✓✓◐

Claude and ChatGPT status comes from Node8's August 2026 capture of both directories. Gmail, Google Drive and Google Calendar in ChatGPT are OpenAI's own built-in apps; Microsoft 365 appears there as four separate apps (Outlook Email, Outlook Calendar, Teams and SharePoint). The Copilot column refers to Microsoft 365 Copilot and Copilot Studio, where an administrator adds third-party MCP servers as tools. Copilot Actions on Windows cannot use any of these.

What this means for a UK business

  • Day-to-day development: Claude Opus 5.5 is now the best capability per pound for PHP, C#, Python and SQL work, at $4 / $20 per million tokens, with Sonnet 5 at $2 / $10 for routine tasks. Keep GPT-6 Astra as a second opinion on the hardest long-running jobs, and Gemini 3.8 Flash for bulk, low-stakes tasks, but budget for its price doubling on 1 January.
  • Security work: cloud models over-block, and Opus 5.5 is stricter on security requests than Opus 5. Use Opus 5 where the cloud is acceptable, look at Anthropic's Cyber Verification Program if this is routine work for you, and use a self-hosted DeepSeek V4.1 Flash or GLM-5.3 for malware triage and vulnerability hunting, which also keeps client code inside the UK.
  • Client data: prefer Business or Enterprise tiers with UK/EU data-processing terms, or open weights on your own hardware. A $20 consumer plan is not a data-processing agreement.
  • Desktop agents: Claude and ChatGPT can now operate your PC and your Microsoft 365 apps; since July, Claude’s Microsoft 365 connector can send mail, manage calendars and create files in OneDrive and SharePoint. Pilot on a test machine with backups, grant access app by app, and keep Copilot Actions to a test device until Microsoft releases it properly.
  • Know who is actually answering: the Kimi allegation, and a fraudulent "discount Claude" reseller described in the same Anthropic report, both make the point that a model name on screen is a claim. Buy AI access through authorised channels, ask which model and which country process your data, and get it in writing.
  • High-volume, low-stakes work: this is where the usage table above should change your shortlist. If you are processing thousands of documents rather than solving hard problems, the cheap open-weight models carrying most of the world's tokens, such as DeepSeek V4.1 Flash, GLM-5.3-Flash and Tencent's Hy4, will do it at a fraction of frontier pricing, and Xiaomi's new MiMo-V2.6-Pro is worth adding to that shortlist.
  • Anything that leaves the building: legal, medical and client-facing text gets a human read every time, whichever model wrote it. Fable 5.1 scored 67% on the hardest published factual test, so even a top model still gets about a third of those questions wrong or unanswered.
  • Marketing assets: GPT Image 2.5 for graphics with words in them, or Ideogram 4.0 if you want open weights. For video, Google's Gemini Omni Flash or Veo 3.1 are the simplest to license, Seedance 2.5 through Higgsfield suits 30-second multi-shot adverts, and fal's MiniMax H3 Max is the pick if cost per minute matters most. Whatever you use, keep real people, brands and film characters out of the prompt.
  • Keep a fallback: models disappear. The Sora 2 API closes on 24 September, and in June Anthropic suspended Fable 5 for nearly three weeks to comply with US export controls, restoring access on 1 July. Build so that switching model is a configuration change, not a rewrite.
  • Expect this page to change: Claude Opus 5.5 arrived on 22 September with Sonnet 5.5 and Haiku 5.5 due in the coming weeks, StepFun says Step 5 weights will follow on 15 October, Google's Gemini 3.5 Pro is still in partner testing, DeepSeek has said a V4.1 Pro is coming, and a Kling 4.0 is widely expected. We will update the figures as they move.

Not sure which model fits your business?

We run these models daily across support, development and security work, and we will tell you plainly which one to use and which to avoid.

Talk to us