Which AI model to use? (Sept 2026)
Confused about which AI model to choose in September 2026? Compare strengths, use cases and performance to pick the right fit for your business.
Every AI lab says its newest model is the best one. Most of them are right about something. This page sets out, category by category, which of the September 2026 models actually leads on medical, coding, legal, cybersecurity, reasoning, images, video and more — with the published scores, the price, and the caveats the launch posts leave out.
The short answer
Best all-round
Claude Opus 5.5, released 22 September, scores 58 on the Artificial Analysis Intelligence Index v4.3 — five points clear of the next models, the widest lead at the top in months. It leads six of the ten tests in the index and costs $4 in / $20 out per million tokens, 20% less than Opus 5.
Close behind
Claude Fable 5.1 and GPT-6 Astra, tied at 53. Astra still leads maths and uses the fewest tokens of any frontier model; Fable 5.1 costs twice Opus 5.5's per-token price. Sonnet 5, now permanently $2 / $10, covers routine work.
Fastest and cheapest
Gemini 3.8 Flash at around 262 tokens per second and $0.75 / $3.75 per million until 31 December, when the price doubles. DeepSeek V4.1 Flash undercuts it at $0.15 / $0.60 off-peak.
Best to self-host
Xiaomi's MiMo-V2.6-Pro (46 on the index, MIT licence, released 21 September) is now the strongest open model, though at 1.02 trillion parameters it needs datacentre-class hardware such as a GB300 station. DeepSeek V4.1 Flash remains the practical pick for code and agents, and GLM-5.3 (45) for general work. Your data never leaves your own servers.
No single model wins everything. The right answer depends on the job.
The models compared
These are the current generally available flagships from each lab, plus the leading open-weight models. Claude Opus 5.5 is the first model in Anthropic's Claude 5.5 family; Anthropic says Sonnet 5.5 and Haiku 5.5 will follow in the coming weeks. Claude Mythos 5.1 is deliberately absent: it is the same underlying model as Fable 5.1 with some safeguards lifted, and it is available only to vetted organisations in Anthropic's trusted-access programmes, not on any public plan or the standard API.
| Model | Maker | Released | Context | API price ($ per million tokens, in / out) |
|---|---|---|---|---|
| Claude Opus 5.5 | Anthropic | 22 Sep 2026 | 1M | $4 / $20 (cache reads $0.20) |
| Claude Fable 5.1 | Anthropic | 1 Sep 2026 | 1M | $10 / $50 |
| Claude Opus 5 | Anthropic | 24 Jul 2026 | 1M | $5 / $25 |
| Claude Sonnet 5 | Anthropic | 30 Jun 2026 | 1M | $2 / $10 (planned rise to $3 / $15 cancelled) |
| GPT-6 Astra | OpenAI | 3 Sep 2026 | 1.05M | $10 / $50 |
| GPT-5.6 Sol | OpenAI | 9 Jul 2026 | 1.05M | $4 / $20 promotional to at least 21 Nov (list $5 / $30) |
| GPT-5.6 Terra | OpenAI | 9 Jul 2026 | 1.05M | $2 / $12 |
| GPT-5.6 Luna | OpenAI | 9 Jul 2026 | 1.05M | $0.20 / $1.20 |
| Gemini 3.8 Flash | 2 Sep 2026 | 1M | $0.75 / $3.75 (doubles to $1.50 / $7.50 on 1 Jan 2027) | |
| Gemini 3.1 Pro | Feb 2026 | 1M | $2 / $12 (prompts up to 200K) | |
| Grok 4.7 | SpaceXAI (xAI) | 21 Sep 2026 | 500K | $2 / $6 |
| Grok 4.6 | SpaceXAI (xAI) | 12 Aug 2026 | 500K | $2 / $6 |
| Muse Spark 1.3 | Meta | 2 Sep 2026 | 1M | $1.25 / $4.25 ($0.10 / $0.20 if Meta may train on your data) |
| MiMo-V2.6-Pro | Xiaomi | 21 Sep 2026 | 1M | $0.435 / $0.87 (open weights, MIT) |
| Kimi K3 | Moonshot AI | 16 Jul 2026 | 1M | $3 / $15 (open weights, custom licence) |
| Qwen3.8-Max | Alibaba | 3 Aug 2026 | ~1M | $2 / $6 (open weights, custom licence) |
| GLM-5.3 | Z.ai | Aug 2026 | 1M | Low; open weights |
| Step 5 Preview | StepFun | 20 Sep 2026 (API) | 1M | $1.00 / $2.70 (weights due 15 Oct) |
| DeepSeek V4.1 Flash | DeepSeek | 10 Sep 2026 | 1M | $0.30 / $1.20 peak, half off-peak (open weights, MIT) |
| DeepSeek V4 Pro | DeepSeek | Aug 2026 | 1M | Open weights (MIT); peak / off-peak API pricing |
| Hy4 preview | Tencent | 28 Aug 2026 | 1M | $0.83 / $2.50 (open weights, Apache 2.0) |
| Mistral Medium 3.5 | Mistral (France) | Apr 2026 | 256K | Low; not fully published |
Category rankings
Reasoning and overall intelligence
The Artificial Analysis Intelligence Index v4.3 combines ten tests, including long-horizon knowledge work, business workflow automation, terminal tasks, document reasoning and factual accuracy, with 45% of the weighting on private test sets so vendors cannot train on them. The current ceiling is 58. Bars here are scaled to the leader.
If you have seen index scores of 57, 61 or 66 quoted for these models, those are older scales from earlier in September — Kimi K3, for example, launched at 57 on version 4.1, and Opus 5 at 61. Only v4.3 numbers are comparable with each other; the figures here are from the current v4.3.2 revision. Opus 5.5 was run at max effort with Anthropic's fallback models switched on, and trails on only three of the ten tests: CritPt, AA-LCR and GDP.pdf. Grok 4.7 was benchmarked on 21 September, the day it launched.
Science knowledge
GPQA Diamond is a set of graduate-level physics, chemistry and biology questions written so that experts outside the field cannot answer them with Google. The frontier is now bunched within six points of the ceiling.
Anthropic did not publish a GPQA Diamond score for Claude Opus 5.5 at launch.
Maths
FrontierMath Tier 4 is research-level mathematics. Note that these are OpenAI-reported figures and that OpenAI funded FrontierMath and had sight of many of its problems, so treat the gap as directional rather than settled.
Coding
The fairest single answer to "can this thing do engineering work" is now Artificial Analysis's Coding Agent Index. Each model runs inside its own maker's coding agent — Claude Code, Codex, Grok Build, Muse Code — on the same three test sets (DeepSWE, Terminal-Bench 4.0 and SWE-Atlas-QnA), scored independently.
Astra ties Fable 5.1 at roughly 60% of the cost per task, because it uses fewer tokens than any other agent on the index. Grok 4.7 gained nine points on Grok 4.6 to take fourth place, overtaking GPT-5.6 Sol. Claude Opus 5.5 had not yet been run through the Coding Agent Index at the time of writing. On Anthropic's own launch figures (vendor-reported), it scores 57.8% on CursorBench 4.0 against 51.8% for Fable 5.1 and 41.7% for GPT-5.6 Sol, and 54.4% on FrontierCode against 53.3% for GPT-6 Astra.
The older shell test is now saturated
Terminal-Bench 2.1, run by Artificial Analysis on the same harness for every model, was the index's shell test until 7 September. Ten models now score 84% or more, so it no longer separates the leaders, but it is still the widest independent comparison that includes the cheaper models.
The harder tests tell a different story
DeepSWE 1.1 puts a model inside a real repository and asks it to fix real issues. Terminal-Bench 4.0 is the much harder shell test that replaced 2.1 in the index. On these the order changes and the field spreads out.
| Model | DeepSWE 1.1 | Terminal-Bench 4.0 (vendor) | Terminal-Bench 4.0 (independent) |
|---|---|---|---|
| Claude Opus 5.5 | — | 66.4% V | 59.6% I |
| Muse Spark 1.3 | 75.4% V | — | — |
| GPT-6 Astra | 74.1% V | 57.7% V | 59.1% I |
| Gemini 3.8 Flash | 73.8% V | 19.1% V | — |
| Claude Opus 5 | 73.6% V | — | 49.0% I |
| GPT-5.6 Sol | 72.7% V | 37.3% V | 39.9% I |
| Claude Fable 5 | 69.9% V | — | — |
| Claude Fable 5.1 | 67.4% V | 55.8% V | 52.0% I |
Opus 5.5's vendor figure is at its xhigh effort setting, with a stated margin of error of about 2.6 points; on Artificial Analysis's independent run it is level with GPT-6 Astra and 11 points ahead of Opus 5. DeepSeek V4 Pro's 80.6% on the older SWE-bench Verified remains the best published open-weight result on that test. Grok 4.7 scored 73% on DeepSWE and 33% on Terminal-Bench 4.0 when Artificial Analysis ran it inside its own Grok Build agent on launch day.
What about PHP, C# and SQL?
Every headline coding score above is Python-biased: SWE-bench Verified is 500 tasks drawn from twelve open-source Python repositories, and DeepSWE and Terminal-Bench lean the same way. Nobody publishes a PHP or C# leaderboard. Two datasets get close, and both point the same way.
| Test | Why it matters for PHP / C# | Leader | Score | Next best |
|---|---|---|---|---|
| SWE-bench Multilingual | Real GitHub issues across nine languages including PHP (C# is not among them). Anthropic is the only frontier lab that publishes a result. | Claude Opus 4.6 | 77.8% V | No published GPT, Gemini or Grok figure. Open-weight Qwen3.6-35B: 67.2% |
| Scale SEAL private commercial repositories | Unseen business codebases that look nothing like popular open-source projects — the closest proxy to a real PHP or .NET application. | Claude Opus 4.6 (thinking) | 47.1% I | Muse Spark 44.7%, GPT-5.4 xHigh 43.4%, Gemini 3.1 Pro 32.2% |
Opus 4.6 dropped only three points moving from Python-only Verified (80.8%) to Multilingual (77.8%), and on the private commercial set Claude degrades less than its competitors when moving from public to unseen repositories. Both figures are from an earlier Claude generation, because the current ones have not been re-run on these tests; the pattern is the point.
Desktop and computer use
OSWorld 2.0 gives a model a real desktop and 108 long workflows that take a human around 1.6 hours each. The headline numbers are partial-credit; the strict "finished the whole job" rate is far lower and shown alongside.
Strict end-to-end completion: Fable 5.1 41.7%, Opus 5 39.6%; no strict figure has been published for Opus 5.5. Anthropic's launch table, run on its current setup, puts Opus 5.5 at 81.8% against 80.7% for Fable 5.1 and 74.0% for Opus 5, so read its lead over Fable as about a point. Fable scores zero on any task where its safety classifiers intervene, which drags its average down on security-adjacent work. Astra completes tasks around 47% faster than GPT-5.6 Sol, so it often wins on wall-clock time despite the lower score.
Global knowledge
Humanity's Last Exam is 2,500 expert-written questions across every academic field. With tools enabled, Claude leads and this is the one academic test GPT-6 Astra loses.
Artificial Analysis's independent run, on its own setup, puts Opus 5.5 at 61.4% against a previous best of 59.1% from Fable 5.1, so the ordering holds. Gemini's figure is on the HLE-Verified variant, so treat it as approximate. Factual accuracy specifically is covered in the hallucination section below.
Biology and life sciences
This is the thinnest category. Only OpenAI publishes head-to-head life-science scores, Anthropic routes biology questions on Fable to Opus 5, and the strongest biology model on either side (Claude Mythos 5.1) is not on public sale: access is limited to vetted organisations in Anthropic's trusted-access programmes. Anthropic says the new Opus 5.5 matches or beats Mythos 5.1 across many areas of biology, which is why it ships with the same biology safeguards as Fable 5.1.
| Model | GeneBench Pro | LifeSciBench | Note |
|---|---|---|---|
| GPT-6 Astra | 37.8% V | 60.3% V | Leads both published tests |
| GPT-5.6 Sol | 28.7% V | 59.9% V | Near-level on LifeSciBench |
| Claude Opus 5.5 | not published | not published | Fable 5.1's biology safeguards; vetted labs and companies can apply to Anthropic's new Life Sciences Verification Program for fuller access |
| Claude Fable 5.1 / Opus 5 | not published | not published | Biology requests on Fable are routed to Opus 5 |
| Grok 4.6 | not published | not published | Won a biosecurity benchmark no other model matched; no percentage released |
| Claude Mythos 5.1 | restricted | restricted | Virology evaluations of 0.81–0.87 reported; vetted organisations only |
Legal (professional)
Legal benchmarks are the most fragmented, and each one is led by a different model. The consistent finding across all of them: models pass around 90% of the individual criteria on a legal task but complete very few whole tasks end to end.
| Benchmark | What it measures | Leader | Score |
|---|---|---|---|
| BigLaw Bench | Drafting and analysis to law-firm standard | Claude (Opus 4.8 reference) | 91.1% I |
| Harvey Bench / GDPval | Professional legal work product | Grok 4.6 | rank only, no % published |
| Legal Research Bench (Vals) | Case-law research accuracy | GPT-5.6 Sol | 48.1% I |
| HAQQ | Legal question answering | DeepSeek V4 Pro | 36.8 / 50 I |
| Harvey LAB | End-to-end agentic legal tasks | Muse Spark 1.1 | 20.0% I |
Reasonable working rule: Claude for drafting and precision, GPT for research, and a human reads everything before it goes out. None of these models should be trusted to complete legal work unsupervised.
Cybersecurity
Two tests matter. ExploitBench measures whether a model can turn a known vulnerability into a working exploit; CyberGym measures whether it can find vulnerabilities in real codebases. A model can be excellent at one and mediocre at the other.
ExploitBench above. On CyberGym (vulnerability discovery) the order flips: DeepSeek V4.1 Flash 88.1%, Gemini 3.8 Flash cyber variant 86.2%, GLM-5.3 84.5%; Xiaomi reports 94.0% for MiMo-V2.6-Pro on its own setup, but only 47.9% on ExploitBench. Claude Fable 5.1 has no published cyber score because penetration-testing, exploit and binary-scanning requests are routed to Opus 4.8. Claude Opus 5.5, which Anthropic rates as comparable to Mythos 5.1 in cybersecurity, has no published score for the same reason: it is the first Opus to ship with Fable-class cyber safeguards. Gemini's cyber variant is only available inside Google's Fairwind programme, and Astra's public build declines advanced offensive work unless your organisation is enrolled in OpenAI's Daybreak scheme.
Tool use and agents
AutomationBench-AA is Artificial Analysis's version of Zapier's business-workflow test: multi-step tasks across real apps. The headline score gives partial credit; the strict figure counts only workflows finished without breaking a single rule.
Artificial Analysis says Claude Opus 5.5 is level with GPT-6 Astra on AutomationBench-AA, but its headline score was not in the launch summary, so it is not charted here. Strict, no-rule-broken completion: Astra 41.6%, Fable 5.1 32.1%, Opus 5 28.3%; Opus 5.5 scored 40.0% in Zapier's own strict run during early access, as published by Anthropic. On AA-Briefcase, the multi-week knowledge-work test, the order reverses: Opus 5.5 now leads at 1,822 Elo, 143 points ahead of Fable 5.1 and the first Anthropic model to beat GPT-5.6 Sol on presentation quality, with Fable 5.1 and Opus 5 behind it and Grok 4.7 joining the frontier just behind them, even though Grok 4.7 slipped 1.1 points against Grok 4.6 on AutomationBench. Grok 4.6 remains the value pick for automation: near the top at a fraction of Astra's price.
Medical and health
Three different questions get asked under "medical", and they have three different answers.
Professional use: answering to a physician's standard
HealthBench Professional is built from 5,000 conversations and physician-written rubrics from 262 doctors in 60 countries, scored on the length-adjusted version so that longer answers do not automatically win. It is OpenAI's benchmark, and all figures are OpenAI-run.
Fable 5.1 scoring below Fable 5 is a safety-routing artefact, not a capability regression: questions its classifiers flag are answered by a smaller model and marked down accordingly. No Gemini, Grok, open-weight or Claude Opus 5.5 score on HealthBench Professional has been published yet.
General medical information for the public
No consumer-grade health benchmark has been published for this generation of models, so the professional ranking above is the best available guide, read together with the hallucination table further down. The practical change this month is on the refusal side: Anthropic reports that Fable 5.1's biology safeguards fire 85% less often on benign elementary biology and medical questions than Fable 5's did, which means fewer "I can't help with that" answers to ordinary health queries.
Scan and imaging interpretation
The imaging benchmarks exist, but they live in academic papers rather than on vendor leaderboards: Radiology's Last Exam (RadLE) tests chatbots against board-certified radiologists on hard spot-diagnosis cases, ReXVQA covers chest X-ray question answering, CXR-LT is a multi-centre chest X-ray challenge with over 145,000 images, and newer sets such as NeuroQA cover 3D brain MRI. No lab publishes scores on any of them, so the only figures come from independent studies — which run a generation behind the current models.
| Modality (RadLE) | Radiologists | Best chatbot | Others |
|---|---|---|---|
| MRI | 98% | GPT-5, 45% | Gemini 2.5 Pro 35%, o3 33%, Grok-4 23%, Claude Opus 4.1 0% |
| X-ray | 89% | GPT-5, 31% | Gemini 2.5 Pro 22%, o3 22%, Grok-4 8%, Claude Opus 4.1 3% |
| CT | 79% | Gemini 2.5 Pro, 29% | GPT-5 22%, o3 19%, Grok-4 8%, Claude Opus 4.1 1% |
Other studies agree on the ordering. On pneumothorax from chest radiographs, ChatGPT-4o was most accurate at 69.6%, then Claude 3.5 at 64.9% and Gemini 2.0 at 57.4% — but every model fell to between 12% and 21% on children, and ChatGPT dropped from 81.6% on large pneumothoraces to 42.2% on small ones. On measuring liver metastases against a radiologist's reference, Gemini reached an agreement score of 0.81, GPT-o3 0.52 and Claude 4 Opus 0.07.
Claude's pattern is consistent: it is the weakest of the majors at reading pixels and the strongest at reading radiology text. On structuring 3,949 head CT reports, Claude was significantly more accurate than GPT and Gemini for intracranial haemorrhage, and on 56 JAMA neuroradiology cases Claude 3.5 achieved the highest accuracy (80.4%) when given the image and the report together.
Explaining a scan
GPT first, Gemini second. Fine for describing findings in plain English to a patient. Not close to a radiologist: the best result on any modality was 45% against 98%.
Working with reports
Claude. Structuring, summarising and extracting findings from radiology text is where it leads, even though it trails badly on the images themselves.
Actually analysing scans
A specialist model, not a chatbot. Google's MedGemma is purpose-built for medicine, reads X-ray, CT and MRI volumes natively, runs locally, and in the ReXVQA reader study scored 83.84% — above every human reader in the panel.
Clinical use
The regulated products deployed in NHS radiology — chest X-ray triage, stroke CT, fracture detection — are UKCA/CE-marked medical devices with published sensitivity per condition. No general chatbot is one.
Two caveats. Every study above tested general models one generation or more behind the current ones — GPT-6 Astra, Claude Opus 5.5, Claude Fable 5.1 and Gemini 3.8 Flash have not yet been run through RadLE. And the RadLE cases are deliberately hard: on routine chest X-rays the chatbots do markedly better, as the pneumothorax study shows, and MedGemma does better still.
Speed
Two different things get called speed. Tokens per second is how fast text appears on screen; time per task is how long an agent takes to finish a job, which depends as much on how many tokens a model needs as on how fast it produces them.
| Model | Output speed (tokens/sec) | Time to first token | Token efficiency |
|---|---|---|---|
| Gemini 3.8 Flash | 262–265 | 3.0s median | Fast but verbose |
| DeepSeek V4.1 Flash | 207.5 | 1.15s | Fast but very verbose: 250M tokens to run the full index, against a 140M median |
| Grok 4.7 | ~188 | — | About 81k output tokens and seven minutes per index task, three times Astra's token count |
| MiMo-V2.6-Pro | ~134 | — | $0.13 per index task, one of the cheapest results at its level; an UltraSpeed variant costs ten times as much |
| GPT-5.6 Terra | ~103 | low | Fastest GPT-5.6 tier |
| GPT-6 Astra | 71.4 | 10s+ | Most efficient at the frontier: about 27k output tokens per index task |
| Claude Fable 5.1 | 69.7 | 10.8s (p95) | Token-hungry: 1.7× the output of Fable 5 |
| Claude Opus 5.5 | not yet measured | — | Anthropic reports output over 30% faster than Opus 5. About 119k output tokens per index task at max, against 73k for Opus 5, but lower prices keep its cost per task level with Opus 5 |
| Muse Spark 1.3 | — | 5.9s (p95) | About 60k output tokens per index task, roughly three times version 1.2 |
| Kimi K3 | — | — | Averages nearly an hour per task on AA-Briefcase |
If you are running agents, token efficiency is the number that sets your bill. Astra is not the fastest per token, but it needs about a third of the output tokens Grok 4.7 uses per task, and under a quarter of Opus 5.5's at max effort, which is why it has the lowest cost per task at its level on the index. Opus 5.5 narrows that gap at lower effort settings: Artificial Analysis finds four of its five effort levels on the best-value frontier among models scoring 50 or more.
Hallucination and accuracy
AA-Omniscience asks 6,000 hard factual questions and tracks two things separately: how often the model is right, and how often it makes something up when it does not know. A model can post a low hallucination rate simply by refusing to answer, so read the two columns together.
| Model | Accuracy (higher is better) | Hallucination rate (lower is better) | Read as |
|---|---|---|---|
| Claude Fable 5.1 | 67% | not in the published snapshot | Highest accuracy in the published snapshot; answers more, so hallucination will not be low |
| Claude Fable 5 | 65% | 63.6% | Very accurate, confidently wrong when it misses |
| Claude Opus 5 | ~61% | 60.8% | First Opus that answers rather than declines |
| GPT-6 Astra | ~55% | 51% | Halved Sol's hallucination rate without refusing more |
| Gemini 3.1 Pro | 54.9% | 50.9% | Balanced |
| Grok 4.7 | 47% | 29% | Lowest measured rate of this month's frontier releases, down from 34% on Grok 4.6 |
| Gemini 3.8 Flash | — | 55.2% | — |
| Kimi K3 | 46–48% | ~53% | Regressed from K2.6 |
| Claude Opus 4.8 | 46.6% | 35.9% | Best-calibrated previous-generation flagship: knows what it doesn't know |
| Grok 4.3 (medium) | — | 16% | Lowest rate of any flagship-class model |
| GPT-5.6 Sol | ~51% | 92.2% | Almost never admits ignorance |
| DeepSeek V4 Flash (July build) | 37% | 84% | Strong coder, still overconfident on facts, though 11 points better than the April build |
On the combined AA-Omniscience Index, which rewards right answers and penalises invented ones, GPT-6 Astra (44, at high effort) sat one point ahead of Fable 5.1 (43) before Opus 5.5 arrived. Artificial Analysis now lists Claude Opus 5.5 as the leader on AA-Omniscience, but its separate accuracy and hallucination figures were not in the launch summary. The very lowest hallucination rates on the board belong to small models that decline most questions. That is not a model you want; it is a model that has learned to say "I don't know". Also worth knowing: reasoning modes hallucinate two to three times more than plain modes on summarisation tests. If the facts matter, turn thinking off or give the model a document to work from.
Refusals and over-blocking
This is the category businesses ask about most and the one with the least data. No leaderboard measures "refused a perfectly reasonable request" across vendors, so the table below collects what has actually been published.
| Model | What is known | Figure |
|---|---|---|
| Claude Opus 5.5 | The first Opus to launch with Fable 5.1-class safeguards on cybersecurity, biology and distillation, each falling back to another model transparently: most cyber tasks go to Opus 4.8. Fixing bugs in your own code is still allowed, and vetted organisations can apply for fuller access through Anthropic's cyber and life-sciences verification programmes. | qualitative |
| Claude Fable 5.1 | Anthropic says its new safeguards block 60% fewer false positives than Fable 5, with around 60% fewer cyber interventions per session, and biology safeguards firing 85% less often on benign medical questions. Pentest, exploit and binary-scan requests still go to Opus 4.8. | −60% / −85% vs Fable 5 |
| Claude Fable 5 | Fell back to Opus on roughly 18% of AutomationBench-AA tasks and 9% of AA-Omniscience questions — the only absolute block rate any vendor's routing has produced. | 9–18% of tasks |
| Older Claude (4.6) | OR-Bench over-refusal on a benign health-robotics set: Opus 4.6 33.5%, Sonnet 4.6 25.9%. Claude has historically been the safest and the most over-cautious. | 26–34% |
| GPT-6 Astra | Public build declines advanced offensive cyber work; wider access through the Daybreak programme for vetted organisations. Alignment testing: 0% unauthorised task completion. | qualitative |
| Gemini | Cyber variant restricted to the Fairwind programme. Older Gemini 2.5 Flash showed 46.3% over-refusal on the same health set. | 46% (older gen) |
| Mistral | Accepts most prompts on OR-Bench; consistently the lowest over-refusal of the majors. | low |
| Open-weight models | Independent bio-research testing found benign-tier over-refusal of 91.5% for Kimi K2.6, 76.6% for Claude Opus 4.7 and 57.9% for GPT-5.5 — and that a model's overall refusal rate is a poor predictor of how sensibly it refuses. | 58–92% (bio prompts) |
A model that refuses to look at malware is no use to the people whose job is looking at malware.
Image generation
Image models are ranked by blind human votes in the Artificial Analysis arena: two images from the same prompt, pick the better one, thousands of times. Elo is the native unit; higher is better and 30 points is a meaningful gap. OpenAI's September release, GPT Image 2.5, now holds the top two places both for creating images and for editing them.
| Model | Maker | Text-to-image Elo | Image editing | Best at | Price per 1,000 images |
|---|---|---|---|---|---|
| GPT Image 2.5 Flare | OpenAI | 1188 (#1) | #2 | The quicker 2.5 tier; photorealism and text in images | $211 |
| GPT Image 2.5 Sunburst | OpenAI | 1182 (#2) | #1 | The slower, more detailed tier; the best editor on the board | $211 |
| GPT Image 2 | OpenAI | 1171 (#3) | #5 | Plans the layout before drawing; strong typography | $211 |
| Grok Imagine Image 2.0 | SpaceXAI | 1154 (#4) | outside top 5 | New in August; top-five quality at under a third of OpenAI's price | $60 |
| MAI-Image-2.6 | Microsoft | 1147 (#5) | #3 | Editing, product and branded design | $39 |
| Reve 2.1 | Reve | 1129 (#6) | outside top 5 | Previously led image editing; now overtaken | $200 |
| Nano Banana 2 (Gemini 3.1 Flash Image) | 1122 (#7) | outside top 5 | Fast, low-cost everyday image model | $67 | |
| Muse Image | Meta | 1111 (#8) | — | Cheapest model in the top ten | $10 |
| Nano Banana Pro (Gemini 3 Pro Image) | 1100 (#11) | outside top 5 | True 4K output, multilingual text, identity lock across edits | $134 | |
| Ideogram 4.0 (Quality) | Ideogram | 1012 (top open-weight) | — | Typography and text-heavy graphics | $100 |
| FLUX.2 [dev] | Black Forest Labs | 1000 | — | Runs locally; best open-weight option for photorealism | $12 |
| Midjourney V8.1 | Midjourney | not on arena | — | Aesthetic and creative ideation | subscription |
Prices are Artificial Analysis's figure for 1,000 images at 1024×1024 on each maker's API. For diagrams, technical illustrations and anything with words in it, no benchmark measures "diagram quality" directly; text-rendering ability is the proxy, and GPT Image 2 and 2.5, Ideogram 4.0 and Nano Banana Pro lead it. For architecture diagrams you will usually get a better result asking a language model for SVG or Mermaid than asking an image model to draw one.
Video generation
Video models are ranked by blind votes in the Artificial Analysis Video Arena, and the board has turned over since the start of September. Google's Gemini Omni Flash now leads for clips with sound; the rest of the top five are built on Chinese models from Alibaba, MiniMax and ByteDance. Kling, which topped several leaderboards earlier in the year, has slipped to twelfth.
| Model | Maker | Text-to-video Elo (with audio) | Image-to-video Elo (with audio) | Output and strengths | API price per minute |
|---|---|---|---|---|---|
| Gemini Omni Flash | 1233 (#1) | 1177 (#3) | Conversational editing and clip extension; 1080p and 4K by upscaling | $6.00 | |
| Wan 3.0 | Alibaba | 1229 (#2) | 1164 | Up to 30-second clips; also leads text-to-video without audio | $12.00 |
| MiniMax H3 Max | fal, built on MiniMax H3 | 1227 (#3) | 1195 (#1) | Top-three quality at the lowest price in the top ten | $2.40 |
| MiniMax H3 | MiniMax | 1220 (#4) | 1181 (#2) | Best open-weight video model | $7.80 |
| Seedance 2.0 | ByteDance | 1210 (#5) | 1174 (#5) | Native audio, multi-shot, up to 12 reference files | $9.07 |
| Seedance 2.5 | ByteDance | not yet ranked | not yet ranked | 30 seconds per pass, extendable; up to 50 references; timestamp and region edits | Higgsfield: ~$3.50 per 10s at 720p |
| HappyHorse 1.1 | Alibaba | 1147 | 1106 | 7-language lip-sync | $9.90 |
| Grok Imagine Video 1.5 | SpaceXAI | not ranked | 1099 | Tuned for animating stills | $8.40 |
| Kling 3.0 Pro (1080p) | Kling AI (Kuaishou) | 1095 (#12) | 1055 | Native 4K, 15-second multi-shot scenes, multilingual lip-sync, motion control | $20.16 |
| Veo 3.1 | 1088 | 1082 | Up to 4K with synchronised dialogue; cheaper Fast ($9.00) and Lite ($4.80) tiers | $24.00 | |
| Runway Gen-4.5 | Runway | not on current board | — | 4K; best motion brushes and scene control for film work | subscription / API |
| Sora 2 | OpenAI | withdrawn | withdrawn | App closed 26 April; API ends 24 September 2026 | — |
Elo scores are for clips with audio, as of 22 September. Without audio the order changes: Wan 3.0 leads text-to-video (1336) and Gemini Omni Flash leads image-to-video (1369). Prices are Artificial Analysis's figure for one minute of video on the maker's own API, except Seedance 2.5, which uses Higgsfield's credit price.
Seedance 2.5 and Higgsfield
ByteDance released Seedance 2.5 on 31 July. It generates up to 30 seconds with sound in a single pass, double Seedance 2.0, and can extend a clip in further rounds while keeping the same characters and setting. One generation can draw on up to 30 images, 10 video clips and 10 audio clips as references, and edits can target a single moment or region instead of re-rendering the whole clip.
It launched in China on Jimeng and Doubao, and outside China it is sold through creative platforms including Higgsfield, Dreamina, Runway, OpenArt and Magnific. Higgsfield, a San Francisco company that bundles more than 50 image, video and audio models under one subscription, is the version most people are talking about, because it wraps the model in direct controls for era, lens, lighting, physics and camera path. It charges 70 credits (about $3.50) for a 10-second clip at 720p and 120 credits (about $6) at 1080p; the model generates up to 1080p, with 4K by upscaling.
There is no blind-vote score yet, and ByteDance itself says complex motion and scenes with several interacting people still need work. Seedance 2.0 also drew legal complaints from Disney and Paramount over famous characters, so keep real people, brands and film characters out of your prompts.
Kling: capable, but no longer on top
Kling 3.0, launched in February, is still Kuaishou's flagship, with a faster Turbo version and an upgraded Omni model added on 17 June. Its strengths are unchanged: native 4K, up to 15 seconds across several shots, multilingual lip-sync and precise motion control. On the blind vote, though, the 1080p Pro model now sits twelfth for text-to-video with sound, and at $20.16 a minute on Kling's own API it costs more than all but one of the models ranked above it.
The business is growing fast. Kling AI was spun out of Kuaishou in July with a $3 billion funding round at an $18 billion valuation, backed by investors including Tencent, Alibaba Cloud and Baidu, and around three-quarters of its revenue comes from outside China. A Kling 4.0 is widely expected but has not been announced, so treat any site already selling "Kling 4" access with suspicion.
Seedance and Kling both come from Chinese companies, so prompts and uploads may fall under Chinese law, just as material sent to a US provider falls under US law. For client footage or unreleased products, check which company processes the data, and where, before you upload anything.
What people actually run
Everything above ranks models by how they score. This ranks them by how much work they are actually given. OpenRouter is a neutral exchange that routes requests across several hundred models from every major lab and publishes its own token volumes, so it is the closest thing to a usage leaderboard that exists. Here is where the tokens went in the seven days to 20 September 2026.
Volumes are prompt plus completion tokens for the seven days to 20 September 2026, so they predate the MiMo-V2.6 and Claude Opus 5.5 launches. OpenRouter lists variants of the same model separately, which is why DeepSeek V4 Flash appears twice. Source: OpenRouter (openrouter.ai/rankings), licensed under CC BY 4.0.
Read that against every ranking above and the mismatch is close to total. Only one model this page names as a category leader appears: DeepSeek V4.1 Flash, top of the CyberGym vulnerability-discovery test, which went from seventh place to first within ten days of its release. Claude Fable 5.1, GPT-6 Astra and Claude Opus 5 are nowhere near the top ten. Eight of the ten entries are Chinese; the two American ones are OpenAI's budget tier and a free Nvidia model. Counted by requests rather than tokens, the week beginning 14 September split DeepSeek 25.4%, Google 18.6%, OpenAI 17.0% and Z.ai 9.4%, with Anthropic on 2.7%.
Benchmarks measure the ceiling. Usage measures the floor, and almost all the work happens on the floor.
The pattern behind it is straightforward once stated. Almost every model in the top ten is open-weight, free or priced near the floor, and an exchange like this accumulates whatever is cheapest to serve. The frontier labs' pricing keeps their best models off it by design. That is worth knowing when you are choosing a stack: the models that win the tables above are bought for the hardest tenth of the work, and something like this list is what handles the rest.
What the work actually is
OpenRouter also classifies requests by task and reports the share of spend each takes. In the week to 12 September the split was a useful corrective to a page organised around medical, legal and scientific benchmarks.
| Category | Share of spend | What sits inside it |
|---|---|---|
| General | 31.8% | Classification, question answering, content writing, conversation, customer support, summarising |
| Code | 30.8% | Code generation, debugging, file operations, shell execution, code review, frontend and UI |
| Agent | 28.2% | Workflow execution, multi-step planning, tool dispatch |
| Data | 9.2% | Data extraction and transformation |
Agent work is already more than a quarter of what people pay for, on a par with writing code. None of the professional benchmarks on this page measure it, which is why the tool-use and computer-use sections matter more than their position in the page suggests. The categories this page covers in most depth, medical and legal, do not appear in the spend breakdown at all.
One reason any of this can be wrong
For two weeks in August 2026 an anonymous listing called Ox Alpha carried 27 trillion tokens through the exchange, 12% to 14% of all traffic, with no lab attached to it. Labs routinely run unreleased models under codenames for a week or two before launch; Z.ai later confirmed that Ox Alpha was GLM-5.3-Flash and published the weights. So a usage table like the one above can be materially wrong about who is winning, for weeks at a time, while being exactly right about the totals.
Why everyone is talking about Kimi
Kimi is the assistant and model family from Moonshot AI, a Beijing start-up founded in 2023. Three things have put it in the headlines since July: a model that rivals the US frontier and can be downloaded, a business growing unusually fast, and a public accusation from a rival.
A frontier-class model you can download
Moonshot released Kimi K3 on 16 July and published the full weights on 27 July: 2.8 trillion parameters with 104 billion active per token, a one-million-token context window and native vision. At launch it took first place in the Frontend Code Arena, ahead of Claude Fable 5 and GPT-5.6 Sol, although Moonshot itself conceded that K3 still trailed those two models overall. The release landed in Washington as a question about how quickly a Chinese start-up had closed the gap. A US congressional committee had been pressing American companies over their use of Kimi since April, and in March Cursor admitted that its Composer 2 coding model was built on an earlier Kimi release.
It is not a bargain model. Moonshot charges $3 per million input tokens and $15 per million output, five times the input price of its predecessor, K2.6, although K3 is free to use in the Kimi app. The weights come under Moonshot's own licence rather than MIT: free for most businesses, but model-hosting companies with more than $20 million in annual revenue need a separate agreement.
A fast-growing business
Bloomberg reports that Moonshot is aiming for $2 billion in annualised revenue by the end of the year, double its August run rate, and OpenRouter shows K3 models generating up to 300 billion tokens a day. On 17 September Moonshot launched Kimi for financial services, wired into data from S&P Global Market Intelligence, Crunchbase, the US SEC's EDGAR filings, the IMF, the World Bank and the Federal Reserve's FRED database. It has reportedly filed confidentially for a Hong Kong listing.
The allegation
On 10 September Anthropic, which makes Claude, published a threat-intelligence report accusing Moonshot of more than copying. It alleged that in one ten-day window Moonshot relayed nearly 300,000 Kimi customer requests to Claude, mostly to Claude Opus, through 5,380 fraudulent accounts that appeared to be in Singapore and Japan; showed Claude's answers to users as if Kimi had written them; and kept the exchanges to train its own models, as part of a campaign in which Anthropic says more than 23 million responses were collected. The same report made distillation allegations against DeepSeek and MiniMax.
What it means if you are weighing Kimi up
What it costs to use them yourself
API prices are in the model table above. For individual chat subscriptions, the whole market has settled on a $20-a-month standard tier, with budget tiers below it and power-user tiers of $100 to $300 above it. Prices are the vendors' US list prices; UK checkouts are billed in sterling with VAT added, so check the vendor page for the exact figure on the day.
| Service | Free tier | Budget | Standard | Power user | What the standard tier gets you |
|---|---|---|---|---|---|
| ChatGPT | Yes, GPT-5.6 Luna, unlimited text; ads in some markets | Go $8 | Plus $20 | Pro $100 / $200 | GPT-5.6 Sol in chat; GPT-6 Astra only inside ChatGPT Work and Codex (Astra in ordinary chat needs Pro) |
| Claude | Yes, daily caps | — | Pro $20 ($17 annual) | Max $100 (5×) / $200 (20×) | Opus 5.5 included, with five-hour usage limits raised at its launch; Fable 5.1 only on pay-as-you-go usage credits (Max includes it for up to half the weekly allowance) |
| Gemini | Yes, 3.6 Flash plus limited 3.1 Pro | AI Plus $4.99 | AI Pro $19.99 | AI Ultra $99.99 / $199.99 | 3.1 Pro with 1M context, plus 3.8 Flash in the app; Ultra was cut from $249.99 at I/O 2026 |
| Grok | Yes, Grok 4.6 with tight limits | X Premium $8 / SuperGrok Lite $10 | SuperGrok $30 | SuperGrok Plus $100 / Heavy $300 | Grok 4.6 and Grok Bot; Grok 4.7 launched on the API, Cursor and Grok Build first, with no app date yet |
| Kimi | Yes, K3 is free in the app | — | from about $19 | up to about $199 | Kimi K3, with context length tiered by plan |
| DeepSeek | Yes, web chat is free | — | no monthly plan | — | Everything else is per-token API ($0.15 / $0.60 off-peak for V4.1 Flash), or self-hosted |
Two things to check before relying on a consumer plan for work. Training terms differ: xAI's individual Grok plans may use your conversations for training while Grok Business ($30 per seat) does not, and Meta's cheapest Muse Spark API tier is cheap precisely because Meta may train on what you send. Availability differs too: some Gemini app features have excluded the UK and EEA at launch, and Claude Mythos 5.1 is not sold on any public plan. Business and Enterprise tiers from every vendor add admin controls, data-processing terms and, in most cases, EU or UK data residency.
Running an AI agent on your own PC
Everything above runs in someone else's data centre. A local agent keeps prompts and files on your own machine, in exchange for a much smaller model. It takes two pieces: a model runner such as Ollama or LM Studio, and an agent layer that plans and uses tools around the model. OpenClaw suits a persistent personal assistant, OpenHands and Cline suit software engineering, and goose suits general local automation. Hermes Agent, from Nous Research, learns reusable skills from each completed task and ships tool-call parsers for local models, so a small Qwen behaves predictably.
Memory decides almost everything. A 4-bit (Q4) model needs roughly 0.6 GB of video memory per billion parameters, before you add room for the conversation. The newest mixture-of-experts models help here: they hold many parameters in memory but only work a few billion of them per word, so they run far faster than their size suggests.
What tokens per second actually feels like
Models write in tokens, not words. In English prose a token is about three-quarters of a word, so 100 tokens is roughly 75 words; code uses more tokens per line because of symbols and indentation. An average adult reads silently at about 238 words a minute, which is only around 5 tokens per second. Anything above 10 or 15 tokens per second therefore outruns your eyes on a normal answer. Speed starts to matter when the model writes far more than you read: reasoning models think privately before answering, and coding agents write whole files and retry until the tests pass. Drag the slider or pick a machine to see the difference.
A token is a chunk of text, usually part of a word. This paragraph is being written at the speed you picked, so you can see how it feels in practice. Above about fifteen tokens per second it outruns most readers, and a short answer like this one feels instant for a short answer. It feels very different when a coding agent has to write a two-hundred-line file, run the tests, read the errors and write the file again, because every one of those steps is paid for in tokens. The faster the machine, the less time you spend waiting for work you will never read.
| What the model writes | Size | Time at this speed |
|---|---|---|
| A short answer to a question | ~150 words, ~200 tokens | 7 sec |
| A detailed, one-page answer | ~600 words, ~800 tokens | 27 sec |
| A 150-line code file | ~1,500 tokens* | 50 sec |
| A reasoning model's hidden thinking, then a short answer | ~2,200 tokens | 1 min 13 sec |
| One agent task, at GPT-6 Astra's average output | ~27,000 tokens | 15 min |
Times assume the speed above holds for the whole answer and ignore the wait before the first word appears. *Assumes about 10 tokens per line of code. The reasoning example uses 2,000 thinking tokens before a 200-token answer; the agent example is the ~27k output tokens Artificial Analysis measured for GPT-6 Astra per index task, against ~119k for Claude Opus 5.5 at max effort.
With that in mind, here is what each build can realistically run. Every PC row assumes the same base machine: an Intel Core Ultra 7 with 64 GB of RAM. The coding level column uses our own four-step scale, explained after the tables.
The Coding vs Opus 5 column shows how close the best coding model each machine runs at a usable speed comes to Claude Opus 5, on one shared yardstick: the Artificial Analysis Coding Index, where Opus 5 scores 78.0. The best coder that fits on a single 24 GB graphics card, Qwen 3.8 27B, scores 68.1, or 87% of Opus 5. The best that any machine on this page holds, Qwen3.8-Flash-Next at about 123 GB, scores 73.0, or 94%. Claude Opus 5.5 was released today and is not on the index yet; it beats Opus 5 on every coding test Anthropic published, and by 11 points on the independent Terminal-Bench 4.0 run, so the gap to Opus 5.5 is wider than these bars show.
| Build | Best all-round model | Best for agents and coding | What to expect | Coding level | Coding vs Opus 5 |
|---|---|---|---|---|---|
| PC, no graphics card | Gemma 4 26B-A4B or Qwen 3.6 35B-A3B, both mixture-of-experts with 3–4B parameters active per word | Qwen 3.6 35B-A3B for coding; gpt-oss 20B for the cleanest tool calls | Workable for private drafting and summarising, too slow for coding agents. A 64 GB machine with no GPU runs a dense 14B model at about 10–15 tokens per second in published tests. gpt-oss 120B will not fit: its weights alone are 63.5 GB. | 1 | 54% |
| PC with an RTX 4070 (12 GB) | Qwen 3.5 9B at Q6 (about 9 GB, around 30 tokens per second) or Gemma 4 12B at Q4 (about 7 GB) | Gemma 4 12B for coding; gpt-oss 20B loads at Q4 but leaves almost no room for context | Snappy chat and light tool use. Anything 14B or larger spills into system RAM and drops to around 7 tokens per second. | 1–2 | 40% |
| PC with an RTX 4090 (24 GB) | Gemma 4 26B-A4B at Q4: about 16 GB, around 85 tokens per second, 256K context | Qwen 3.8 27B (about 17 GB at Q4), the strongest coder that fits one card. Qwen 3.6 27B is faster per token, and Poolside's Laguna XS 2.1 is a lighter agentic option | The sweet spot, and the first tier where a local coding agent is genuinely useful. Skip 70B dense models: they do not fit at Q4. | 3 | 87% |
| PC with 2× RTX 3090 (48 GB in total) | Gemma 4 31B or Qwen 3.6 27B at higher precision with a longer context | Qwen 3.8 27B for coding on one card, with an assistant model on the other | Each 3090 runs Qwen 3.6 27B at about 35 tokens per second. 70B models fit at Q4, but the newer 27–31B models outscore them. Use vLLM for NVLink-aware splitting of one model across both cards. Still short of the 60 GB gpt-oss 120B wants. | 3 | 87% |
| PC with 2× RTX 5090 (64 GB in total) | Qwen 3.6 35B-A3B at Q6 on one card (about 28 GB, around 80 tokens per second) with Gemma 4 31B on the other | Qwen 3.8 27B at high precision; Mistral Small 4 (119B) fits across both cards but codes worse | The fastest consumer build, but more memory is not more intelligence: Mistral Small 4 scores 43 on BenchLM's leaderboard against 55 for Gemma 4 31B, which fits on one card. | 3 | 87% |
Apple Mac Studio and MacBook Pro
Macs work differently. The processor and graphics share one pool of unified memory, so a Mac can load models far larger than any consumer graphics card holds. The trade-off is speed: writing each token is limited by memory bandwidth, where Nvidia cards are faster, while reading your prompt is limited by raw compute, where the M5 generation made its biggest jump. In Apple's own measurements, M5 produced the first token 3.3 to 4 times faster than M4 but subsequent tokens only 19–27% faster. For coding agents, which send long prompts full of code, that first-token speed is the one you feel.
Two supply notes. There has never been an M3 Max Mac Studio: that chip shipped only in the MacBook Pro, so it is listed here as a laptop. And the memory shortage has hit Macs hard: Apple removed the 512 GB M3 Ultra option in March 2026, and by May the M3 Ultra Mac Studio was sold only with 96 GB. The M5 Max and M5 Ultra Mac Studio, announced on 26 August, restore the 512 GB tier.
| Mac | Memory and bandwidth | What it can run | Speed | Closest cloud equivalent | Coding level | Coding vs Opus 5 |
|---|---|---|---|---|---|---|
| M3 Max (MacBook Pro only) | up to 128 GB, 400 GB/s | Qwen 3.6 27B and Gemma 4 31B comfortably; gpt-oss 120B squeezes in with 128 GB | Slow on dense models: the M1 Max, which has the same 400 GB/s, measured 15.5 tokens per second on Qwen 3.6 27B | Close to early-2026 flagships on standard coding tests; well behind today's frontier on hard tasks | 2–3 | 87% |
| M4 Max (Mac Studio 2025) | up to 128 GB, 546 GB/s | gpt-oss 120B (about 65 GB) and 122B-class mixture-of-experts models | 16.6 tokens per second in a published Q4 run of Qwen 3.6 27B; far faster on mixture-of-experts models | As above; the larger MoE models add knowledge, not frontier reasoning | 3 | 87% |
| M3 Ultra (Mac Studio 2025) | 96 GB (256 and 512 GB withdrawn), 819 GB/s | At 96 GB, slightly less room than a 128 GB M4 Max; used 256–512 GB units can hold 235B-class models | 28.6 tokens per second measured on Qwen 3.6 27B via Ollama | As above, with the fastest dense-model speeds of the 2025 Macs | 3 | 87% |
| M5 Max (Mac Studio 2026) | up to 128 GB, 614 GB/s | gpt-oss 120B; DeepSeek V4 Flash (284B) at 2-bit | 76–88 tokens per second measured on gpt-oss 120B in MLX (on an M5 Max MacBook Pro); about 39 on DeepSeek V4 Flash at 2-bit (community figure) | As the M4 Max, but noticeably faster: the most practical big-model Mac | 3 | 87% |
| M5 Ultra (Mac Studio 2026) | 96–512 GB, 1.2 TB/s | At 512 GB (about 435 GB usable): Llama 4 Maverick 400B, several 120B models at once, and by our sizing arithmetic DeepSeek V4.1 Flash at Q4 | Estimates: 20–25 tokens per second on a 70B dense model, 70–100 on a mixture-of-experts model with 20B active parameters | With DeepSeek V4.1 Flash loaded: roughly level with GPT-5.6 Luna and Gemini 3.8 Flash on the intelligence index | 3 | 94% |
M5 Ultra prices start at $5,499 with 96 GB and run to $18,299 fully loaded; the 96 and 256 GB models ship from 22 September and the 512 GB model in late October. M5 Ultra speeds are published estimates, not measurements, until shipped units are benchmarked. Measured figures depend on the runtime (MLX, Ollama, llama.cpp), quantisation and context length, so treat them as a guide.
Dedicated AI computers
A new class of machine sits between the gaming PC and the datacentre. Some copy Apple's approach of one large pool of shared memory, but use AI chips instead of desktop processors. Others stack professional graphics cards with 96 GB each. All of them run the same open-weight models as the builds above. What changes is how big a model fits, how fast it runs, and how many people it can serve at once.
| Machine | Memory and bandwidth | What it can run | Coding level | Coding vs Opus 5 | Estimated cost (US) | Status |
|---|---|---|---|---|---|---|
| NVIDIA DGX Spark deskside appliance | 128 GB unified, 273 GB/s. GB10 Grace Blackwell chip with a 20-core Arm CPU | NVIDIA quotes models up to about 200B parameters, or about 405B with two units linked. Its sweet spot is mixture-of-experts models such as gpt-oss 120B; dense 70B models run slowly on this bandwidth | 3 | 87% | $4,699 (launched at $3,999) | On sale. Runs NVIDIA's Linux-based DGX OS, not Windows |
| HP ZGX Nano G1n deskside appliance | 128 GB unified. Same GB10 chip as the DGX Spark | The same as the DGX Spark | 3 | 87% | $6,499–7,399 | On sale; the dearest GB10 box reviewed so far |
| HP Z2 Mini G1a x86 mini workstation | 128 GB unified, 273 GB/s. AMD Ryzen AI Max+ PRO 395 | gpt-oss 120B on its integrated Radeon graphics, no discrete card. The same chip runs gpt-oss 120B at about 30 tokens per second. Runs normal Windows 11 Pro | 3 | 87% | About $3,300–4,900; prices have climbed with memory costs | On sale. AMD's next chip, the Ryzen AI Max Pro 400, raises the ceiling to 192 GB |
| HP OmniBook Ultra 16 and X 14 NVIDIA RTX Spark laptops | Up to 128 GB unified memory, up to 1 petaflop of FP4 AI compute | The same model sizes as a DGX Spark, on battery, in a Windows laptop | 3 (expected) | 87% | Not announced | Announced 4 September; due this autumn. An always-on HP OmniDesk desktop follows |
| Dell Precision 7875 multi-GPU tower | 2× RTX PRO 6000 Blackwell Max-Q: 192 GB of graphics memory at 1.8 TB/s per card | 120B-class models at high precision with fast replies; DeepSeek V4 Flash (284B) at Q4 by our sizing arithmetic | 3 | 94% | About $40,000–55,000 (our estimate: each card now costs $14,000–16,000) | On sale with Threadripper PRO 9000 processors |
| HP Z8 Fury G6i multi-GPU tower | Up to 4× RTX PRO 6000 Max-Q: 384 GB of graphics memory | DeepSeek V4.1 Flash at Q4 held entirely in fast graphics memory (by our arithmetic), fast enough to serve a team | 3 | 94% | About $70,000–90,000 fully loaded (our estimate; HP publishes no price) | On sale |
| MSI XpertStation WS300 and HP ZGX Fury GB300 deskside station | 748 GB coherent: 252 GB of HBM3e at 7.1 TB/s plus 496 GB of LPDDR5X | Trillion-parameter models at FP4, so the top open model, MiMo-V2.6-Pro, fits. The only machines on this page that can run it | 3 | 94% | $85,000–99,999 for the MSI; HP has not published a price | MSI shipping; HP orderable |
| AMD Threadripper Halo Station x86 accelerator workstation | 96-core Threadripper PRO, 2 TB DDR5, 2× Instinct MI350P, each with 144 GB of HBM3E at 4 TB/s (288 GB in total; 576 GB with four) | AMD claims trillion-parameter models. That is a capacity claim; expect it to be usable for experiments rather than fast serving | 3 | 94% | No price. Parts alone exceed $100,000; a finished system could pass $150,000 | Prototype shown at IFA on 4 September; no date |
Every machine here tops out at coding level 3. Even the $85,000 GB300 stations run the best open model, which scores 46 on the intelligence index, against 58 for Claude Opus 5.5 in the cloud. What the expensive machines buy is capacity and speed for a whole team, with nothing leaving the building, rather than frontier intelligence. Bandwidth decides single-user speed: the GB10 and Ryzen AI Max boxes share the same 273 GB/s, about a third of an M3 Ultra, so they favour mixture-of-experts models over dense ones.
What each option costs
Prices are US list or street prices seen in August and September 2026, or our own estimates where no price has been published. Memory and graphics card shortages have pushed almost every figure up this year, so check the day's price before you buy. UK checkouts are in sterling with VAT added. Electricity, software and support are not included. The two new columns show the ceiling of each option: the strongest model it can hold, with its Artificial Analysis Intelligence Index score where one exists, the coding level from our four-step scale, and the fastest published speed. Speeds are for different models, so compare them with care: a fast figure on a small model is not the same as a fast figure on a large one. All PC rows share the Core Ultra 7, 64 GB base. The Coding vs Opus 5 column is explained above the first local table.
| Option | Estimated cost | Max intelligence (best model it can hold) | Coding vs Opus 5 | Max speed (fastest published figure) | Basis |
|---|---|---|---|---|---|
| PC, no graphics card | About $1,800–2,500 | Gemma 4 26B-A4B or Qwen 3.6 35B-A3B. Coding level 1 | 54% | 10–15 tok/s (dense 14B) | Our estimate. A 64 GB DDR5 kit alone now costs $769–929 |
| PC with an RTX 4070 (12 GB) | About $2,300–3,200 | Qwen 3.5 9B or Gemma 4 12B. Level 1–2 | 40% | ~30 tok/s (Qwen 3.5 9B) | Our estimate. The card is discontinued, so remaining or used stock |
| PC with an RTX 4090 (24 GB) | About $3,200–6,500 | Qwen 3.8 27B: index 34. Level 3 | 87% | ~85 tok/s (Gemma 4 26B-A4B) | A used 4090 costs $1,400–2,000, a new one $3,700–4,300 |
| PC with 2× RTX 3090 (48 GB) | About $4,000–5,500 | Qwen 3.8 27B at higher precision: index 34. Level 3 | 87% | ~35 tok/s (Qwen 3.6 27B) | Used 3090s cost $1,000–1,300 each; allow for a larger power supply |
| PC with 2× RTX 5090 (64 GB) | About $9,000–10,500 | Qwen 3.8 27B at high precision: index 34. Level 3 | 87% | ~80 tok/s (Qwen 3.6 35B-A3B) | About $3,500 per card, plus a workstation-class power supply |
| MacBook Pro, M3 Max, 128 GB | Resale only | Qwen 3.8 27B; gpt-oss 120B fits. Level 2–3 | 87% | ~15 tok/s (Qwen 3.6 27B, est.) | Discontinued; no reliable current price |
| Mac Studio, M4 Max, 128 GB | About $4,000–7,000 | Qwen 3.8 27B; gpt-oss 120B and 122B-class models. Level 3 | 87% | 16.6 tok/s (Qwen 3.6 27B); faster on MoE | Replaced by M5; €4,399 at a European retailer, $4,000–7,000 on US resale during the shortage |
| Mac Studio, M3 Ultra, 96 GB | $3,999 list | Qwen 3.8 27B; gpt-oss 120B. Level 3 | 87% | 28.6 tok/s (Qwen 3.6 27B) | Replaced by M5; larger memory options were withdrawn |
| Mac Studio, M5 Max | From $2,499 (£2,499) | Qwen 3.8 27B; DeepSeek V4 Flash at 2-bit (128 GB). Level 3 | 87% | 76–88 tok/s (gpt-oss 120B) | Apple list price for 36 GB; the 128 GB model costs more |
| Mac Studio, M5 Ultra | $5,499–18,299 | DeepSeek V4.1 Flash at 512 GB: index 39. Qwen3.8-Flash-Next for coding. Level 3 | 94% | 70–100 tok/s (MoE, 20B active, est.) | Apple list: 96 GB to fully loaded 512 GB |
| HP Z2 Mini G1a, 128 GB | About $3,300–4,900 | Qwen 3.8 27B; gpt-oss 120B. Level 3 | 87% | ~30 tok/s (gpt-oss 120B) | Retail listings |
| NVIDIA DGX Spark | $4,699 | Qwen 3.8 27B; models up to ~200B. Level 3 | 87% | ~30 tok/s (gpt-oss 120B, est. from matching bandwidth) | NVIDIA US store, August 2026 |
| HP ZGX Nano G1n | $6,499–7,399 | As DGX Spark. Level 3 | 87% | As DGX Spark | HP direct, 2 TB to 4 TB |
| HP OmniBook RTX Spark laptops | Not announced | As DGX Spark (up to 128 GB). Level 3 expected | 87% | Not published | Due this autumn |
| Dell Precision 7875, 2× RTX PRO 6000 | About $40,000–55,000 | Qwen3.8-Flash-Next (~123 GB) for coding. Level 3 | 94% | Not published; 1.8 TB/s per card | Our estimate from card prices |
| HP Z8 Fury G6i, 4× RTX PRO 6000 | About $70,000–90,000 | DeepSeek V4.1 Flash: index 39. Qwen3.8-Flash-Next for coding. Level 3 | 94% | Not published; 1.8 TB/s per card | Our estimate from card prices |
| MSI XpertStation WS300 (GB300) | $85,000–99,999 | MiMo-V2.6-Pro: index 46, the best open model. Level 3 | 94% | Not published; 7.1 TB/s HBM3e | MSRP and retail listings |
| AMD Threadripper Halo Station | $100,000–150,000+ | Trillion-parameter models (AMD claim). Level 3 | 94% | Not published; 4 TB/s per accelerator | Press estimates from component prices; AMD has set no price |
| For comparison: a cloud plan | $20 a month | Claude Opus 5.5: index 58. Level 4 | 100%+ | ~262 tok/s (Gemini 3.8 Flash API); 71.4 (GPT-6 Astra) | List price; Artificial Analysis measurements |
For scale: $4,699, the price of a DGX Spark, would pay for a $20-a-month Claude or ChatGPT plan for about 19½ years, or buy about 235 million Claude Opus 5.5 output tokens at $20 per million. Local hardware pays back through privacy, predictable cost and very high volume, not by being cheaper for occasional use.
The best setup for your budget
Pulling the tables together, here is what we would buy at each price point today, and what is worth waiting for. Prices are US prices in September 2026. "A PC you already own" assumes a desktop with a free graphics slot and a power supply big enough for the card. The Coding vs Opus 5 bar is the same yardstick used in the tables above.
| Budget | Best pick | What it gets you | Coding vs Opus 5 | Alternatives | Coming later this year |
|---|---|---|---|---|---|
| Under $1,000 | Add an RTX 5060 Ti 16 GB to a PC you already own: about $670–805 now (it launched at $429) | Runs Gemma 4 12B and gpt-oss 20B comfortably, and a 27B model squeezed to 3-bit if you accept some quality loss. Coding level 1–2 | 40% | No PC? A Mac mini M6 with 16 GB ($899) runs the same class of model, quietly, but on slower memory (153 GB/s) | Card prices are rising, not falling, so buy when you need it |
| $1,000–2,000 | Add a used RTX 3090 (24 GB) to a PC you already own: $1,000–1,300 | The cheapest route to Qwen 3.8 27B, the strongest coder that fits on one card, at around 35 tokens per second on 27B models. Level 3 | 87% | Starting from scratch: a Mac mini M6 with 32 GB ($1,299) runs the same model, but slowly (170 GB/s). A used RTX 4090 ($1,400–2,000) is about 40% faster than a 3090 | The Mac mini M5 Pro starts at $1,699, but its 24 GB base is tight; 48 GB costs $2,299 |
| $2,000–4,000 | A complete PC with a used RTX 4090: about $3,200–4,500 | The fastest single-card setup: Gemma 4 26B at around 85 tokens per second, and Qwen 3.8 27B for coding. Upgradable later. Level 3 | 87% | Mac Studio M5 Max with 48 GB ($3,099) or 64 GB ($3,499): silent, 614 GB/s. HP Z2 Mini G1a with 128 GB ($3,300–4,900) if you need 120B-class models | HP's OmniBook RTX Spark laptops, with up to 128 GB, are due this autumn; no price yet |
| $4,000–6,000 | Mac Studio M5 Max with 128 GB: $5,099 | Holds gpt-oss 120B, measured at 76–88 tokens per second on this chip, plus DeepSeek V4 Flash at 2-bit, with Qwen 3.8 27B for coding. Level 3 | 87% | M5 Ultra with 96 GB ($5,499) for the fastest speeds on 27B-class models. NVIDIA DGX Spark ($4,699) if you need NVIDIA's CUDA software. A PC with 2× used RTX 3090 ($4,000–5,500) | — |
| $6,000–10,000 | A PC with 2× RTX 5090: about $9,000–10,500, at the top of the band | The fastest consumer setup: around 80 tokens per second on Qwen 3.6 35B-A3B, with Qwen 3.8 27B at high precision on one card and an assistant on the other. Level 3 | 87% | Two linked DGX Sparks (about $9,400) share 256 GB, which NVIDIA says handles models of about 405B. That is enough for Qwen3.8-Flash-Next (94%), at modest speed. An M5 Ultra with the 36-core chip and 96 GB costs $6,800 | The 256 GB M5 Ultra sits just above this band |
| $10,000–15,000 | Mac Studio M5 Ultra with 256 GB: about $10,800, or $11,299 with 2 TB | Holds Qwen3.8-Flash-Next (about 123 GB), the best local coder on this page, with 1.2 TB/s of bandwidth and room for a second model. Level 3 | 94% | A single RTX PRO 6000 (96 GB) costs $14,000–16,000 for the card alone, so it does not fit this band as a full system | The 512 GB M5 Ultra arrives in late October, unpriced; it adds DeepSeek V4.1 Flash (index 39) |
The curve flattens fast. $1,000–2,000 already reaches 87% of Opus 5 on the Coding Index, and a further $9,000 adds only another seven points. What bigger budgets mainly buy is speed, room for larger general-knowledge models and capacity for a team. No budget here reaches Claude Opus 5.5, which a $20-a-month plan includes.
Local against the cloud: speed
The fastest cloud models write several times faster than any local machine, and they do it while running much larger models. The local figures below are for the best model each machine runs comfortably, so they are not like-for-like: a local 27B model is smaller than anything in the cloud rows.
Tokens per second, single user. Bars are scaled to Gemini 3.8 Flash. Cloud figures are Artificial Analysis output-speed measurements; Claude Opus 5.5 has not yet been measured independently. A local machine serves one person at full speed; a cloud API serves you at this speed however many of your staff use it at once.
Local against the cloud: intelligence and coding
This is where the gap is widest. On the Artificial Analysis Intelligence Index, the best model a machine under $20,000 can hold is DeepSeek V4.1 Flash, on a 512 GB M5 Ultra. It scores 39, level with the budget cloud tier. Only the $85,000-plus GB300 stations can hold the strongest open model, MiMo-V2.6-Pro, and no local machine can run the frontier models at all.
Artificial Analysis Intelligence Index v4.3. Qwen 3.8 27B, the strongest model that fits on a single 24 GB graphics card, scores 34 at its highest effort setting; its predecessor Qwen 3.6 27B scores 21. On coding the picture is more mixed. On the Artificial Analysis Coding Index, Qwen 3.8 27B scores 68.1 against 78.0 for Claude Opus 5, and on the older SWE-bench Verified test Qwen 3.6 27B's 77.2% is within four points of Claude Opus 4.6's 80.8%. On the harder Terminal-Bench 4.0, the best open model, MiMo-V2.6-Pro, reports 34.9% against 59.6% for Claude Opus 5.5.
Level 1: helper
Explains code, writes snippets and small scripts that you copy in and run yourself. Any machine on this page manages it.
Level 2: single-file
Writes or edits one file or a short program reliably, but agent loops that run tests and retry are slow enough to be frustrating.
Level 3: repository agent
Plans and edits across a real codebase and runs the tests, using Qwen 3.6 27B-class models at a usable speed. The ceiling for local hardware today.
Level 4: frontier
Long, unattended work on large codebases, such as migrations and audits that run for hours. Cloud only: Claude Opus 5.5, GPT-6 Astra and Claude Fable 5.1.
Desktop coworkers compared: Claude, ChatGPT and Copilot
All three vendors now sell an assistant that works across your files, apps and desktop rather than just answering questions. Anthropic merged Claude's chat and Cowork modes into one window on 16 September. OpenAI rebuilt its desktop app around ChatGPT Work on 9 July, combining Chat, Work and Codex. Microsoft's Copilot Actions carries out tasks in a separate agent workspace on Windows, but as of September 2026 it has not reached general availability and remains an experimental preview.
1. Claude (Cowork in Claude Desktop)
The most capable and the most careful. Claude Opus 5.5 leads the desktop-task benchmark, the connector directory is deep, and scheduled tasks keep running in the cloud with your PC off. It is slower than ChatGPT and explains more of its working.
2. ChatGPT desktop app (Work)
The fastest and the widest. It has the largest app directory and the only event triggers of the three, so it can react to a new email rather than waiting for a timer. It matched Claude's accuracy in independent tests while finishing sooner.
3. Microsoft Copilot Actions
The most cautious design: each agent gets its own Windows account and workspace. But it is an experimental preview with no scheduler and folder access limited by default. For Microsoft 365 automation today, Copilot Studio is the more mature route.
| Capability | Claude (Cowork / Claude Desktop) | ChatGPT desktop app | Microsoft Copilot Actions |
|---|---|---|---|
| Connector directory (MCP) | 1,253 connectors listed in August 2026, 474 of them partner-verified | 2,289 apps listed in August 2026 | Agent connectors registered in the Windows On-Device Registry (preview). Copilot Studio separately offers 1,400+ connectors plus MCP servers |
| Your own MCP servers | Yes: custom connectors, plus local servers in Claude Desktop | Remote servers only, via Developer mode, set up from chatgpt.com in a browser | Via Copilot Studio (Add a tool, then Model Context Protocol), not in the Copilot app itself |
| Full desktop control without MCP | Yes, in beta for Pro and Max on macOS and Windows. Uses a connector first, then the browser, then the screen. On macOS 15 or later it works in background windows without taking over your pointer | Yes. Computer use in the desktop app on Mac and Windows, demonstrated organising Apple Notes, plus voice control of the computer since 23 July | Yes, but inside a separate agent workspace with its own Windows account. By default it can reach only Documents, Downloads, Desktop, Music, Pictures and Videos |
| Desktop task score (OSWorld 2.0, partial credit) | 81.8% with Opus 5.5 (vendor) | 72.6% with GPT-6 Astra (vendor) | Not published |
| How fast | Slower. In a 22 September head-to-head it took 1:44 against 1:17 on a research task and about 4 minutes against 28 seconds on a Drive-to-spreadsheet job, with identical accuracy. Screen control is its slowest route by design | Fastest on all three tests in the same comparison. GPT-6 Astra completes desktop tasks around 47% faster than GPT-5.6 Sol | No published timings. It runs in a parallel session, so it does not block you while it works |
| Runs without re-prompting? | Yes, on a timer. Scheduled tasks run hourly, daily, weekly or on weekdays, remotely, even with the PC asleep or the app closed. Tasks that need local files or apps run on your machine, which must be on. No event triggers | Yes, on a timer or an event. Scheduled tasks plus triggers from new Gmail, Slack or GitHub activity on Plus plans and above. A task can pause if it needs your action, and sending messages may need approval | No. Each task starts from a prompt and runs while its conversation is open; Windows will not sleep until you close it. No scheduler |
| Paid from your monthly plan or API credits? | Your plan first. Cowork shares one five-hour and weekly allowance with chat and Claude Code on Pro, Max, Team and seat-based Enterprise. Past the limit you can wait, upgrade, or switch on optional usage credits billed at standard API rates. Fable 5.1 runs only on usage credits outside Max. Separate Console API credits are used only if you choose them | Your plan first. On Plus and Pro, Work, Codex and ChatGPT for Excel draw on one shared agentic allowance. Past it you can buy optional flexible credits, with automatic reload if you want it; OpenAI states these are not API credits. Business and Enterprise use a pooled workspace credit balance | No separate charge has been published for the Windows preview. For business agents, Microsoft 365 Copilot is a per-user licence from $30, and internal agents used by licensed staff are largely zero-rated. Copilot Studio agents, and Microsoft 365 Copilot's own Cowork tasks, consume Copilot Credits at $0.01 each pay-as-you-go or $200 per 25,000 a month |
Connector counts come from Node8's August 2026 capture of both public directories and change weekly. OSWorld scores are each vendor's own and use different setups. Timings are from The New Stack's three-task comparison published on 22 September.
Sixty common business apps and where they connect
Most business software now publishes an MCP server, the plug that lets an AI assistant read and act inside it. What differs is whether the app is a one-click listing or something your IT team has to wire in. These are sixty of the most widely used apps with MCP servers, not a ranking by usage.
| Productivity, files and meetings | Claude | ChatGPT | Copilot |
|---|---|---|---|
| Gmail | ✓ | ✓ | ◐ |
| Google Drive | ✓ | ✓ | ◐ |
| Google Calendar | ✓ | ✓ | ◐ |
| Microsoft 365 (Outlook, Teams, SharePoint) | ✓ | ✓ | Native |
| Slack | ✓ | ✓ | ◐ |
| Zoom | ✓ | ✓ | ◐ |
| Notion | ✓ | ✓ | ◐ |
| Atlassian (Jira, Confluence) | ✓ | ✓ | ◐ |
| Asana | ✓ | ✓ | ◐ |
| monday.com | ✓ | ✓ | ◐ |
| ClickUp | ✓ | ✓ | ◐ |
| Trello | ✓ | ✓ | ◐ |
| Wrike | ✓ | ✓ | ◐ |
| Teamwork.com | ✓ | ✓ | ◐ |
| Todoist | ✓ | ✓ | ◐ |
| Linear | ✓ | ✓ | ◐ |
| Smartsheet | ✓ | ✓ | ◐ |
| Airtable | ✓ | ✓ | ◐ |
| Dropbox | ✓ | ✓ | ◐ |
| Box | ✓ | ✓ | ◐ |
| Egnyte | ✓ | ✓ | ◐ |
| Docusign | ✓ | ✓ | ◐ |
| Jotform | ✓ | ✓ | ◐ |
| Calendly | ✓ | ✓ | ◐ |
| Fireflies | ✓ | ✓ | ◐ |
| Otter.ai | ✓ | ✓ | ◐ |
| Fathom | ✓ | ✓ | ◐ |
| Granola | ✓ | ✓ | ◐ |
| Gamma | ✓ | ✓ | ◐ |
| Whimsical | ✓ | ✓ | ◐ |
| Sales, finance, IT and development | Claude | ChatGPT | Copilot |
|---|---|---|---|
| HubSpot | ✓ | ✓ | ◐ |
| Zoho CRM | ✓ | ✓ | ◐ |
| Pipedrive | ◐ | ✓ | ◐ |
| Intercom | ✓ | ✓ | ◐ |
| Mailchimp | ✓ | ✓ | ◐ |
| Klaviyo | ✓ | ✓ | ◐ |
| Semrush | ✓ | ✓ | ◐ |
| Ahrefs | ✓ | ✓ | ◐ |
| Stripe | ✓ | ✓ | ◐ |
| PayPal | ✓ | ✓ | ◐ |
| QuickBooks | ✓ | ✓ | ◐ |
| Xero | ✓ | ◐ | ◐ |
| Shopify | ✓ | ✓ | ◐ |
| Canva | ✓ | ✓ | ◐ |
| Figma | ✓ | ✓ | ◐ |
| Miro | ✓ | ✓ | ◐ |
| WordPress.com | ✓ | ✓ | ◐ |
| Wix | ✓ | ✓ | ◐ |
| Zapier | ✓ | ◐ | ◐ |
| ServiceNow | ✓ | ◐ | ◐ |
| Freshservice | ✓ | ◐ | ◐ |
| PagerDuty | ✓ | ◐ | ◐ |
| Malwarebytes | ✓ | ✓ | ◐ |
| GitHub | ◐ | ✓ | ◐ |
| Postman | ✓ | ◐ | ◐ |
| Sentry | ✓ | ✓ | ◐ |
| Supabase | ✓ | ✓ | ◐ |
| Vercel | ✓ | ✓ | ◐ |
| Cloudflare | ✓ | ✓ | ◐ |
| Datadog | ✓ | ✓ | ◐ |
Claude and ChatGPT status comes from Node8's August 2026 capture of both directories. Gmail, Google Drive and Google Calendar in ChatGPT are OpenAI's own built-in apps; Microsoft 365 appears there as four separate apps (Outlook Email, Outlook Calendar, Teams and SharePoint). The Copilot column refers to Microsoft 365 Copilot and Copilot Studio, where an administrator adds third-party MCP servers as tools. Copilot Actions on Windows cannot use any of these.
What this means for a UK business
Not sure which model fits your business?
We run these models daily across support, development and security work, and we will tell you plainly which one to use and which to avoid.