LLM & Model Insights Published 11 September 2026 8 min read

OpenAI vs Anthropic vs xAI vs Google: Which US Provider Actually Suits a School?

In September 2026 the four leading US flagships range from $0.75 to $10 per 1M input tokens — a 13x spread. This head-to-head compares them on the six axes a school actually cares about (price, context, Chinese-language handling, content safety, platform availability and cost cliffs) and explains why the strongest model is usually not the one a school should be running.

Edor.ai Education Team
Former teachers, edtech consultants and AI engineers
以繁體中文閱讀

An IT coordinator at a secondary school fed the same 22-page school assessment policy to all four US flagships and asked each for a summary. The most expensive run cost 13 times the cheapest — and when the four summaries were put in front of the panel heads, nobody could say which one was worth the difference. This is not an anecdote: on September 2026 published pricing, OpenAI's GPT-6 Astra costs $10 per 1M input tokens and Google's Gemini 3.8 Flash costs $0.75, a gap of exactly 13.3x.

The question a school has to answer is never "which is strongest". It is "which is good enough on the tasks we actually run, at a cost we can control". Those two questions rarely have the same answer.

Three terms that get conflated

Before any comparison, three words in plain language — because almost every bad selection decision starts with these three being treated as one.

Context window is the model's working memory. It sets how much material the model can hold in mind at once. A million tokens is roughly a full term of teaching materials plus the school policy set. But a big memory is not the same as remembering well: tip the whole box in and the model is more likely to latch onto the wrong detail, and, as below, long inputs often trigger doubled billing.

Parameter count is the number of brain cells. More brain cells means more latent capability, but a Primary 5 Chinese worksheet does not need a doctoral brain. Using a flagship to generate fill-in-the-blanks is hiring a barrister to do your photocopying.

Reasoning effort is the dial for how long the model thinks before answering. All four converged on this design in 2026: OpenAI's reasoning.effort, Anthropic's effort, Google's thinking_level, xAI's reasoning_effort. More thinking gives better answers, but the thinking itself is billed at output rates. There is more on this in the reasoning revolution.

Head-to-head on six axes that matter to a school

All figures below are the specifications and pricing published by each vendor in September 2026, in USD per 1M tokens.

AxisOpenAI GPT-6 AstraAnthropic Claude Fable 5.1Google Gemini 3.8 FlashxAI Grok 4.6
Release date2026-09-032026-09-012026-09-022026-08-12
Input / output price10 / 5010 / 500.75 / 3.75 (introductory)2 / 6
Cache read price1.000.250.075 (introductory)0.50
Context window1,050,0001,000,0001,048,576500,000
Max output128,000128,00065,536No stated limit
Knowledge cutoff30 April 2026June 2026Not stated on pricing page1 February 2026
Long-input cost cliffAbove 272K tokens the whole request bills at 2x input and 1.5x outputNo equivalent clauseNo equivalent clauseAt 200K tokens the whole request bills at 4 / 12
Open weightsNoNoNo (the Gemma family is separate)No
Reachable in Edor.aiNative (default)NativeVia Poe onlyVia Poe only

Four things in that table deserve a second read:

  • Gemini's $0.75 / $3.75 has an expiry date. Google's pricing page states this rate holds through 31 December 2026 and becomes $1.50 / $7.50 on 1 January 2027. A school building a three-year budget should use the later figure.
  • The cache gap is wider than the headline gap. A cache read on Claude Fable 5.1 costs $0.25, a quarter of GPT-6 Astra's. For any workflow that reloads the same school curriculum outline on every call, this column drives the bill more than the headline rate does.
  • Two vendors have a long-input cliff. OpenAI and xAI use the same mechanic: once a request crosses the threshold, it is not only the excess that is repriced — the entire request bills at the higher rate. Pasting a whole term of materials in one go can silently double a bill.

On speed, none of the four publishes tokens-per-second figures that can be compared across vendors. What Anthropic does publish on its model pages is a relative latency band (Fable 5.1 marked Slower, Opus 5 Moderate, Sonnet 5 Fast, Haiku 4.5 Fastest) — a vendor-reported relative description, not a measured benchmark. Any marketing material claiming "A is X times faster than B" should be met with a request for the test method.

Provider by provider: what it means in a school

OpenAI: the only one that supplies all three model types

What makes OpenAI hard to replace is not that GPT is strongest, but that one vendor supplies chat models, embedding models and speech models. A school knowledge base needs embeddings; oral practice needs speech recognition and synthesis. Lose either and you no longer have a complete platform. That is why OpenAI is the Edor.ai default. Sponsoring bodies already on Azure can route the same models through Azure OpenAI and manage data location under their own subscription. See the OpenAI model page.

Anthropic: the steady choice for Chinese writing feedback and long documents

Adaptive thinking on Claude Fable 5.1 is always on, with effort defaulting to high — it thinks before it answers by default. That helps on tasks requiring judgement, such as composition feedback or checking work against a marking rubric. The cost is higher latency and an output price of $50, level with GPT-6 Astra. Note that Anthropic supplies no embedding or speech models, so a school that drops OpenAI entirely must arrange those separately. Where budget matters, Claude Sonnet 5 at $2 / $10 — confirmed as permanent pricing on 10 August 2026 — is the practical tier for schools. See the Anthropic Claude model page.

Google: cheapest for bulk documents, but access is the obstacle

Gemini 3.8 Flash is the only one of the four with native text, image, audio and video input, at an introductory price far below the rest. For grunt work such as turning two years of scanned notices into structured data, the value is obvious. The obstacle is access: Edor.ai has no native Gemini provider, so calls go through Poe, and Poe offers no embeddings or speech. Schools already on Google Workspace for Education can run both side by side. See the Google Gemini model page.

xAI: mid-priced, but content policy needs your own assessment

Grok 4.6 at $2 / $6 sits between Gemini and the other two, its 500K context is more than enough for any school document, and the xhigh reasoning level is in place. What a school must judge for itself is content posture: Grok is tightly integrated with X platform content, which in a student-facing setting is something to assess first and then write into the school AI policy. It is also Poe-only. See the xAI Grok model page.

Limits and blind spots the quotation will not mention

All four hallucinate, and they are worst on local detail. Hong Kong curriculum codes, subject reference numbers, EDB circular numbers, a school's full Chinese and English name — these are the thinnest parts of any training set and, inconveniently, exactly what a teacher is most likely to paste straight into a notice. We have watched a model invent a curriculum code with perfect formatting that does not exist. Anything going to parents or students needs teacher review.

Knowledge cutoffs differ, and they are earlier than people assume. GPT-6 Astra is 30 April 2026, Grok 4.6 is 1 February 2026, Claude Sonnet 5 is January 2026. Ask about "this school year" and the model may well answer about last year, in exactly the same confident tone. Anywhere current facts matter, the answer must come from the school's own knowledge base rather than model memory.

Prices move, and in both directions. OpenAI cut GPT-5.6 Luna by 80% on 30 July 2026 and Sol by over 20% on 21 August 2026; Gemini 3.8 Flash has already announced a doubling for January 2027. A school wiring up the API itself cannot lock a budget — which is precisely why we absorb model price movement inside a fixed annual fee.

A long context does not mean good comprehension. The million-token headline does not mean you will get a good answer by pushing a million tokens in. The right approach is RAG, retrieving only the relevant passages, which is both more accurate and avoids the long-input surcharge. See the school knowledge base.

"Native provider" is a capability distinction, not marketing. Reaching Gemini or Grok through Poe gives you chat, not embeddings and speech. That distinction is invisible in a demo and surfaces in week three of deployment.

Actionable guidance: four things to do this week

One, separate the tasks that need a flagship. In our experience roughly eight in ten requests — question generation, summarising, reformatting, first-pass marking, notice drafts — are well served by a mid tier. Only cross-document reasoning, multi-step maths and science explanation, and tasks cross-referencing several policies justify escalation. Getting the routing right saves more than getting the vendor right.

Two, run a blind test on your own materials rather than reading benchmarks. The method is simple: take three real artefacts (a Chinese composition, an extract of the school curriculum, a draft parent notice), run one prompt across all four, mask the vendor names and have three panel heads score them. That record also makes a useful annex for AI funding applications and reporting.

Three, match yourself to one of these cases:

  • You want one vendor covering chat, knowledge base and oral practice → OpenAI (or Azure OpenAI).
  • You prioritise steady Chinese writing feedback and long-document judgement → Anthropic Claude, on Sonnet 5 daily with escalation to Fable 5.1.
  • You need the lowest cost on bulk documents or media and already use Google Workspace → Gemini, budgeted at 2027 prices.
  • You want a middle point on cost and capability and have completed a content-safety assessment → Grok.
  • Data cannot leave the campus network → none of the four applies; read open weights versus closed API instead.

Four, write switchability into the procurement terms. In the first nine months of 2026 all four vendors replaced a flagship at least once. Any contract that locks a school to one provider becomes a liability inside two years. The questions to ask are: when we switch, does the teacher interface change, must the knowledge base be rebuilt, do historical records survive? All three answers should be no.

Takeaway

The capability gap between the four US providers in September 2026 is much narrower than the price gap between them. The real selection skill for a school is not picking the strongest vendor but building a routing pattern — cheap by default, expensive for hard problems, local for sensitive data — that survives the next generation of models untouched.

For the full picture across all 15 providers, including Chinese and European vendors, see the 2026 all-model roundup. To get the vocabulary straight first, start at lesson one of the LLM classroom.

FAQ

Usually not. For lesson planning, worksheet generation and first-pass marking, teachers rarely see a quality difference between a mid tier and a flagship, yet the price gap runs from 5x to 13x. The sensible pattern is mid tier by default and escalation only for hard tasks, routed automatically by the platform.

Two caveats. First, the $0.75 / $3.75 rate for Gemini 3.8 Flash is introductory; Google's own pricing page states it rises to $1.50 / $7.50 on 1 January 2027. Second, Edor.ai has no native Gemini provider, so it is reached through Poe, and Poe supplies no embedding or speech models — the knowledge base and oral practice still need an OpenAI or Azure key.

Have the subject panel head and IT coordinator assess it first. Grok is tightly integrated with X platform content and takes a different line on content safety from the other three, so a school should test it and record that assessment before any student-facing use.

None leads across every Chinese task. In our school deployments Claude is the steadiest for Chinese writing feedback and long-document work, OpenAI has the more mature Cantonese speech recognition, and Gemini is the most economical for bulk document processing. Rather than betting on one vendor, choose a platform that lets you switch.

For the school it should be one setting in the admin panel. An administrator switches provider or assigns a model per module, and the teacher and student interfaces, the school knowledge base and all historical records are unaffected.

Sources, trust labels and disclaimers

Prices and specifications in this article are current as of 2026-09

  • · All prices, features and specifications follow the official documentation linked above. Vendors may change them at any time — verify before you purchase.
  • · Product names and trademarks mentioned belong to their respective owners. Edor.ai has no partnership, agency or sponsorship relationship with these companies.
  • · This article is an independent review compiled for educational purposes and is not procurement advice or legal advice.
  • · For any use involving student personal data, assess it against your school policy and the Personal Data (Privacy) Ordinance (PDPO) before rollout.
model comparisonOpenAIAnthropicGoogle GeminixAI Grok
Subscribe to the AI in Education newsletter

One email a month: practical AI teaching articles for Hong Kong schools, platform updates and grant news. Unsubscribe any time.

We only use this address for the newsletter and never share it.

Related articles

LLM & Model Insights 11 September 2026 8 min read

The 2026 All-Model Roundup: 15 Providers and One Selection Framework for Schools

From OpenAI, Anthropic, Google, xAI, Meta and Mistral to DeepSeek, Qwen, Kimi, GLM, MiniMax, Doubao, Hunyuan and ERNIE, plus the coding-only Cursor Composer. One school usage model, applied to every vendor's published pricing, puts all 15 on a single table — followed by a four-question framework for choosing between them.

Read more
LLM & Model Insights 11 September 2026 8 min read

China's LLMs Compared: DeepSeek, Qwen, Kimi, GLM and Five More — What Can a Hong Kong School Actually Use?

DeepSeek costs one sixty-seventh of GPT-6 Astra per input token and is no weaker in Chinese. So why does a school platform not simply plug into it? This roundup compares DeepSeek, Qwen, Kimi, GLM, MiniMax, Doubao, Hunyuan and ERNIE, explains the crucial difference between open weights and a cloud API, and sets out the three routes that are genuinely workable under the Personal Data (Privacy) Ordinance.

Read more
LLM & Model Insights 11 September 2026 7 min read

From Chatbot to AI Agent: Tool Calling, Multi-Step Automation, MCP, and What a School Should Not Automate

A chatbot only talks. An agent acts — it looks things up, reads documents, calls systems and runs several steps in sequence. This article explains tool calling and the MCP standard in terms a teacher can use, compares published tool-call charges, and offers a should-automate and should-not-automate list. The question is not what AI can do, but which step a person must confirm.

Read more