Anthropic Claude (Fable 5.1 / Opus 5 / Sonnet 5): the steadiest writing feedback, and its two gaps
Anthropic's top model Claude Fable 5.1 launched on 1 September 2026 with a 1M-token context window, 128K output and pricing of $10 / $50 per 1M tokens. This page sets out how Fable 5.1, Opus 5, Sonnet 5 and Haiku 4.5 divide the work, what teachers can genuinely use them for, and what it means for a school platform that Claude has no embedding or speech models.
Native provider. Claude is reliable for Chinese writing feedback and long-document work, but it has no embedding or speech models, so the knowledge base and oral practice still need an OpenAI or Azure key.
Specifications
- Vendor
- Anthropic · United States
- Representative model
claude-fable-5-1- Released
- 2026-09-01
- Context window
- 1,000K tokens
- Max output
- 128K tokens
- Modalities
- Text / Image
- Open weights
- No
- API pricing
- $10 / $50 — USD per 1M tokens (input / output)
- Free tier available
- Yes
Short answer: of the four large American providers, Claude handles long documents and writing feedback most reliably, and it is a native provider you can switch to inside Edor.ai. But it has no embedding and no speech models, so it can be a school's main chat model without being able to run a school platform on its own.
What this is: a vendor that sells safety as the product
Anthropic was founded in 2021 by researchers who had left OpenAI, and it has one product line: Claude. What distinguishes it from other vendors is that model behaviour and safety are positioned as the product rather than as a clause in the terms. It publishes system cards and safety evaluations, publishes retirement commitments for every model, and even publishes the system prompt used on claude.ai. For a school that has to answer to an incorporated management committee, that verifiability has value in itself.
In the classroom the recurring teacher comment is that Claude is willing to write at length and willing to give reasons. Marking a Secondary 3 English composition, it tends to explain each deduction; summarising a forty-page school assessment policy, it is less prone to blending requirements from different chapters. That steadiness comes from how it handles long context rather than from sheer scale.
One thing should be said plainly, though. Anthropic makes language models only. There is no embedding model and no speech recognition or synthesis. A complete school platform needs all three kinds, so inside Edor.ai, Claude is a chat model you can run as your main choice, not a provider that can stand alone.
The current line-up: three flagship tiers plus a small one
The September 2026 line-up is easiest to read from the hardest work downwards.
Claude Fable 5.1 is the current top tier, released on 1 September 2026 with a 1M-token context window, 128K output, and pricing of $10 per 1M input tokens and $50 per 1M output tokens. Its adaptive thinking is always on, meaning the model decides how long to think and there is no manual switch. Worth noting is that cache reads fall to $0.25 per 1M tokens, the lowest in the family, which matters when the same school document is queried repeatedly.
Claude Opus 5 (24 July 2026) is the tier Anthropic itself recommends as a default, at $5 / $25 — half the price of Fable 5.1 for close to the same intelligence. Claude Sonnet 5 (30 June 2026) offers the same 1M-token context at $2 / $10 and is the best value tier as well as the default on the Free and Pro plans; on 10 August 2026 Anthropic confirmed that rate as permanent rather than introductory. At the bottom, Claude Haiku 4.5 is the fast $1 / $5 tier with a 200K context window, suited to classification, tagging and other high-frequency small tasks.
One older model deserves a mention: Claude 3.7 Sonnet was the first mainstream model to expose extended thinking as a switch, and every vendor's reasoning mode today descends from it. The full list of versions, dates and prices is in the automatically generated series table further down this page.
Strengths: the four that matter most to a school
- Long documents hold up. The 1M-token context window is the same across all three flagship tiers, so length alone never forces an upgrade to the most expensive option. Consolidating a year's curriculum outline or comparing old and new versions of a school policy is well within Sonnet 5.
- Writing feedback comes with evidence. When marking, Claude tends to quote the pupil's own sentence before explaining the problem rather than issuing a bare verdict. That saves time both for teacher review and for the student's understanding.
- Cache reads are cheap. When the same document is queried repeatedly across a lesson the difference is visible, and Fable 5.1's $0.25 rate is the lowest in the family.
- Documentation and lifecycle are transparent. Retirement dates and knowledge cutoffs are published for every model, which gives an IT coordinator something official to cite in a risk assessment.
Limits: what to say before anyone signs anything
- There is no embedding model. This is the most concrete gap. The vector index behind a school knowledge base cannot be built with Claude, so any school that wants retrieval over its own documents must keep an OpenAI or Azure OpenAI key.
- There are no speech models. Cantonese reading assessment, English oral practice and model pronunciation audio are all outside what Claude can do. For primary Chinese and English oral work that is a hard constraint.
- The flagship costs the same as OpenAI's. Fable 5.1's $10 / $50 matches GPT-6 Astra, so a school wiring up the API itself faces the same budgeting difficulty. That is why we charge a fixed annual fee and absorb the price movement.
- It is occasionally over-cautious. The conservative safety training means legitimate teaching material on war history, violence in literature or adolescent health education can be declined or heavily softened. Stating the purpose and year group in the prompt usually fixes it, but teachers should expect it.
- Hallucination has not gone away. Like every language model, Claude states local details confidently and sometimes wrongly — clause numbers in Hong Kong curriculum documents, school names, assessment specifics. Outward-facing documents and marking results need teacher review.
Where each tier fits
Teachers: marking long compositions, consolidating policy across documents, drafting letters and circulars, summarising meeting notes, and rewriting worksheets for differentiated groups. A Sonnet 5 class model is enough for all of these; the most expensive tier is unnecessary.
Students: only inside a gated, monitored environment. Tutor mode restricts Claude to hints rather than answers, with the hint ceiling set by the teacher. Claude's willingness to explain is an advantage here, but it can also say too much at once, so teachers should put a length limit in the prompt as well.
IT coordinators: before switching to Claude, confirm that an OpenAI or Azure key is still in place behind the knowledge base and oral practice. If the school intends to disable OpenAI entirely, those two modules stop working and the replacement arrangement needs to be written into the project plan first.
Prompts worth testing
Copy these into a Edor.ai teacher tool or any Claude interface. The constraint lines are what make the output usable:
You are a Hong Kong secondary Chinese-language teacher. Below is a Secondary 3 student's narrative composition. Score Content and Ideas, Structure and Organisation, and Language and Rhetoric out of 6 each.
Output: for each criterion quote one or two sentences from the student as evidence, then write two specific comments; finish with three practice priorities, each paired with one action the student can take immediately.
Constraints: write in Hong Kong Chinese conventions; do not rewrite any paragraph on the student's behalf; keep the total comment length under 400 characters; if a cultural allusion or quotation source is uncertain, mark it "teacher to verify" rather than inferring it.
[paste composition]
You are a Hong Kong primary General Studies panel head. Design three 35-minute lessons on "public transport in Hong Kong" for Primary 5.
Output: the learning focus of each lesson, a minute-by-minute lesson flow, one photocopiable worksheet, and a differentiation plan stating what less confident and more confident pupils each do.
Constraints: write in Hong Kong English conventions; present the worksheet as a table; do not include an answer key; for any fare, frequency or route figure, write "teacher to verify" instead of filling in a number from memory.
Availability inside Edor.ai
Anthropic is a native provider. Once an administrator adds one school Anthropic API key, teacher tools, the AI tutor and student conversations can all run on Claude, and it can be set per module — for example Chinese and English on Claude with everything else on OpenAI.
There is one prerequisite to settle first: Claude has no embedding or speech models. The vector index behind the school knowledge base and both Cantonese and English oral practice need an OpenAI or Azure OpenAI key behind them. A school can make Claude its main chat model, but it cannot add an Anthropic key alone and switch the other providers off. We suggest stating this plainly in procurement papers and in the technical annex of an EDB funding application, so that a missing module is not discovered at acceptance testing.
Everything else matches the other providers. The key belongs to the school and can be rotated or revoked, every call passes through four-layer gating and ai_audit_logs, and the data belongs to the school with a free one-click full export at any time.
Alternatives
- For one vendor covering chat, embedding and speech together, see OpenAI.
- For the lowest cost on high document volumes, especially on Google Workspace, see Google Gemini — but note it is reachable here only through Poe.
- To keep data entirely inside the campus network, see Meta open weights and read the open versus closed cost and privacy trade-off.
- To understand adaptive thinking, caching and hallucination first, start with the LLM Classroom or read the four-provider comparison.
Different tiers from the same vendor
Most vendors keep flagship, workhorse, lightweight and reasoning lines running at once, and prices can differ tenfold.
| Model | Tier | Released | Context window | In / Out | Notes |
|---|---|---|---|---|---|
| Claude Fable 5.1 | Flagship | 2026-09-01 | 1,000K | $10 / $50 | The current top tier with adaptive thinking always on. Cache reads dropped to $0.25 per 1M tokens, the lowest in the family. |
| Claude Opus 5 | Flagship | 2026-07-24 | 1,000K | $5 / $25 | Anthropic's recommended default: close to Fable 5's frontier intelligence at half the price, with a Fast mode available. |
| Claude Sonnet 5 | Workhorse | 2026-06-30 | 1,000K | $2 / $10 | The best value tier and the default on Free and Pro plans. On 10 August 2026 the $2 / $10 rate was confirmed as permanent. |
| Claude Fable 5 / Mythos 5 | Flagship | 2026-06-09 | — | — | The start of the fifth generation and the first Claude able to sustain multi-day, asynchronous tasks. |
| Claude Haiku 4.5 | Lightweight | 2025-10-15 | 200K | $1 / $5 | The fastest and cheapest tier with a 200K context window, suited to classification, tagging and high-frequency small tasks. |
| Claude Opus 4.5 / 4.6 / 4.7 / 4.8 | Flagship | 2025-11-24 | — | — | Four flagship iterations between November 2025 and May 2026, the direct predecessors of Opus 5. |
| Claude 3.7 Sonnet | Reasoning | 2025-02-24 | — | — | The first mainstream model to expose extended thinking as a switch, and the milestone that popularised chain-of-thought reasoning. |
| Claude 3.5 Sonnet | Workhorse | 2024-06-20 | — | — | The most-cited grading baseline of 2024 and an early workhorse for LLM-as-a-Judge evaluation. |
FAQ
Both can be switched in the admin panel, so this is not an either-or decision. In practice the split is that Claude tends to be steadier on long compositions, cross-document school policy work and letters to parents where tone matters, while knowledge-base retrieval and oral practice need embedding and speech models, which only OpenAI or Azure OpenAI can supply.
A school knowledge base works by chunking your documents, turning each chunk into a vector, then retrieving the relevant passages to feed the model. The embedding model does that conversion. If a school adds only an Anthropic key and no OpenAI or Azure key, the knowledge base and oral practice modules cannot run, though the other teacher tools are unaffected.
No. That is Anthropic's developer API list price, not a school's spend. Edor.ai is a per-school annual subscription with AI usage included, no token charges and no per-seat pricing.
Anthropic's safety training is comparatively cautious, so material on war history, violence in literature or health education can occasionally be declined or heavily softened. Stating the teaching purpose and the year group in the prompt usually resolves it, and a school can also switch an individual module to another model.
We call the API under no-training terms, which is different from a teacher using the consumer claude.ai product. The data belongs to the school and can be exported in full at any time as JSON, CSV or the original files.
Prices and specifications in this article are current as of 2026-09
- Anthropic — Claude models overview (specifications and model IDs)
- Anthropic — Claude Fable 5.1 model page
- Anthropic — Pricing, including caching and batch discounts
- Anthropic — Claude product page
- · All prices, features and specifications follow the official documentation linked above. Vendors may change them at any time — verify before you purchase.
- · Product names and trademarks mentioned belong to their respective owners. Edor.ai has no partnership, agency or sponsorship relationship with these companies.
- · This article is an independent review compiled for educational purposes and is not procurement advice or legal advice.
- · For any use involving student personal data, assess it against your school policy and the Personal Data (Privacy) Ordinance (PDPO) before rollout.
Related reading
- OpenAI vs Anthropic vs xAI vs Google: Which US Provider Actually Suits a School?
In September 2026 the four leading US flagships range from $0.75 to $10 per 1M input tokens — a 13x spread. This head-to-head compares them on the six axes a school actually cares about (price, context, Chinese-language handling, content safety, platform availability and cost cliffs) and explains why the strongest model is usually not the one a school should be running.
- The 2026 All-Model Roundup: 15 Providers and One Selection Framework for Schools
From OpenAI, Anthropic, Google, xAI, Meta and Mistral to DeepSeek, Qwen, Kimi, GLM, MiniMax, Doubao, Hunyuan and ERNIE, plus the coding-only Cursor Composer. One school usage model, applied to every vendor's published pricing, puts all 15 on a single table — followed by a four-question framework for choosing between them.
- The Reasoning Revolution: From Chain-of-Thought to Adaptive Thinking, and What It Changes for Marking, Maths and Science
Models in 2024 blurted out an answer. Models in 2026 think first, and decide for themselves how long to think. This article explains chain-of-thought, reasoning effort and adaptive thinking in terms a teacher can use, compares the four vendors' effort dials, and sets out what has genuinely changed for marking, maths and science teaching — and what has not.