DeepSeek (V4.1-Flash): the cheapest open-weight flagship, and the two realistic routes for a Hong Kong school
DeepSeek's current model V4.1-Flash launched on 10 September 2026 with a 1M-token context window, 384K output and off-peak pricing of $0.15 / $0.60 per 1M tokens, with weights published under MIT. This page explains why its MoE architecture is so cheap, what self-hosting really demands in hardware, and why calling the official API is a cross-border transfer that needs a privacy assessment first.
Not a native provider. The practical route is to run a distilled or quantised build on a school server through Ollama so no data leaves the campus network. Calling DeepSeek's own API is a cross-border transfer and needs a privacy assessment first.
Specifications
- Vendor
- DeepSeek · China
- Representative model
deepseek-flash- Released
- 2026-09-10
- Context window
- 1,000K tokens
- Max output
- 384K tokens
- Modalities
- Text / Image
- Open weights
- Yes (MIT)
- API pricing
- $0.15 / $0.6 — USD per 1M tokens (input / output)
- Free tier available
- No
Short answer: DeepSeek publishes its flagship weights under MIT at the lowest unit price of any major provider, but it has no native integration in Edor.ai, so a school has two realistic routes — self-host a distilled build on Ollama, or complete a privacy assessment before calling the official API, which is a cross-border transfer.
What this is: a vendor that competes on engineering efficiency
DeepSeek was founded in China in 2023. What brought it international attention was DeepSeek-R1 in January 2025: an open-weight reasoning model showing that near-frontier logic did not require the training budgets people had assumed. That release changed a lot of minds about whether open models could only do second-tier work.
DeepSeek's technical direction has stayed on one theme, which is buying down inference cost through architecture. The current V4.1-Flash is a 552B-parameter mixture-of-experts model that activates only 8B parameters during prefill and 16B during decode, with native image input. Its model card puts the emphasis on KV cache compression, which is precisely where long-context cost comes from. The result is a model with a 1M-token context window at a fraction of the mainstream flagship rates.
For a school, the licence is the other headline: the weights are published under MIT. That is the most permissive of the common open licences and it matters to a sponsoring body that wants full autonomy over its deployment. But being able to download a model and being able to run it are two different things, as set out below.
The current line-up: one Flash trunk and two milestones
DeepSeek-V4.1-Flash is the current model, released on 10 September 2026 with a 1M-token context window, 384K output and native image handling. Pricing is $0.15 per 1M input tokens and $0.60 per 1M output tokens — but that is the off-peak rate, and peak hours double it, and DeepSeek has no free tier.
Behind it sit DeepSeek-V4-Pro and V4-Flash (16 August 2026), the generation that introduced peak and off-peak pricing. Two earlier releases explain the whole trajectory: DeepSeek-R1 (20 January 2025) is the open-weight reasoning model that reset industry expectations and is still the one most often named in school self-hosting discussions, and DeepSeek-V3 (26 December 2024) brought MoE architectures into the mainstream conversation.
The full list of versions, dates and prices is in the automatically generated series table further down this page.
Strengths: the four that matter most to a school
- The licence is clear. MIT is the easiest permissive licence to explain to an incorporated management committee: commercial use, modification and redistribution are all allowed. Compared with the bespoke licences some competitors use, with revenue thresholds or usage restrictions attached, the review effort is far lower.
- The lowest unit price. At an off-peak $0.15 / $0.60 it sits at the bottom of the range among major providers, which is a real advantage on high-volume work of moderate difficulty.
- Long context is cheap. A 1M-token window combined with the compressed cache design makes reading a whole school document in one pass affordable in a way it is not with other providers.
- A mature distillation ecosystem. The R1 distillations come in sizes from 1.5B to 70B on Ollama, so a school can pick one that fits the hardware it already owns and run a proof of concept before buying anything.
Limits: what to say before anyone signs anything
- The official API is a cross-border transfer. DeepSeek's servers are in mainland China, so any request containing a student's name, class or written work leaves Hong Kong. A school should complete and file an assessment against the Data Protection Principles before deciding to enable it. This is the most important point on this page.
- Self-hosting the flagship is not realistic. The full 552B weights need a multi-GPU server environment beyond what a typical school server room can carry. What is feasible is a distilled or quantised build, and the capability gap to the flagship is substantial — flagship benchmark figures should never be used to sell an on-premises deployment.
- Distilled builds need licence checks one by one. The R1 weights are MIT, but some distillations derive from Qwen 2.5 under Apache 2.0 and others from Llama 3.1 and 3.3 under their own licences. Confirm which build you actually downloaded before writing "MIT" into procurement papers.
- No embedding or speech models. DeepSeek makes language models only. The vector index behind a school knowledge base and all oral practice still come from OpenAI or Azure OpenAI.
- Peak pricing doubles and there is no free tier. The $0.15 / $0.60 rate applies off-peak only, and a Hong Kong school's peak usage, during the school day, will not necessarily fall in the off-peak window. Do not build a cost estimate by multiplying usage by the lowest rate.
- Localisation and hallucination. Like every model, DeepSeek gets Hong Kong curriculum specifics, Traditional Chinese usage conventions and local assessment terminology wrong, and some phrasing leans towards mainland usage. All output needs teacher review.
Where each tier fits
Teachers: once a school has completed its assessment and enabled a route, this suits high-volume work that involves no student personal data — rewriting English material at different difficulty levels, compiling background on a topic from public sources, drafting practice questions. Marking that involves student work and names is better left on an assessed native provider.
Students: only consider it on the self-hosted route. With a distilled build on Ollama, nothing leaves the campus network, and combined with the platform's four-layer gating and tutor mode this is the most controllable option. Have teachers test the distilled build's Chinese feedback quality first, then decide which subjects it serves.
IT coordinators: two things come first. Write a privacy assessment setting out the data categories, the destination, the level of protection and the alternatives, and have management sign it. Then run a mid-sized distillation from Ollama on an existing workstation as a proof of concept, measuring Traditional Chinese output quality and response speed, before deciding whether the hardware purchase is justified.
Prompts worth testing
Copy these into a Edor.ai teacher tool or any DeepSeek interface. The constraint lines are what make the output usable:
You are a Hong Kong primary General Studies teacher. Design a 25-minute classroom activity on "weather and climate" for Primary 4, using only simple equipment a school would already have.
Output: three activity objectives, a procedure written in short sentences pupils can follow, an observation record table, and two extension questions.
Constraints: write in Hong Kong English conventions and use the vocabulary Hong Kong schools use, for example "worksheet"; do not give away answers; for any specific Hong Kong temperature or rainfall figure, write "teacher to verify" and say which Hong Kong Observatory resource should be consulted.
Below is a Secondary 2 Chinese-language reading passage and five draft comprehension questions. Please review the quality of the questions.
Output: for each question, say which skill it tests (retrieval, inference or evaluation), identify any question that can be answered without reading the passage, and propose one rewritten version of each question that needs improvement.
Constraints: write your commentary in Hong Kong Chinese conventions; do not rewrite the passage; do not introduce material beyond the passage; if you are unsure whether a word suits the Hong Kong primary or junior secondary level, write "teacher to verify".
[paste passage and draft questions]
Availability inside Edor.ai
DeepSeek is not a native provider in Edor.ai. The five native routes are OpenAI, Azure OpenAI, Anthropic, Poe and local Ollama, and none of them is a dedicated DeepSeek connector. In practice a school has two options.
The first is self-hosting through local Ollama, which is the route we recommend. A distilled or quantised build is downloaded onto a school server, inference happens inside the campus network, no student data goes out, and the platform's four-layer gating, teacher monitoring and ai_audit_logs trail all operate normally. The costs are the hardware and the capability gap between a distillation and the flagship.
The second is calling DeepSeek's own API, which is a cross-border data transfer. The school must first complete a privacy assessment stating what data would be sent, where it goes, what protections apply and whether a lower-risk alternative exists, signed off and filed by management. We would not recommend enabling any student-facing use before that step is finished.
On either route, DeepSeek supplies no embedding or speech models, so the school knowledge base and oral practice still need an OpenAI or Azure OpenAI key. If the school's real goal is keeping data inside the campus network rather than using DeepSeek specifically, it is worth comparing the other open-weight options alongside it and reading how the platform handles data and security.
Alternatives
- For the same self-hosting route with more established Chinese-language strength, see Alibaba Qwen.
- For the same route at a lower hardware threshold, see Meta open weights or Mistral.
- For a natively integrated provider that works out of the box, see OpenAI or Anthropic Claude.
- To compare the mainland Chinese providers and their risks in one place, read DeepSeek, Kimi, Qwen and GLM compared and the open versus closed cost and privacy trade-off.
Different tiers from the same vendor
Most vendors keep flagship, workhorse, lightweight and reasoning lines running at once, and prices can differ tenfold.
| Model | Tier | Released | Context window | In / Out | Notes |
|---|---|---|---|---|---|
| DeepSeek-V4.1-Flash | Flagship | 2026-09-10 | 1,000K | $0.15 / $0.6 | The current model: a 552B-parameter MoE that activates only 8B (prefill) to 16B (decode) at inference, with native image support. The $0.15 / $0.60 rate is off-peak; peak hours double it. |
| DeepSeek-V4-Pro / V4-Flash | Flagship | 2026-08-16 | — | — | The V4 generation that introduced peak / off-peak pricing. From 14 September 2026 V4-Pro requests route to V4.1-Flash. |
| DeepSeek-R1 | Reasoning | 2025-01-20 | — | — | An open-weight reasoning model that showed near-frontier logic could be trained at modest cost — one of the most consequential releases of 2025. |
| DeepSeek-V3 | Open weights | 2024-12-26 | — | — | The version that brought MoE architectures into the mainstream conversation and still the starting point for many self-hosting discussions. |
FAQ
Realistically, no. V4.1-Flash is a 552B-parameter MoE, and although only 8B to 16B parameters are active at inference, the full weights still need a multi-GPU server environment. The practical choice for a school is a distilled or quantised build, such as the DeepSeek-R1 distilled family on Ollama at sizes from 1.5B to 70B, which is noticeably weaker than the flagship.
It means personal data leaving Hong Kong to be processed on servers elsewhere. DeepSeek's official API is in mainland China, so any request containing a student's name, class or written work falls into that category. What a school has to assess is the categories of data, the purpose, the level of protection at the destination and whether a lower-risk alternative exists — and then record the reasoning.
Because an MoE activates only a small fraction of its parameters at inference, and DeepSeek has engineered heavily around KV cache compression, so the actual computation is far below a dense model of comparable size. Note, though, that this is the off-peak rate, peak hours double it, and DeepSeek has no free tier.
The V4.1-Flash weights are indeed published under MIT, which is genuinely permissive. Be careful with distilled builds, though: some of the R1 distillations on Ollama derive from Qwen 2.5 under Apache 2.0 and others from Llama 3.1 and 3.3 under their own Llama licences. Check the actual licence of each model before writing "MIT-licensed model" into a procurement document.
The risk is much lower, but two things still apply. Teaching materials and examination items may carry the school's own intellectual property, and once a habit forms it is easy for a teacher to paste in a student's work without thinking. It is safer for the school to provide one assessed route centrally than to leave teachers using personal accounts.
Prices and specifications in this article are current as of 2026-09
- DeepSeek — official API documentation (model names and usage)
- Hugging Face — DeepSeek-V4.1-Flash model card (architecture, parameters, MIT licence)
- Hugging Face — DeepSeek organisation page (weights by generation)
- Ollama — DeepSeek-R1 library (self-hosting sizes and licence notes)
- Hong Kong Office of the Privacy Commissioner — Six Data Protection Principles
- · All prices, features and specifications follow the official documentation linked above. Vendors may change them at any time — verify before you purchase.
- · Product names and trademarks mentioned belong to their respective owners. Edor.ai has no partnership, agency or sponsorship relationship with these companies.
- · This article is an independent review compiled for educational purposes and is not procurement advice or legal advice.
- · For any use involving student personal data, assess it against your school policy and the Personal Data (Privacy) Ordinance (PDPO) before rollout.
Related reading
- China's LLMs Compared: DeepSeek, Qwen, Kimi, GLM and Five More — What Can a Hong Kong School Actually Use?
DeepSeek costs one sixty-seventh of GPT-6 Astra per input token and is no weaker in Chinese. So why does a school platform not simply plug into it? This roundup compares DeepSeek, Qwen, Kimi, GLM, MiniMax, Doubao, Hunyuan and ERNIE, explains the crucial difference between open weights and a cloud API, and sets out the three routes that are genuinely workable under the Personal Data (Privacy) Ordinance.
- Open Weights or Closed API? The Real Cost Maths of Self-Hosting, and Where Privacy Flips the Answer
Many schools assume that buying a machine to run an open model must be cheaper. Put real usage through September 2026 published pricing and the answer usually reverses — the annual cloud API bill is too small for hardware to beat. This article works through both sides honestly, lists the open models that are realistic on school hardware, and identifies the point at which data policy tips the balance towards self-hosting.
- The 2026 All-Model Roundup: 15 Providers and One Selection Framework for Schools
From OpenAI, Anthropic, Google, xAI, Meta and Mistral to DeepSeek, Qwen, Kimi, GLM, MiniMax, Doubao, Hunyuan and ERNIE, plus the coding-only Cursor Composer. One school usage model, applied to every vendor's published pricing, puts all 15 on a single table — followed by a four-question framework for choosing between them.