Self-hosted via OllamaAlibaba Qwen · China

Alibaba Qwen (Qwen3.8): the open family most respected for Chinese, and the two routes open to Hong Kong schools

Qwen3.8-Max reached general availability on 3 August 2026 and the 0902 snapshot shipped on 3 September, a 2.4T-parameter mixture-of-experts model activating about 95B per token with a 1M-token context window at $2 / $6 per 1M tokens. On the open side, Qwen3.8-27B ships under Apache 2.0 and is documented to run in 24GB of VRAM. This page compares the cloud API and on-premises routes and explains availability through local Ollama in Edor.ai.

Edor.ai support status
Self-hosted via OllamaOpen weights run on a school server so data never leaves the campus network

Not a native provider. Qwen is the open family most respected for Chinese-language work, and Edor.ai's Ollama suggestions already include `qwen2.5:14b`. The newer Qwen3.8-27B ships under Apache 2.0 and is documented to run in 24GB of VRAM, making it a practical on-premises choice for Chinese material. The flagship Qwen3.8-Max is a closed cloud API and a cross-border transfer, so student data needs assessment first.

以繁體中文閱讀

Specifications

Vendor
Alibaba Qwen · China
Representative model
qwen3.8-max-0902
Released
2026-09-03
Context window
1,000K tokens
Max output
131K tokens
Modalities
Text / Image / Video
Open weights
Yes (自訂授權 / Apache 2.0(視型號))
API pricing
$2 / $6 — USD per 1M tokens (input / output)
Free tier available
Yes

Short answer: Qwen is the open family most respected for Chinese, and a Hong Kong school has two ways to use it — Qwen3.8-27B on a server inside the school, where no data leaves the campus network, or the Qwen3.8-Max cloud API, which is a cross-border transfer and needs a privacy assessment first.

What this is: one Chinese-language family running both an open and a closed track

Qwen is Alibaba's model family. What most distinguishes it from other Chinese vendors is that each generation ships both a cloud-only flagship and a set of downloadable open weights, with an unusually complete range of sizes — from models that fit on a laptop up to ones that need a data centre.

For a Hong Kong school that structure maps neatly onto two entirely different decisions. Go to the cloud and you get the strongest Chinese performance and the longest context window, at the cost of personal data leaving the territory. Self-host and you buy a graphics card and take on the maintenance, but the whole compliance question disappears. Neither route is better in the abstract; the question is which cost the school prefers to carry.

One caveat worth stating: Qwen's training context leans towards mainland usage. Its Chinese is reliable, but the output may use mainland rather than Hong Kong terminology and may assume mainland grade names and examination systems. Prompts must ask explicitly for Traditional Chinese in Hong Kong conventions, and a teacher must check the result.

The current line-up: one closed flagship, two open sizes, one cheap small model

  • Qwen3.8-Max (generally available 3 August 2026, with the 0902 snapshot on 3 September): a 2.4T-parameter mixture-of-experts model activating about 95B per token, with a 1M-token context window and text, image and video input. Pricing is $2 per 1M input tokens and $6 per 1M output tokens, with cached input at about $0.25. The Max endpoint is closed and its weights are not published.
  • Qwen3.8-2.4T-A95B (12 August 2026): the open-weight build of the same architecture, but under a bespoke licence — the public release omits image input, and providers above US$50M annual revenue need a commercial licence. At 2.4T parameters it needs institutional hardware.
  • Qwen3.8-27B (14 August 2026): a 27.8B dense vision-language model under Apache 2.0, documented to run in 24GB of VRAM. This is the one in the batch best suited to school self-hosting.
  • Qwen3.8-Flash (26 August 2026): a small mixture-of-experts model with 125B total and 6B active parameters at $0.15 / $0.47 per 1M tokens, suited to high-volume batch work.
  • The Qwen2.5 series (19 September 2024): still the most-used Chinese open generation on Ollama, and the generation qwen2.5:14b in our on-premises suggestion list belongs to.

The full comparison of versions, dates and prices is in the automatically generated series table further down this page.

Strengths: the four that matter most to a school

  1. The steadiest Chinese in its open-weight class. Chinese reading comprehension, document summarising, script conversion and understanding of Chinese prompts are all clearly ahead of English-first open models of similar size, which helps most in Chinese, General Studies and Chinese History planning.
  2. There is a size that genuinely runs. 24GB of VRAM is one high-end consumer card, so Qwen3.8-27B turns "self-host a competent Chinese model in school" from a theory into a purchase order.
  3. Apache 2.0 lets a sponsoring body replicate the deployment. The 27B licence has no user cap, so one set of weights can serve several schools with no additional licence fee.
  4. Cloud prices are low, so evaluation is cheap. The flagship at $2 / $6 and Flash at $0.15 / $0.47 mean a teacher can run capability tests, on material containing no personal data, at almost no cost.

Limits: what to say plainly to a principal

  • The Max endpoint is a cross-border data transfer. This is the point that matters most. Sending a student's composition or name to Qwen's cloud API moves personal data out of Hong Kong. The school either completes a privacy impact assessment with informed parental consent, or self-hosts instead.
  • The open flagship is not Apache 2.0. Qwen3.8-2.4T-A95B carries a bespoke licence, the public build omits image input, and providers above US$50M annual revenue need a commercial licence. School use is unlikely to be affected, but offering a service externally requires reading the terms.
  • Terminology needs calibrating. Output often carries mainland vocabulary and assumptions about the mainland education system. That is a context problem rather than a capability one, and the fix is prompt templates plus the school knowledge base.
  • Answers on sensitive and political topics are guarded or skewed. On history, politics and social issues, expect answers that are not neutral. It is usable for gathering material in humanities subjects, but should never be the only source.
  • The hidden costs of self-hosting remain. A graphics card is the start; power, cooling, backups and version upgrades follow. Without IT staffing headroom, the total may not beat a cloud subscription.

Where each tier fits

Teachers: Chinese-language lesson plans, classical and modern Chinese comparisons, variants of Chinese worksheet questions, summarising long Chinese documents. Qwen3.8-27B is enough for all of this; the flagship is unnecessary.

Students: suitable for Chinese reading comprehension hints and vocabulary explanations, in tutor mode only, where the platform holds it to hints rather than answers and keeps conversations a teacher can review. Questions involving value judgements should be led by a teacher rather than handed to the model.

IT coordinators: confirm you have 24GB of VRAM, then choose between qwen2.5:14b and Qwen3.8-27B. Running both in parallel for a term and measuring latency, concurrent users and teacher satisfaction is the cleanest way to settle on one.

Prompts worth testing

Paste these into a Edor.ai teacher tool or any interface connected to local Qwen weights. The terminology constraint is the important part — without it the output mixes in mainland usage.

You are a Hong Kong primary Chinese-language teacher. Design a 30-minute reading and writing activity on the origins and customs of the Mid-Autumn Festival for Primary 5.
Output: one passage under 350 words, a vocabulary table of six terms with example sentences, four comprehension questions, two writing prompts, and three teacher observation points.
Constraints: write everything in Traditional Chinese using Hong Kong conventions, using Hong Kong grade and material names rather than mainland ones; avoid phrasing typical of mainland textbooks; if any historical date or allusion is uncertain, write "teacher to verify" rather than inventing it.
Below is a Secondary 3 student's Chinese narrative composition. Score Content, Expression and Structure out of 5 each, write two specific comments per criterion, then list three things this student should practise next.
Constraints: write in Traditional Chinese using Hong Kong conventions; do not rewrite the composition, give feedback only; quote the student's own sentences unchanged; point to specific paragraphs in your comments; if you are unsure about Hong Kong Chinese-subject marking conventions, mark it "teacher to verify".
[paste composition]

Availability inside Edor.ai

Qwen is not a native provider. Edor.ai natively supports OpenAI, Azure OpenAI, Anthropic, Poe and local Ollama, and Qwen arrives through the last of those. A school runs Qwen3.8-27B or qwen2.5:14b with Ollama on its own server and sets the provider to local Ollama in the admin panel. Our on-premises suggestion list already includes qwen2.5:14b, so an IT coordinator is not starting from nothing.

This is how a school keeps student data inside the campus network, and with Qwen it also settles the cross-border question. Once self-hosted, conversations reach no external cloud, while four-layer gating, teacher monitoring, the ai_audit_logs trail and knowledge-base retrieval all work as normal with identical features to the cloud version. The data still belongs to the school and can be exported in full at any time. Insisting on the Qwen3.8-Max cloud API instead means accepting a cross-border transfer and completing a privacy assessment first.

A school that wants to go further can fine-tune the open weights on its own Chinese lesson plans and model answers, so the output matches school conventions directly. LoRA, QLoRA and INT4 quantisation are what make that possible on a single consumer GPU, and they are the same techniques taught in the LLM Classroom — our on-premises option and the course cover the same concepts.

Alternatives

Different tiers from the same vendor

Most vendors keep flagship, workhorse, lightweight and reasoning lines running at once, and prices can differ tenfold.

ModelTierReleasedContext windowIn / OutNotes
Qwen3.8-Max-0902Flagship2026-09-031,000K$2 / $6The newest post-trained snapshot of the flagship: a 2.4T-parameter MoE activating about 95B per token, with text, image and video input. This snapshot is cloud-API only and its weights are not published.
Qwen3.8-MaxFlagship2026-08-031,000K$2 / $6The flagship reached general availability with per-token pricing; cached input is about $0.25 per 1M tokens.
Qwen3.8-2.4T-A95B(開源權重)Open weights2026-08-121,000KThe open-weight build of the same architecture under a bespoke licence (the public release omits image input, and providers above US$50M annual revenue need a commercial licence). The September 2026 post-training was not open-sourced.
Qwen3.8-27BOpen weights2026-08-141,000KA 27.8B dense vision-language model under Apache 2.0, documented to run in 24GB of VRAM — the most realistic of this batch for school self-hosting.
Qwen3.8-FlashLightweight2026-08-26$0.15 / $0.47A small MoE with 125B total and 6B active parameters, extending context from 256K to 1M at a very low price.
Qwen2.5 系列Open weights2024-09-19Still the most-used Chinese open-model generation on Ollama; the `qwen2.5:14b` in Edor.ai's on-premises suggestions belongs to it.

FAQ

Among open models of comparable size, Qwen is the most widely respected for Chinese, particularly on simplified and traditional conversion, idiom and Chinese document comprehension. Note that its training context leans towards mainland usage, so prompts for Hong Kong lesson plans and circulars must ask explicitly for Traditional Chinese in Hong Kong conventions, with teacher review afterwards.

Qwen3.8-2.4T-A95B does publish weights, but 2.4T parameters need institutional hardware that a school cannot run. The public build also omits image input and uses a bespoke licence rather than Apache 2.0. Qwen3.8-27B is a 27.8B dense vision-language model documented to run in 24GB of VRAM, which is the realistic school choice.

Qwen3.8-2.4T-A95B carries a bespoke licence under which providers above US$50M annual revenue need a separate commercial licence. An ordinary Hong Kong school will not reach that threshold, but a sponsoring body or vendor offering a service externally should read the terms first. Qwen3.8-27B is Apache 2.0 and considerably cleaner.

The Max endpoint is a closed cloud service, so sending a student's composition or name to it transfers personal data out of Hong Kong. Under normal Personal Data (Privacy) Ordinance practice a school should complete a privacy impact assessment and obtain informed parental consent, or simply self-host and avoid the question altogether, which is usually simpler.

There is no hurry. It is small, well documented and still adequate for everyday question-answering and retrieval. Adding Qwen3.8-27B alongside it for a term's comparison is worth it only if the school already has a 24GB card and needs image understanding or a longer context window.

Sources, trust labels and disclaimers

Prices and specifications in this article are current as of 2026-09

  • · All prices, features and specifications follow the official documentation linked above. Vendors may change them at any time — verify before you purchase.
  • · Product names and trademarks mentioned belong to their respective owners. Edor.ai has no partnership, agency or sponsorship relationship with these companies.
  • · This article is an independent review compiled for educational purposes and is not procurement advice or legal advice.
  • · For any use involving student personal data, assess it against your school policy and the Personal Data (Privacy) Ordinance (PDPO) before rollout.

Related reading

  • China's LLMs Compared: DeepSeek, Qwen, Kimi, GLM and Five More — What Can a Hong Kong School Actually Use?

    DeepSeek costs one sixty-seventh of GPT-6 Astra per input token and is no weaker in Chinese. So why does a school platform not simply plug into it? This roundup compares DeepSeek, Qwen, Kimi, GLM, MiniMax, Doubao, Hunyuan and ERNIE, explains the crucial difference between open weights and a cloud API, and sets out the three routes that are genuinely workable under the Personal Data (Privacy) Ordinance.

  • Open Weights or Closed API? The Real Cost Maths of Self-Hosting, and Where Privacy Flips the Answer

    Many schools assume that buying a machine to run an open model must be cheaper. Put real usage through September 2026 published pricing and the answer usually reverses — the annual cloud API bill is too small for hardware to beat. This article works through both sides honestly, lists the open models that are realistic on school hardware, and identifies the point at which data policy tips the balance towards self-hosting.

  • The 2026 All-Model Roundup: 15 Providers and One Selection Framework for Schools

    From OpenAI, Anthropic, Google, xAI, Meta and Mistral to DeepSeek, Qwen, Kimi, GLM, MiniMax, Doubao, Hunyuan and ERNIE, plus the coding-only Cursor Composer. One school usage model, applied to every vendor's published pricing, puts all 15 on a single table — followed by a four-question framework for choosing between them.

Subscribe to the AI in Education newsletter

One email a month: practical AI teaching articles for Hong Kong schools, platform updates and grant news. Unsubscribe any time.

We only use this address for the newsletter and never share it.