Moonshot Kimi (K3 / K2.6): the largest open-weight model there is, and why a school still cannot use it
Kimi K3 launched on 16 July 2026 and its 2.8T-parameter weights were published on 27 July 2026, with native vision, a 1M-token context window and API pricing of $3 / $15 per 1M tokens. This page sets out where K3 and K2.6 sit, what the long-context strength is actually worth to a teacher, the bespoke licence and data-residency limits, and why Edor.ai has not wired Kimi up.
Not wired up today. Kimi K3 publishes weights, but at 2.8T parameters (about 104B active) it is far beyond what a school server can host, so self-hosting is unrealistic. Its official API sits in mainland China and is a cross-border transfer. Schools wanting to try its long-document strength should have teachers use a personal account on de-identified material only.
Specifications
- Vendor
- Moonshot AI (Kimi) · China
- Representative model
kimi-k3- Released
- 2026-07-16
- Context window
- 1,000K tokens
- Max output
- Not published
- Modalities
- Text / Image
- Open weights
- Yes (Kimi K3 License)
- API pricing
- $3 / $15 — USD per 1M tokens (input / output)
- Free tier available
- Yes
Short answer: Kimi K3 is the largest open-weight model in existence and the engineering deserves real respect, but for a Hong Kong school the self-hosting bar at trillion-parameter scale and the cross-border transfer to a mainland API keep it in the "worth knowing about" column rather than the "can deploy" one.
What this is
Moonshot AI is a Beijing model team with two product lines: the consumer-facing Kimi assistant and the developer API at platform.kimi.ai. Its best-known label is long context — since 2023 the pitch has been "read the whole book in one go", and that positioning runs straight through to today's 1M tokens.
Moonshot's other distinguishing habit is publishing weights. From K2 onwards every flagship generation has gone up on Hugging Face in full, and K3 is the first open-weight model in the 3-trillion-parameter class. For research institutions and companies able to host it, that is a rare resource. Open, however, is not the same as runnable in a school, and the gap between the two is the subject of the rest of this page.
The current line-up
- Kimi K3 (released 16 July 2026, weights published 27 July 2026): a 2.8T-parameter MoE activating about 104B parameters per token, with native image input and a 1M-token context window. The API model id is
kimi-k3at $3 per 1M input tokens and $15 per 1M output tokens, with cache hits at $0.30. The weights carry Moonshot's own "Kimi K3 License". - Kimi K2.6 (April 2026): the previous flagship under a modified MIT licence at $0.95 / $4.00 per 1M tokens, still a common choice for cost-sensitive batch work.
- Kimi K2 (July 2025): the release that brought Kimi to international attention and made long context a signature selling point for Chinese models.
The full version and price comparison is in the automatically generated series table further down this page.
Strengths
- Genuinely frontier-class open weights. Publishing the strongest generation in full was still unusual in 2026, and it is a real contribution to research, to local fine-tuning, and to any organisation that does not want to be locked to one vendor. K3 is the largest of this cohort.
- Steady long-document and long-conversation behaviour. A million tokens is not just a number: a full curriculum guide, a term of coursework records or a long research report can stay in one conversation without repeated pasting. Teachers notice it most on cross-unit comparisons and long-form summarising.
- Flagship capability at mid-range pricing. At $3 / $15 per 1M tokens it sits in the lower half of the flagship field, and cache hits at $0.30 make repeated reads of the same document set unusually cheap.
- Natural Chinese. This is the common strength of Chinese model teams. Chinese writing, rewriting and polishing read more naturally than in several Western models; Traditional-character output and Hong Kong usage still need checking, but the base is sound.
Limits
- The self-hosting bar is out of reach. 2.8T parameters require multiple nodes of high-end GPUs. University labs have to plan for it carefully; a school server room is simply not in the conversation. "It is open, so we can run it ourselves" does not hold at this scale.
- The licence is not a standard open licence. The Kimi K3 License is bespoke and differs from MIT or Apache 2.0 on commercial scope, territory and attribution. K2.6's modified MIT is more permissive but is still not plain MIT.
- The API is in mainland China. For a Hong Kong school, calling it directly is a cross-border data transfer, and anything involving personal data needs an assessment first. That is not a technical problem but a compliance responsibility the school carries itself.
- Content policy is for the school to judge. The model's safety boundaries are designed within the mainland framework and will not always match how the Hong Kong curriculum treats history, current affairs or civic education. Any school considering student-facing use should test this itself and decide.
- There is almost no education ecosystem around it. No Hong Kong curriculum prompt templates, teacher community or classroom case studies, so an IT coordinator starts from zero and training costs run well above the mainstream providers.
How a Hong Kong school should view it
Start by separating two questions: is this model good, and may our school use it. K3 answers the first well; schools have to answer the second.
Three practical steps. First, treat Kimi as reference reading: it lets an IT coordinator and panel chairs see how far open weights have come, which is useful context when judging other options. Second, if an individual teacher wants to try it, treat it as a public website — generic planning material only, no pupil names, student numbers, marks or un-redacted school documents. Third, if a sponsoring body seriously considers wiring any Chinese model into a school system, do two pieces of work first: a cross-border transfer assessment under the Personal Data (Privacy) Ordinance setting out where data goes, how long it is kept and who can read it; and a legal opinion on the licence terms covering commercial use and redistribution.
As for the very common wish — strong Chinese without data leaving the campus — the realistic answer is not Kimi but a sensibly sized open model hosted on site. See the alternatives below.
Availability inside Edor.ai
Not wired up today. Edor.ai natively supports five providers — OpenAI, Azure OpenAI, Anthropic, Poe and local Ollama — and Kimi is not among them, so administrators will not see the option. The reasons are the two above: 2.8T parameters cannot be hosted by a school, and the official API sits in mainland China as a cross-border transfer.
We also have no partnership of any kind with Moonshot AI. This page exists to set out verifiable facts so a school can judge for itself, not to sell anything.
Alternatives
- For strong Chinese with data kept entirely on campus, look at the 27B-class open weights covered under Alibaba Qwen, served through local Ollama — the most realistic route today.
- To compare the licences, prices and availability of the Chinese models side by side, read the DeepSeek, Kimi, Qwen and GLM comparison.
- To understand the real cost and privacy trade-off between self-hosting and cloud, read the open versus closed cost and privacy trade-off.
- If you simply need a provider that works in the school platform today, start with OpenAI or Anthropic Claude, or go back to the model overview.
Different tiers from the same vendor
Most vendors keep flagship, workhorse, lightweight and reasoning lines running at once, and prices can differ tenfold.
| Model | Tier | Released | Context window | In / Out | Notes |
|---|---|---|---|---|---|
| Kimi K3 | Flagship | 2026-07-16 | 1,000K | $3 / $15 | A 2.8T-parameter MoE (about 104B active per token) with native vision and a 1M-token context window — currently the largest open-weight model; weights were published on 27 July 2026. Cache hits cost $0.30 per 1M tokens. |
| Kimi K2.6 | Flagship | 2026-04 | — | $0.95 / $4 | The previous flagship under a modified MIT licence, far cheaper than K3 and still a common choice for cost-sensitive work. |
| Kimi K2 | Open weights | 2025-07-11 | — | — | The release that brought Kimi to international attention and made long context a signature selling point for Chinese models. |
FAQ
Open weights and runnable are two different things. K3 has 2.8T parameters with about 104B active per token; even at low precision, inference needs multiple nodes of high-end GPUs plus the power and cooling to match. A school considering self-hosting should start at the 27B class, not at a trillion-parameter flagship.
Technically yes, but treat it as a public website. Do not enter pupil names, student numbers, marks, family circumstances or anything else identifying, and do not upload school documents that have not been de-identified. Generic planning material carries low risk; anything involving personal data belongs on a gated, audited school platform.
MIT is a very short standard licence with few conditions. K3 uses Moonshot's own terms, which have to be read clause by clause — commercial scope, any scale or territorial conditions, and attribution requirements. If a sponsoring body plans to embed it in a school system, its legal adviser should read it first rather than assuming that open means unrestricted.
Most directly, reading a whole textbook unit, a year's worth of coursework records or a long report in one pass for a combined analysis. Two caveats: long input is billed, and models still miss details deep inside very long contexts. In practice, retrieving the relevant sections and then asking is usually more accurate than pasting everything.
We only add a provider when three things hold at once — the data-processing location and terms can be explained to a school, the licence is clear, and the model performs reliably in education settings. Kimi still has open questions on the first two for Hong Kong schools, so the answer today is no. If that changes, the date at the top of this page will change with it.
Prices and specifications in this article are current as of 2026-09
- Kimi Open Platform — model inference pricing
- Hugging Face — moonshotai/Kimi-K3 model card (architecture, licence, deployment)
- Hugging Face — Moonshot AI organisation page
- Kimi official site
- · All prices, features and specifications follow the official documentation linked above. Vendors may change them at any time — verify before you purchase.
- · Product names and trademarks mentioned belong to their respective owners. Edor.ai has no partnership, agency or sponsorship relationship with these companies.
- · This article is an independent review compiled for educational purposes and is not procurement advice or legal advice.
- · For any use involving student personal data, assess it against your school policy and the Personal Data (Privacy) Ordinance (PDPO) before rollout.
Related reading
- China's LLMs Compared: DeepSeek, Qwen, Kimi, GLM and Five More — What Can a Hong Kong School Actually Use?
DeepSeek costs one sixty-seventh of GPT-6 Astra per input token and is no weaker in Chinese. So why does a school platform not simply plug into it? This roundup compares DeepSeek, Qwen, Kimi, GLM, MiniMax, Doubao, Hunyuan and ERNIE, explains the crucial difference between open weights and a cloud API, and sets out the three routes that are genuinely workable under the Personal Data (Privacy) Ordinance.
- The 2026 All-Model Roundup: 15 Providers and One Selection Framework for Schools
From OpenAI, Anthropic, Google, xAI, Meta and Mistral to DeepSeek, Qwen, Kimi, GLM, MiniMax, Doubao, Hunyuan and ERNIE, plus the coding-only Cursor Composer. One school usage model, applied to every vendor's published pricing, puts all 15 on a single table — followed by a four-question framework for choosing between them.