Self-hosted via OllamaZhipu AI / Z.ai (GLM) · China

Zhipu AI / Z.ai (GLM-5.3, 5.2, 5.3-Flash): the team that puts large models under MIT

GLM-5.3 arrived on 14 August 2026 at $1.40 / $4.40 per 1M tokens with a 1M-token context window. GLM-5.2 is a 744B MoE with MIT-licensed open weights, and GLM-5.3-Flash is a natively multimodal MoE, also under MIT. This page covers the licensing advantage, the Hong Kong listing, and the hardware and compliance questions a school has to answer before self-hosting.

Edor.ai support status
Self-hosted via OllamaOpen weights run on a school server so data never leaves the campus network

Not a native provider. GLM is one of the few teams publishing large models under MIT (GLM-5.2 and GLM-5.3-Flash), which appeals to sponsoring bodies wanting full autonomy, though the parameter counts still need server-class hardware and weight releases do not land at the same time for every version — verify against the official Hugging Face page before committing.

以繁體中文閱讀

Specifications

Vendor
Zhipu AI / Z.ai (GLM) · China
Representative model
GLM-5.3
Released
2026-08-14
Context window
1,000K tokens
Max output
Not published
Modalities
Text / Image
Open weights
Yes (MIT(GLM-5.2 與 5.3-Flash))
API pricing
$1.4 / $4.4 — USD per 1M tokens (input / output)
Free tier available
Yes

Short answer: GLM is one of the few teams releasing large models under MIT, which makes it the most interesting option for a sponsoring body that wants genuine self-deployment — but the appeal stops at the licence, because the real barrier is still hardware, and the GLM-5.3 flagship weights remain unpublished.

What this is

Zhipu AI, which trades internationally as Z.ai, grew out of a Tsinghua University research group, and GLM is the model family it has developed continuously since 2021. Compared with other Chinese model teams, its most consistent trait is publishing weights under a standard licence: GLM-4.5 and GLM-4.6 built its international profile, and GLM-5.2 went straight to MIT.

One piece of background comes up often: Zhipu listed in Hong Kong in January 2026, the first foundation-model company to IPO there. That matters to local investment and industry discussion, but schools should note one thing — where a company is listed says nothing about where it processes data.

The current line-up

  • GLM-5.3 (14 August 2026): the current flagship, aimed at coding and agentic work, with a 1M-token context window at $1.40 per 1M input tokens and $4.40 per 1M output tokens. Same price as GLM-5.2, but the weights were not public when this page was updated.
  • GLM-5.3-Flash (26 August 2026): a natively multimodal MoE with 320B total and 18B active parameters, released under MIT at a promotional $0.075 / $0.25 per 1M tokens against a list price of $0.15 / $0.50.
  • GLM-5.2 (16 June 2026): a 744B-parameter MoE with MIT-licensed open weights, the most-cited fully self-deployable large model of 2026.
  • GLM-4.5 / 4.6 (from July 2025): the generation that built GLM's international profile, still widely deployed.

The full version and price comparison is in the automatically generated series table further down this page.

Strengths

  1. The cleanest licence terms in the field. Against a landscape of bespoke vendor licences, MIT costs almost nothing to review. For a school that has to explain legal risk to its management committee, that is more useful than a benchmark table.
  2. Clearly separated price tiers. The flagship at $1.40 / $4.40 already sits low among its peers, and Flash goes down to a promotional $0.075 / $0.25. For high-volume routine document work, the gap is measured in multiples.
  3. Native multimodality even in the small tier. GLM-5.3-Flash activates only 18B parameters yet handles text and images natively, which is an unusual combination for schools wanting to process scanned worksheets and text on limited hardware.
  4. Solid Chinese foundations. Chinese comprehension, summarising and rewriting are dependable. Traditional-character output still needs checking for Taiwanese or mainland usage, but the overall quality is already usable for Chinese and humanities lesson preparation.

Limits

  • The flagship weights are not published. The MIT benefit applies to 5.2 and 5.3-Flash. Using the strongest 5.3 tier means calling the cloud API, so the self-deployment argument does not hold at the top of the range.
  • Self-hosting still needs server-class hardware. 744B parameters are beyond a campus workstation, and even the 320B Flash needs serious specification. Licence freedom and deployment feasibility are separate questions.
  • The cloud API is a cross-border transfer. Calling Z.ai directly takes data out of Hong Kong, so the school must complete its own privacy assessment and parent notification arrangements.
  • No embedding or speech companion for education. Even after self-hosting a chat model, the vector index behind a school knowledge base and Cantonese oral practice need separate provision. One model does not cover everything.
  • Content policy needs school-based testing. Safety boundaries follow the mainland framework, so a school should sample-test history, current affairs and civic education topics before opening it to pupils.

How a Hong Kong school should view it

Of the Chinese models in this set, GLM is the one most worth a sponsoring body's serious evaluation, not because of benchmarks but because MIT licensing makes full self-deployment legally coherent for the first time.

Turning that into a workable school plan runs in a particular order. Start with hardware. An inference server able to run a 320B MoE is a combination of capital expenditure, server-room space, power and maintenance staffing, which usually only balances when shared across a sponsoring body. A single school without that scale will almost always end up choosing a smaller model. Next comes compliance. The whole point of self-hosting is that data never leaves the campus network, which removes the cross-border question; if the decision instead lands on the Z.ai cloud API, do the Personal Data (Privacy) Ordinance assessment properly and add extra gating for student-facing use. Third is division of labour: even with GLM handling conversation, embedding and speech still need a second arrangement, so do not leave them out of the plan.

It is worth adding that a self-hosting proposal can be presented to a management committee as capital funding now against subscription savings later — but the maintenance cost has to be stated honestly, or the server becomes unmaintained within a year.

Availability inside Edor.ai

Not a native provider, but reachable through local Ollama if you self-host. Edor.ai natively supports OpenAI, Azure OpenAI, Anthropic, Poe and local Ollama. We do not build in a Z.ai cloud connector and we have no partnership with Zhipu. However, if a school or sponsoring body stands up its own inference server and exposes an OpenAI-compatible endpoint through Ollama, the platform can point at it and teacher tools and the AI tutor carry on working.

One caveat: on that route the hardware, updates and fault handling belong to the school, while we handle platform-side integration and safety gating. The knowledge base still needs an embedding model and oral practice still needs speech models, so an OpenAI or Azure OpenAI key or a separate local deployment remains necessary. Start by assessing feasibility with us through contact us.

Alternatives

Different tiers from the same vendor

Most vendors keep flagship, workhorse, lightweight and reasoning lines running at once, and prices can differ tenfold.

ModelTierReleasedContext windowIn / OutNotes
GLM-5.3Flagship2026-08-141,000K$1.4 / $4.4The current flagship aimed at coding and agentic work, priced the same as GLM-5.2. Check the official Hugging Face page for the current status of the flagship weights.
GLM-5.3-FlashLightweight2026-08-261,000K$0.075 / $0.25A natively multimodal MoE with 320B total and 18B active parameters under MIT, at a promotional $0.075 / $0.25 per 1M tokens (list $0.15 / $0.50).
GLM-5.2Flagship2026-06-16A 744B-parameter MoE with MIT-licensed open weights — the most-cited fully self-deployable large model of 2026.
GLM-4.5 / 4.6Open weights2025-07-28The generation that built GLM's international profile and the mainstay before Zhipu's Hong Kong listing in January 2026.

FAQ

MIT is a one-paragraph standard open licence permitting commercial use, modification and redistribution, essentially requiring only that the copyright notice is kept. The difference is review cost. A bespoke licence has to be read clause by clause for territorial, scale or use restrictions; MIT is clear on one reading. For schools and sponsoring bodies with limited resources, that gap is very real.

Technically yes, if the hardware is there. GLM-5.2 has 744B parameters and needs a server-class GPU cluster with matching facilities even after quantisation, which usually only makes sense at sponsoring-body level. GLM-5.3-Flash, at 320B total and 18B active, lowers the bar but still exceeds a single consumer graphics card. Assess the hardware first, then pick the model.

No. Where a company is listed and where it processes data are different things. Zhipu listed in Hong Kong in January 2026, the first foundation-model company to do so, but that does not change where its cloud API runs. What matters is the actual service terms and server location.

A school can, but only after an assessment. At $1.40 / $4.40 the price is competitive among flagships, yet calling the cloud API means data crosses the border, and because the GLM-5.3 weights are not published, self-hosting is not an option at that tier. Schools that value self-deployment should look at the MIT-licensed GLM-5.2 or GLM-5.3-Flash instead.

Three steps: the school or sponsoring body provides a working inference server, loads the weights there with Ollama or an equivalent, and points the platform's Ollama endpoint at it in the admin panel. Features carry on as normal, but embedding and speech still need separate arrangements. Talk to us first so we can assess feasibility together.

Sources, trust labels and disclaimers

Prices and specifications in this article are current as of 2026-09

  • · All prices, features and specifications follow the official documentation linked above. Vendors may change them at any time — verify before you purchase.
  • · Product names and trademarks mentioned belong to their respective owners. Edor.ai has no partnership, agency or sponsorship relationship with these companies.
  • · This article is an independent review compiled for educational purposes and is not procurement advice or legal advice.
  • · For any use involving student personal data, assess it against your school policy and the Personal Data (Privacy) Ordinance (PDPO) before rollout.

Related reading

  • China's LLMs Compared: DeepSeek, Qwen, Kimi, GLM and Five More — What Can a Hong Kong School Actually Use?

    DeepSeek costs one sixty-seventh of GPT-6 Astra per input token and is no weaker in Chinese. So why does a school platform not simply plug into it? This roundup compares DeepSeek, Qwen, Kimi, GLM, MiniMax, Doubao, Hunyuan and ERNIE, explains the crucial difference between open weights and a cloud API, and sets out the three routes that are genuinely workable under the Personal Data (Privacy) Ordinance.

  • Open Weights or Closed API? The Real Cost Maths of Self-Hosting, and Where Privacy Flips the Answer

    Many schools assume that buying a machine to run an open model must be cheaper. Put real usage through September 2026 published pricing and the answer usually reverses — the annual cloud API bill is too small for hardware to beat. This article works through both sides honestly, lists the open models that are realistic on school hardware, and identifies the point at which data policy tips the balance towards self-hosting.

  • The 2026 All-Model Roundup: 15 Providers and One Selection Framework for Schools

    From OpenAI, Anthropic, Google, xAI, Meta and Mistral to DeepSeek, Qwen, Kimi, GLM, MiniMax, Doubao, Hunyuan and ERNIE, plus the coding-only Cursor Composer. One school usage model, applied to every vendor's published pricing, puts all 15 on a single table — followed by a four-question framework for choosing between them.

Subscribe to the AI in Education newsletter

One email a month: practical AI teaching articles for Hong Kong schools, platform updates and grant news. Unsubscribe any time.

We only use this address for the newsletter and never share it.