TL;DR: GLM is the model family from Z.ai, the Beijing lab formerly known as Zhipu AI. Most of its releases ship with open weights under the permissive MIT license, and the family has built its reputation on agentic coding - including a subscription plan that plugs its models into Claude Code and the other popular coding agents. You can download the weights, call Z.ai's API, or use third-party hosts.
What it is and how you use it
Z.ai began in 2019 as Zhipu AI, a spin-out of Tsinghua University's knowledge-engineering group. It took the Z.ai name for its international business in 2025 and listed on the Hong Kong stock exchange in January 2026. One governance fact belongs in any procurement conversation: the company has been on the US Commerce Department's Entity List since January 2025, which restricts what US firms may export to it but does not by itself restrict running its open weights. The GLM name predates the chatbot era - it stands for General Language Model, the architecture the Tsinghua group published in 2022 - and the family has kept it through every generation since.
The modern line began with GLM-4.5 in mid-2025 and has moved quickly through the GLM-5 series, each generation a large mixture-of-experts model with a smaller Flash sibling, plus a vision branch and smaller models for OCR, speech, and image generation. The licensing story is unusually clean for an open family: every release through the GLM-5.2 generation, and the current GLM-5.3-Flash, is plain MIT. The one exception is the current flagship, GLM-5.3, which carries its own license - the same broad grant with a single carve-out requiring very large model-as-a-service operators to clear deployment with Z.ai first. Check which license the checkpoint you deploy carries; the answer is MIT more often than not.
What the family is known for is agentic coding. Z.ai positions each release as an engine for long-horizon coding agents, exposes a reasoning-effort setting on the current generation, and leads its launches with the agentic benchmark suites - Terminal-Bench, SWE-bench Pro, and tool-using variants of Humanity's Last Exam - as vendor claims. The distinctive move is distribution: the GLM Coding Plan is a flat-rate subscription whose Anthropic-compatible and OpenAI-compatible endpoints are documented as drop-in backends for Claude Code, Cline, Kilo Code, OpenCode, and a long list of other agents, with a one-line environment change. Z.ai also ships its own Apache-licensed coding harness, ZCode. None of this is endorsed by the makers of those agents; it works because the request shapes are compatible.
Access follows the open-family pattern. Z.ai's international API is OpenAI-style and sits alongside a separate mainland platform; the Coding Plan adds the agent-facing endpoints above. Western inference providers such as Together and Fireworks host the open weights, OpenRouter lists them, Microsoft Foundry carries the flagship through Fireworks, and Amazon Bedrock has picked up earlier generations. Self-hosting means vLLM or SGLang on a multi-GPU node - the full-size models are far too large for a workstation, and Ollama's GLM entry routes to a hosted endpoint rather than local weights. Where it fits: coding-agent backends where a flat subscription beats per-token billing, and self-hosted deployments that want a permissive license at frontier scale. The counterweights are the ones every China-based first-party API carries - jurisdiction and data governance - plus, for US buyers, the Entity List question above.
Where it sits in the AI stack
GLM sits at the model layer as an open family, with one extra door: a flat-rate plan whose endpoints slot directly behind the coding agent you already use:
Key tools and implementations
-
Open-weight downloads
The GLM-4.5 through GLM-5 series on Hugging Face - MIT for nearly every checkpoint, with the current flagship under Z.ai's own license.
-
Z.ai API and GLM Coding Plan
The OpenAI-style first-party API, plus a flat-rate plan with Anthropic-compatible and OpenAI-compatible endpoints built to sit behind coding agents.
-
Coding-agent integrations and ZCode
Documented setup for Claude Code, Cline, Kilo Code, OpenCode and others via a base-URL change, and Z.ai's own Apache-licensed terminal harness.
-
Providers and serving runtimes
Together, Fireworks, OpenRouter and Microsoft Foundry serve the weights on Western infrastructure; vLLM and SGLang run them on self-managed GPUs.
Related entries
- Qwen Alibaba's family of open-weight models spanning many sizes, strong in multilingual and coding tasks.
- Kimi Moonshot AI's open-weight model family, known for very long context windows and agentic coding, served by its own API and third-party hosts.
- DeepSeek A Chinese AI lab known for open-weight models with strong reasoning and unusually low training costs.
- AI coding agent Software that plans and executes multi-step coding tasks - reading files, editing code, and running tests with minimal supervision.
- Open-weights models Models whose trained weights are published for anyone to download, run locally, and fine-tune.