model
Zhipu GLM
Direct answerZhipu currently recommends GLM-5.2 for long-running text and coding, GLM-5V-Turbo for visual coding, and GLM-Image for image generation, alongside specialized models.
Updated · Reviewed
Current official portfolio
Zhipu's overview currently recommends GLM-5.2, GLM-5V-Turbo, and GLM-Image and separately lists text, vision, image, video, audiovisual, embedding, and character models. These categories are not switches on one chat model.
Capability split
GLM-5.2 targets long-running tasks and software delivery; vision models handle image or video understanding and multimodal tools; GLM-Image, CogVideoX, TTS, and ASR use dedicated generation or processing APIs. Official tables publish context and output per model.
Choosing within the family
Evaluate the flagship text model for complex coding and agents, vision for graphical interfaces or visual code, image models for Chinese typography, and specialized resources for audio/video. Free or Flash variants still require evaluation for high-value work.
API use
Text chat uses /v1/chat/completions and a Bearer token, while media or asynchronous jobs use task-specific routes. Thinking, tools, structured output, and context caching are enabled per model.
Lifecycle
The official overview marks models approaching retirement. Verify exact version, context and output, thinking controls, media limits, concurrency, billing, regulatory filing, and data terms.
Use cases
- Agents and coding
- Long context
- Vision generation
API protocols
/v1/chat/completions
FAQ
Is the model name permanent?
No. Official aliases can move and enabled API IDs can change; pin a tested ID in production and verify availability separately in the model marketplace.
Can GLM text, vision, image, and audio-video models replace one another?
No. Text and agents, visual understanding, image generation, video, and speech are separate model categories with different modalities, context, and output protocols. Select a dedicated model and validate quality, latency, and limits on real samples.
Can GLM's OpenAI-compatible chat call every media task?
No. Text conversation can use Chat Completions, while image, video, TTS, ASR, and asynchronous work use their own interfaces. Thinking, tools, structured output, and caching must also be enabled only for the exact model.
Related guides
Official sources
- GLM Overview Official
兔子API