Skip to content
Dashboard

GLM 4.5 Air

GLM 4.5 Air is Z.AI's efficiency-focused model released July 28, 2025. It delivers fast inference for high-volume workloads while keeping reasoning and coding capability at reduced cost compared to the full GLM-4.5.

Input and output price
Input $0.20, Output $1.10, Per 1M tokens
24h uptime
Loading AI Gateway uptime
import { streamText } from 'ai'
const result = streamText({
model: 'zai/glm-4.5-air',
prompt: 'Why is the sky blue?'
})
Read docs

Copy link to headingProviders

Route requests across multiple providers. Copy a provider slug to set your preference. Visit the docs for more info. Using a provider means you agree to their terms, listed under Legal.

Provider
Context
Max Output
Latency
Throughput
Input
Output
Cache
Web Search
Capabilities
ZDR
No Training
Free Tier
Release Date
128K96K0.5 s74 tps
$0.20/M
$1.10/M
Read$0.03/M
07/28/2025

Copy link to headingPlayground

Try out GLM 4.5 Air by Z.AI. Usage is billed to your team at API rates. Free users (those who haven't made a payment) get $5 of credits every 30 days.

zai logo
zai logo

GLM 4.5 Air

Copy link to headingUptime

Direct request success rate on AI Gateway and per-provider. Visit the docs for more info.

Copy link to headingThroughput

P50 throughput on live AI Gateway traffic, in tokens per second (TPS). Visit the docs for more info.

Copy link to headingLatency

P50 time to first token (TTFT) on live AI Gateway traffic, in milliseconds. View the docs for more info.

Copy link to headingMore models by Z.AI

Model
Context
Latency
Throughput
Input
Output
Cache
Web Search
Capabilities
Providers
ZDR
No Training
Free Tier
Release Date
1M2.3 s179 tps
$0.70/M
$2.20/M
Read$0.13/M
digitalocean logo
09/02/2026
1M0.3 s455 tps
$0.07/M
$0.24/M
Read$0.01/M
+1
baseten logo
deepinfra logo
digitalocean logo
+14
08/26/2026
1M0.2 s387 tps
$0.70/M+1 more
$2.20/M+1 more
Read$0.12/M
baseten logo
deepinfra logo
digitalocean logo
+11
08/18/2026
1M0.5 s208 tps
$2.10/M
$6.60/M
Read$0.21/M
alibaba logo
baseten logo
fireworks logo
06/23/2026
1M0.2 s503 tps
$0.70/M+1 more
$2.20/M+1 more
Read$0.11/M
alibaba logo
baseten logo
crusoe logo
+15
06/16/2026
203K0.5 s125 tps
$1/M
$3.20/M
Read$0.20/M
bedrock logo
novita logo
zai logo
02/12/2026

Copy link to headingAbout GLM 4.5 Air

GLM 4.5 Air was released July 28, 2025 as the efficiency-optimized variant in Z.AI's GLM-4.5 generation. Where GLM-4.5 targets maximum capability, GLM 4.5 Air trades a degree of depth for faster inference and lower per-token cost, making it practical for high-throughput production pipelines.

The model retains the core reasoning, coding, and agentic capabilities of the GLM-4.5 family while operating at reduced computational overhead. This positions it for workloads where response latency and cost per request are primary constraints: classification, extraction, summarization, and conversational applications that process high volumes of requests.

GLM 4.5 Air supports the same context window of 128K tokens as the full GLM-4.5 model. Through AI Gateway, it benefits from unified API access, built-in observability, and intelligent provider routing with automatic retries.

Copy link to headingWhat To Consider When Choosing a Provider

  • Configuration: GLM 4.5 Air is optimized for speed. For tasks requiring deep multi-step reasoning, the full GLM-4.5 or later models like GLM-5 may produce better results.
  • Configuration: At $0.2 input and $1.1 output per million tokens, GLM 4.5 Air is designed for workloads where unit economics matter. Estimate your monthly token volume to compare total cost against heavier alternatives.
  • Configuration: GLM 4.5 Air uses the same API interface as GLM-4.5, so switching between the two requires only changing the model identifier.
  • Zero Data Retention: Zero Data Retention is offered on a per-provider and model basis. See the documentation for details.
  • Authentication: AI Gateway authenticates requests using an API key or OIDC token. You do not need to manage provider credentials directly.

Copy link to headingWhen to Use GLM 4.5 Air

Best for

  • High-volume production pipelines: Low latency and cost efficiency per request outweigh peak reasoning depth
  • Classification and extraction tasks: Competent language understanding without extended chain-of-thought overhead
  • Conversational applications: Many concurrent users where response speed directly affects user experience
  • Summarization workflows: Large document sets where throughput determines pipeline feasibility
  • Development and prototyping: Fast iteration cycles benefit from quick model responses

Consider alternatives when

  • Deep reasoning needed: The full GLM-4.5 or GLM-5 provides deeper deliberation capabilities for complex multi-step planning
  • Vision or image understanding: GLM-4.5V builds on GLM-4.5-Air with multimodal input support
  • Advanced code generation: GLM-4.6 and GLM-4.7 include targeted coding improvements
  • Lowest cost simple tasks: Evaluate flash-tier models in the GLM lineup when further capability tradeoffs are acceptable

GLM 4.5 Air fills the efficiency tier in Z.AI's GLM-4.5 generation. It offers the practical balance teams need when deploying language models at scale: broad general capability with the speed and cost profile that high-volume production demands.

Your use is subject to Z.AI's Terms & Privacy Policies.