Z.ai reports first half results, sales surge

Published September 1, 2026

Z.ai, the Chinese AI company behind the leading open GLM models, reported its first half results. Z.ai launched its IPO on the Hong Kong Stock Exchange in January.

In a press release, Z.ai reported first half revenue of $141.95 million, up nearly 400%, with a net loss of $308.33 million. The loss for the period was lower than the $350.87 million reported a year ago.

The company's R&D expenses were $317.14 million, up 33.6% from a year ago.

Z.ai said its open platform and API cloud business accounted for $122.8 million, or 86% of revenue. The company's enterprise general model business, which covers on-premises deployments was $9.98 million, down 54.6% from a year ago.

The company touted its GLM-5.3 inference models and said it has more than 7.4 million enterprise and developer users on its model-as-a-service platform. Z.ai also said it scaled an inference cluster with more than 100,000 Chinese AI chips and reduced token costs by 80%.

Key excerpts from the statement:

  • "Since the beginning of 2026, the Company has fully implemented an All-in-Infra strategy, investing in both training and inference, and integrating domestic chips as a primary inference compute resource. It has achieved large-scale, low-cost inference across a cluster of over 100,000 domestic chips, with unit token inference costs declining by 80% from the beginning of the year. Given the persistent global supply constraints on compute, this path ensures that the Company’s cost curve remains within its own control."
  • "Industry evaluation criteria are also shifting: from single-turn response quality and token consumption volume to long-horizon task success rate, unit intelligence cost, result verifiability, and the actual business outcomes delivered to customers."
  • "We have focused essentially on one thing: training more powerful models. When the industry was hot, we did not divert resources to roll out multiple product lines; when the industry cooled, we did not pause model iteration."
  • "Leading a point-in-time leaderboard is no longer the most important thing. What truly matters is who can sustainably bring more capable models to market at a faster pace and with better cost-efficiency."
  • "The inference of GLM-5.3-Flash marks the first time the company has fully deployed a domestic chip cluster to serve ultra-large-scale real-world traffic, with chips interconnected via a self-developed high-bandwidth network. The way this work was accomplished is itself worth recording: the entire inference engine was substantially accelerated by an infra agent powered by GLM-5.3, which assisted engineers in developing and optimizing operators, diagnosing performance bottlenecks, and improving the deployment service stack. The model optimizes the system, and the system hosts the model – this loop ran internally for the first time. Operator development cycles were shortened by half."