Results

Hugging Face released its open model report and it's a must read. One big headline is how Qwen-based models account for more than 151,000 derivatives.

According to the report, which is a must read:

"Qwen has become one of the largest foundations in the open model ecosystem. Qwen-based models now account for 151,448 derivatives on the Hub, 2.6× Meta’s total footprint and 4.7× the Llama repositories specifically. Google follows with 82,506 derivatives. The third-largest source is Unsloth, a community account publishing quantized and fine-tuning-ready builds, many of which further extend the Qwen ecosystem.

Qwen derivatives have increased at roughly 180–210 new repositories per day throughout the first seven months of 2026, showing that adoption is not driven only by individual launches. Qwen has become part of the default workflow for developers deciding what models to fine-tune and deploy."

Z.ai released GLM-5.3, a model post-trained on GLM-5.2 to get better coding and cybersecurity capabilities.

In a blog post, Z.ai revealed that GLM-5.3 outperformed GLM-5.2 on multiple benchmarks. The company said every advance in GLM-5.3 was due to post training.

"For GLM-5.3, we pushed environment scaling toward tasks that look less like coding exercises and more like real units of expert work. The environments now cover a much broader range of production workflows, with tasks designed around how engineering and research work is actually carried out in practice."

Google released Gemini 3.7 Flash, its flagship workhorse model. Gemini 3.7 Flash is designed for coding, knowledge work and agents.

The company outlined the usual benchmarks, but the biggest takeaway is the frontier models are competing on price now as enterprises freak out over AI budgets.

Note that Google's pricing in the chart below is introductory. But until Jan. 1, 2027, Gemini 3.7 Flash aims to undercut rivals.

Gemini 3.7 Flash


OpenAI named Dali Rajic Chief Revenue Officer. Rajic, formerly of Google's Wiz unit, replaces Denise Dresser, who joined the company in December to build out the enterprise business.

In a blog post, OpenAI said Dresser is leaving to pursue other opportunities. Before OpenAI, Dresser was a long-time Salesforce executive. Rajic was President and Chief Operating Officer of Wiz, which is now owned by Google. He had previous stints at Zscaler and AppDynamics.

Writer, which specializes in marketing and revenue AI workflows, released its Palmyra X6 flagship model and upgrades to its Writer Agent harness. A few points stuck out in the announcement.

  • Writer is targeting token spending. The company said Writer Agent operates at a 52% lower cost with 48% improvement in speed when using Palmyra X6.
  • Palmyra X6 was trained on top of GLM-5.2.
  • Writer benchmarked its latest model and harness compared to frontier models and outperformed.
  • Palmyra X6 completes tasks in 26 seconds on average, generates 82 tokens per second and can work unattended for up to 8 hours.

MongoDB launched a set of tools that bring automated embeddings to its Atlas platform via Voyage AI. The company said the Atlas Embedding and Reranking API, voyage-code-4, and vector search in Atlas Stream Processing all are designed to improve retrieval accuracy. MongoDB also announced Atlas Managed MCP Server, a fully hosted service that connects Claude Code, Codex, Grok Build, and Devin to MongoDB Atlas.

Mistral said it will expand its platform to open models beyond its own. Mistral's platform will support third-pary open models starting with Z.ai's GLM-5.2. The third-party open models will run on Mistral's infrastructure, regional controls and service commitment used by its own models.

The effort was outlined in a post detailing Mistral's sovereign AI plans. Notably, Mistral said its Regional Endpoints are generally available. The service lets customers choose whether their inference runs in the US or Europe. The company also said its Mistral Priority Tier is now in public preview.

Cerebras CEO Andrew Feldman had interesting quotes on the company's second quarter earnings call. It's the AI equivalent of the "speed kills" phrase in sports. He said:

  • "The market is realizing that speed is not a benchmark item. Speed changes user engagement, it changes agentic performance, and it changes AI productivity. Fast inference unlocks new applications and new markets."
  • "On the capabilities front, in the second quarter, we delivered support for OpenAI's GPT-5.6 Sol, the largest and most capable of the frontier models. In fact, Cerebras serves 5.6 Sol at a speed that is 10x faster. With GPT-5.6 Sol, this lays to rest any of the remaining concerns regarding our ability to support large frontier models."
  • "Speed is critical for user experience. Throughput is critical for inference economics."
  • "Cerebras wants more throughput without giving up speed. Herein is the strength of our disaggregated solution. It delivers Cerebras speed with 5x higher throughput. Increasing throughput by 5x while keeping our industry-leading speed has a profound impact on the economics of token generation. It means up to 5x as many high-speed, high-value tokens are made by each Cerebras system. More tokens per system at lower cost means more revenue and more gross margin."

In the big picture, Cerebras' quarter was a solid building block. Wall Street wanted more an shares were down 16% premarket.