Name: Towards AI Legal Name: Towards AI, Inc. Description: Towards AI is the world's leading artificial intelligence (AI) and technology publication. Read by thought-leaders and decision-makers around the world. Phone Number: +1-650-246-9381 Email: pub@towardsai.net
228 Park Avenue South New York, NY 10003 United States
Website: Publisher: https://towardsai.net/#publisher Diversity Policy: https://towardsai.net/about Ethics Policy: https://towardsai.net/about Masthead: https://towardsai.net/about
Name: Towards AI Legal Name: Towards AI, Inc. Description: Towards AI is the world's leading artificial intelligence (AI) and technology publication. Founders: Roberto Iriondo, , Job Title: Co-founder and Advisor Works for: Towards AI, Inc. Follow Roberto: X, LinkedIn, GitHub, Google Scholar, Towards AI Profile, Medium, ML@CMU, FreeCodeCamp, Crunchbase, Bloomberg, Roberto Iriondo, Generative AI Lab, Generative AI Lab VeloxTrend Ultrarix Capital Partners Denis Piffaretti, Job Title: Co-founder Works for: Towards AI, Inc. Louie Peters, Job Title: Co-founder Works for: Towards AI, Inc. Louis-François Bouchard, Job Title: Co-founder Works for: Towards AI, Inc. Cover:
Towards AI Cover
Logo:
Towards AI Logo
Areas Served: Worldwide Alternate Name: Towards AI, Inc. Alternate Name: Towards AI Co. Alternate Name: towards ai Alternate Name: towardsai Alternate Name: towards.ai Alternate Name: tai Alternate Name: toward ai Alternate Name: toward.ai Alternate Name: Towards AI, Inc. Alternate Name: towardsai.net Alternate Name: pub.towardsai.net
5 stars – based on 497 reviews

Frequently Used, Contextual References

TODO: Remember to copy unique IDs whenever it needs used. i.e., URL: 304b2e42315e

Resources

Free: 6-day Agentic AI Engineering Email Guide.
Learnings from Towards AI's hands-on work with real clients.
Cheap, Fast, and Good: How Chinese AI Models Broke the Pick-Two Rule
Latest   Machine Learning

Cheap, Fast, and Good: How Chinese AI Models Broke the Pick-Two Rule

Last Updated on July 27, 2026 by Editorial Team

Author(s): Saurabh Singh

Originally published on Towards AI.

Cheap, Fast, and Good: How Chinese AI Models Broke the Pick-Two Rule

Cheap, Fast, and Good: How Chinese AI Models Broke the Pick-Two Rule

The project-management triangle, updated for 2026.

Every engineer knows the triangle. Cheap, fast, good — pick two. It’s held for decades across construction, software, and manufacturing, and for the first years of the LLM era it held for AI too: frontier quality meant frontier prices and frontier wait times.

Then, over eighteen months, a handful of Chinese labs — DeepSeek, Alibaba’s Qwen, Moonshot’s Kimi, Zhipu’s GLM, MiniMax — shipped models that are an order of magnitude cheaper than Western flagships, served fast, iterated faster, and landed within a few points of the best models on Earth. Not one of these claims requires squinting. Each has a leaderboard, a pricing page, or a US-government evaluation behind it.

This article walks through the evidence for all three legs of the triangle, the caveats (there are real ones), what the biggest names in the industry have said about it on the record, and how it is visibly bending the strategies of OpenAI, Anthropic, and Washington.

The moment the triangle cracked

On January 20, 2025, DeepSeek — a Hangzhou lab spun out of a quant hedge fund — released R1, a reasoning model that matched OpenAI’s o1 on math benchmarks (79.8% vs ~79.2% on AIME 2024, per its technical report) while charging $0.55 per million input tokens against o1’s $15 — roughly 27x cheaper.

A week later, on January 27, Nvidia lost $589 billion in market capitalization in a single day — the largest one-day loss for any company in history. The Nasdaq fell 3.1%. DeepSeek’s app displaced ChatGPT at #1 on the US App Store.

The reactions came fast, and from the very top:

“deepseek’s r1 is an impressive model, particularly around what they’re able to deliver for the price.” — Sam Altman, OpenAI CEO, X, Jan 27, 2025

“Deepseek R1 is AI’s Sputnik moment.” — Marc Andreessen, X, Jan 26, 2025

“Jevons paradox strikes again! As AI gets more efficient and accessible, we will see its use skyrocket, turning it into a commodity we just can’t get enough of.” — Satya Nadella, Microsoft CEO, X, Jan 27, 2025

“DeepSeek (Chinese AI co) making it look easy today with an open weights release of a frontier-grade LLM trained on a joke of a budget (2048 GPUs for 2 months, $6M).” — Andrej Karpathy, X, Dec 26, 2024, on DeepSeek-V3

Eighteen months later this is no longer a shock story. It’s a structural story. Let’s take the triangle one leg at a time.

Leg one: Cheap

Here is what a million output tokens costs across the current flagship generation, as of July 2026:

Official API list prices, July 2026. Sources: lab pricing pages; trackers for Qwen, MiniMax, and Gemini.

DeepSeek V4 Pro — an MIT-licensed 1.6-trillion-parameter model with a 1M-token context window — charges $0.435 per million input tokens and $0.87 per million output tokens. Claude Opus 5 charges $5 and $25. GPT-5.6 Sol charges $5 and $30. That’s a 10–30x spread between models a single Artificial Analysis intelligence tier apart (more on quality below). And DeepSeek’s cache-hit input price is $0.003625 — effectively free.

Raw list price isn’t the whole story, because reasoning models burn different amounts of tokens per task. So the better chart is intelligence against blended price:

Artificial Analysis Intelligence Index v4.1 vs blended price, July 2026. The bottom-left cluster is the story.

MiniMax M3 and DeepSeek V4 Pro sit at an Intelligence Index of 44 for $0.12–$0.18 per million blended tokens. Artificial Analysis’s cost-to-run numbers for agent tasks tell the same story at the top end: Kimi K3 completes its AutomationBench tasks at $0.94 per task versus $1.80 for Claude Opus 4.8, while scoring within four index points of it.

“But the $6M training number was fake”

Partly, and it’s worth being precise, because the correction strengthens the cheap story rather than killing it.

DeepSeek’s V3 technical report claimed $5.576M — but explicitly for the final pre-training run only, excluding research, ablations, and infrastructure. SemiAnalysis estimated the company’s total hardware spend at $1.3–1.6 billion, and Anthropic CEO Dario Amodei wrote in his January 2025 essay:

“DeepSeek does not ‘do for $6M what cost US AI companies billions’… DeepSeek’s total spend as a company (as distinct from spend to train an individual model) is not vastly different from US AI labs.”

But Amodei also conceded the part that matters:

“DeepSeek’s team did this via some genuine and impressive innovations, mostly focused on engineering efficiency.”

The honest version: training was never $6M all-in, but the efficiency per dollar was real — sparse mixture-of-experts architectures (V4 Pro activates 49B of 1.6T parameters per token), aggressive caching, and inference engineering that lets these labs profitably charge 10–30x less at the API. CNBC later reported Moonshot’s Kimi K2 Thinking trained for ~$4.6M (a figure Moonshot’s CEO declined to confirm). Whatever the true numbers, the prices are published, and anyone can pay them.

Leg two: Fast

“Fast” means three different things, and Chinese models score on all three.

Fast to serve. On Artificial Analysis’s own measurements, GLM-5.2 streams at 191 tokens/sec with 1.35s time-to-first-token, and Qwen3.7 Max at 197 tokens/sec — faster than Claude Opus 5 (57 tok/s) and GPT-5.6 Sol (66 tok/s), which spend long seconds thinking before the first answer token.

Fast because open. This is the structural advantage. Because the weights are downloadable, anyone can serve them on specialized silicon. Cerebras serves Kimi K2.6 at a measured 981 tokens per second — 6.7x faster than the next-fastest GPU cloud. A closed model’s speed is whatever its one provider offers. An open model’s speed is a market.

Speed is now a hosting decision, not a model property. Note the amber bar.

Become a Medium member

Fast to ship. Count flagship releases with pinned public dates from January 2025 through July 2026: the five Chinese labs shipped 29; OpenAI, Anthropic, and Google DeepMind shipped 15. DeepSeek alone went V3 → R1 → R1–0528 → V3.1 → V3.2 → V4 in eighteen months, cutting prices 50% mid-stream.

Every dot is a dated flagship release. Chinese labs iterate roughly twice as fast in aggregate.

Leg three: Good

This is the leg skeptics doubt, so let’s use only neutral scoreboards.

The best open-weight model sits four points off the global frontier.

  • Artificial Analysis Intelligence Index (July 2026): Kimi K3 scores 57 — behind only Claude Opus 5/Fable 5 (61/60) and GPT-5.6 Sol (59), and ahead of every Google model. Artificial Analysis’s own headline: “Kimi K3 achieves #3 in the Artificial Analysis Intelligence Index, comparable to Opus 4.8 and GPT-5.5.”
  • LMArena (July 2026): Kimi K3 sits #10 overall among all models by human preference — and holds #1 on the Frontend Code arena, ahead of Claude Fable 5.
  • Humanity’s Last Exam: Kimi K2 Thinking’s 44.9% with tools in November 2025 was the first time an open-weights model claimed a state-of-the-art frontier benchmark; Moonshot’s K2.6 table (vendor-reported) later showed 54.0 against GPT-5.4’s 52.1.
  • GPQA Diamond is effectively saturated: NIST’s independent evaluation put DeepSeek V4 Pro at 90% vs GPT-5.5’s 96% — a knowledge gap of a few points, not a generation.
  • Stanford’s AI Index 2026: the performance gap between the top US and top Chinese model was 2.7% as of March 2026, down from double digits in May 2023. Epoch AI puts the average time lag at ~7 months — and notes the US-China gap is now nearly identical to the closed-vs-open gap, because almost every leading Chinese model ships its weights.

And the ecosystem has voted. Qwen became the first model family in history to pass 1 billion Hugging Face downloads, holds over 50% of global open-model downloads, and underlies ~40% of all new fine-tune derivatives on the Hub. Nathan Lambert of Interconnects put it flatly: “Qwen alone is roughly matching the entire American open model ecosystem.”

“We’re relying a lot on Alibaba’s Qwen model. It’s very good. It’s also fast and cheap.” — Brian Chesky, Airbnb CEO, Bloomberg interview, Oct 2025 — explaining why Airbnb’s production customer-service agent doesn’t primarily run on OpenAI

That quote is the entire trilemma, from the CEO of a $100B American company, in eleven words. (By May 2026, Congress was formally demanding answers from Chesky about it.)

The honest part: where the triangle still holds

A credible version of this article has to show the other side, because it exists and it’s measurable.

On neutral agentic harnesses, the gap is real: ~25 points on Terminal-Bench 2.1, ~30 on ARC-AGI-2.

  • Long-horizon agentic work. On the official Terminal-Bench 2.1 leaderboard, Claude Code + Fable 5 scores 83.8% and the best listed Chinese entry (GLM-5.1) scores 58.7%. NIST’s CAISI evaluation of DeepSeek V4 Pro found it ~8 months behind the US frontier overall, with the widest gaps exactly where models act rather than answer: ARC-AGI-2 (46 vs 79), cyber CTFs (32 vs 71), agentic software engineering (44 vs 78). Simon Willison’s July 2026 verdict on Kimi K3 was the same in prose: the remaining gap is robust tool-calling reliability across long conversations.
  • Hallucination and safety. On Artificial Analysis’s Omniscience testing, DeepSeek V4 Pro answers even when it doesn’t know — a 94% hallucination rate on unknown-answer questions, versus ~36% for Claude Opus 4.7. NIST’s earlier evaluation found R1–0528 complied with 94% of overtly malicious requests under a common jailbreak, versus 8% for US reference models.
  • Benchmark hygiene. Vendor-run tables and neutral harnesses diverge hardest on agentic benchmarks (GLM-5.2’s vendor table claims 81.0 on Terminal-Bench; the official board shows no such entry). Contamination-filtered suites like SWE-rebench have documented score drops for several Chinese models on fresh tasks — though GLM held up, and contamination cuts both ways: one study found Claude had absorbed ~12% of SWE-bench Pro tasks into training data. OpenAI has abandoned SWE-bench Verified entirely.
  • Cheap is eroding at the top. Kimi K3 launched at $3/$15 — Western-band pricing. The-decoder called it “the end of super-cheap Chinese AI.” Moonshot and DeepSeek now have real revenue ($100M→$200M ARR and ~$220M ARR respectively) and investor valuations to defend.

So the precise claim is not “Chinese models are the best.” It is: for the fat middle of real workloads — chat, extraction, RAG, coding assistance, high-volume agents — models exist that are simultaneously ~95% as good, several times faster to first token, and 10–30x cheaper. That combination was supposed to be impossible.

How they did it (in one paragraph)

No single trick. Sparse MoE architectures that activate 2–4% of parameters per token; training-efficiency innovations Amodei himself called “genuine and impressive”; a brutal domestic price war among five labs shipping every quarter; export controls that made compute scarcity a forcing function for efficiency; a policy of releasing weights, which recruits the world’s inference providers, fine-tuners, and researchers as an unpaid distribution and R&D arm; and — per Anthropic’s accusations against Qwen, DeepSeek, Moonshot, and MiniMax — distillation of frontier US models’ outputs, which Nvidia’s Jensen Huang shrugs off as “fundamental to intelligence” and US AI-czar David Sacks calls theft. Yann LeCun’s January 2025 framing remains the cleanest lens:

“To people who see the performance of DeepSeek and think: ‘China is surpassing the US in AI.’ You are reading this wrong. The correct reading is: ‘Open source models are surpassing proprietary ones.’”

What it’s doing to OpenAI and Anthropic

The competitive response is the strongest evidence that the trilemma break is real. Companies don’t restructure pricing and philosophy over vibes.

OpenAI reversed a six-year closed-weights policy. Days after R1, Altman told Reddit: “I personally think we have been on the wrong side of history here and need to figure out a different open source strategy.” In August 2025, OpenAI shipped gpt-oss-120b and gpt-oss-20b — its first open-weight models since GPT-2 — under Apache 2.0, framed explicitly as keeping the world “building on an open AI stack created in the United States.” GPT-5 launched the same week at $1.25/$10, far below prior flagship pricing; the budget GPT-5.6 Luna tier now matches Chinese cost-per-intelligence almost exactly. That tier exists because the floor moved.

Anthropic cut Opus pricing 67% (from $15/$75 to $5/$25 with Opus 4.5, November 2025) while going the opposite direction on access: banning Chinese-controlled entities from Claude, and in June 2026 accusing Alibaba’s Qwen team of running ~28.8 million Claude conversations through ~25,000 fake accounts to distill coding capabilities. Amodei has spent 2026 arguing capable open-weight models are “a serious concern” and that chip exports to China are, in his Davos phrasing, “a bit like selling nuclear weapons to North Korea and bragging that Boeing made the casings.”

The market share moved under both of them. The a16z/OpenRouter study of 100T+ routed tokens found US models fell from ~70% of token volume in June 2025 to ~30% in June 2026; in the week of February 9–15, 2026, Chinese models processed more tokens than American ones for the first time. DeepSeek alone is the single largest provider on the router at 16.3%. The crucial caveat: revenue skews the opposite way — US closed models still capture most of the actual spend. Chinese models won the token war; they have not yet won the money war.

Tokens flipped; dollars haven’t. Both facts matter.

Washington became a market participant. The H20 chip ban of April 2025 was reversed by July 2025 (with Nvidia paying the US Treasury 15% of China revenue); January 2026 brought case-by-case H200 export licenses plus a 25% surcharge; June 2026 brought the “Gold Eagle” program giving the government pre-release testing access — under which the White House temporarily blocked US frontier releases even as CNBC noted the crackdown “opens door for Chinese model makers to close gap.”

And Jensen Huang, whose company sits underneath all of it, has completed a striking arc: from “China is right behind us” (April 2025), to “China is going to win the AI race” (November 2025, walked back within hours to “nanoseconds behind”), to July 2026, post-Kimi-K3, defending the other side outright: “These Chinese models are excellent. Open-source models that are excellent should be used.”

Where this goes

Three trend lines worth watching, all sourced above:

  1. The gap is a lag, not a wall. Every neutral measure — Epoch (~7 months), NIST (~8 months), Stanford (2.7% Elo) — describes a delay, and the July 2026 Kimi K3 release (2.8T parameters, #3 on the Intelligence Index, a second Nvidia-rattling market moment) suggests the delay is stable or shrinking. The bet that Chinese models stay permanently a tier behind has lost every six-month checkpoint since January 2025.
  2. Open weights are becoming Chinese infrastructure. With Llama below 1% of routed volume and Meta pivoted closed, the default substrate for the world’s fine-tuners, researchers, and cost-sensitive enterprises is Qwen/DeepSeek/GLM. The US answer so far is gpt-oss and the ATOM Project’s call for a serious American open-model lab — backed by OpenAI’s CSO, Hugging Face’s CEO, and PyTorch’s creator. Whether that consolidates into real investment is the biggest open question of 2026–2027.
  3. The trilemma is reforming one level up. Frontier agentic capability — the thing Terminal-Bench and ARC-AGI-2 measure — still commands frontier prices, and even Chinese labs now price it that way (K3 at $3/$15). The pick-two rule didn’t die; it retreated to the frontier. Everything behind the frontier is now cheap, fast, and good — and the frontier keeps moving outward every seven months, with last year’s impossible trilemma-break becoming this year’s commodity.
  4. The triangle held for decades because physics and payroll enforced it. It broke because software costs collapse when weights are free, silicon competes for open models, and five labs in a price war ship every quarter. Satya Nadella saw the consequence within hours of the DeepSeek shock, and it remains the best one-line summary of the whole story: when intelligence gets cheap enough, you don’t buy less of it. You buy vastly, insatiably more.

All benchmark figures, prices, and quotes verified against the sources below as of July 26, 2026. Prices are list prices; benchmark scores are pinned to the dates shown, because this field re-writes its own leaderboard roughly every seven months.

Sources

Join thousands of data leaders on the AI newsletter. Join over 80,000 subscribers and keep up to date with the latest developments in AI. From research to projects and ideas. If you are building an AI startup, an AI-related product, or a service, we invite you to consider becoming a sponsor.

Published via Towards AI


Towards AI Academy

We Build Enterprise-Grade AI. We'll Teach You to Master It Too.

15 engineers. 100,000+ students. Towards AI Academy teaches what actually survives production.

Start free — no commitment:

6-Day Agentic AI Engineering Email Guide — one practical lesson per day

Agents Architecture Cheatsheet — 3 years of architecture decisions in 6 pages

Our courses:

AI Engineering Certification — 90+ lessons from project selection to deployed product. The most comprehensive practical LLM course out there.

Agent Engineering Course — Hands on with production agent architectures, memory, routing, and eval frameworks — built from real enterprise engagements.

AI for Work — Understand, evaluate, and apply AI for complex work tasks.

Note: Article content contains the views of the contributing authors and not Towards AI.