LLM Vendor Capacity and Integration Timeline: A Complete Breakdown ⚠️ 先说实话 / Truth Be Told

Precise capacity data for 300 vendors does not exist in any public source. While llm-stats.com claims to track “300+ AI models,” this refers to the number of model variants, not independent vendors. Across three major tracking platforms (ZeroEval, Artificial Analysis, llm-stats.com), I identified only 28 independent vendors serving as the global backbone for API access. To reach the number 300, one would need to include every fine-tuned variant, quantized version, and minor player in China’s “War of a Hundred Models,” whose GPU counts remain strictly guarded trade secrets. Below is the verifiable public data:

I. 28 Major Global LLM Vendors (Tracked by ZeroEval)

厂商 / Vendor 国籍 / Country 主力模型 / Flagship Models

Method OpenAI

GPT-5.5, GPT-5.1, o4-mini API Anthropic

Claude Mythos, Opus 4.8, Sonnet 4.6 API Google DeepMind

Gemini 3.5 Flash, 3.1 Pro API Meta

Llama 4, Llama 4 Behemoth

Open Weights Microsoft

API / Product xAI

Grok 4.3, Grok 4 API Amazon

API NVIDIA

Nemotron-3 Ultra 550B

API + Open Source Cohere

Command R+, Aya API AI21 Labs

Jamba 2, Jamba

API + Open Source IBM

Granite

Open Source Mistral AI

Mistral Medium 3.5

API + Open Source Nous Research

Hermes 3, Nous-Hermes

Open Source 大模型厂商产能与接入时间表 / LLM Vendor Capacity & Timeline Breakdown

厂商 / Vendor 国籍 / Country 主力模型 / Flagship Models

Method

V4 Pro, R1-0528

API + Open Source Alibaba / Qwen (通义千问)

Qwen 3.7 Max, 3.7 Plus

API + Open Source

API (Volcengine) Moonshot AI (月之暗面)

Kimi K2.6, K2 Think

API + Partial OS Zhipu AI (智谱 AI)

GLM 5.1, ChatGLM

API + Open Source Baidu (百度)

API MiniMax

MiniMax-M3, M2.7, Hailuo API

Step 3.7 Flash, Step-2 API Xiaomi (小米)

MiLM

Product Embedded Meituan (美团)

Internal Use Only

MiniCPM 5, CPM-Bee

Open Source LG AI Research

EXAONE API Sarvam AI

OpenHathi, Sarvam-1 API Inception 🇦酋 / UAE Jais

Open Source Unisound (云知声)

UniGPT API

II. GPU Capacity Estimation (Public Data, 2025)

Sources: Corporate financial reports, SemiAnalysis, The Information, public procurement orders, Omdia estimates. Exact numbers are trade secrets; ranges represent third-party estimates. 大模型厂商产能与接入时间表 / LLM Vendor Capacity & Timeline Breakdown

Rank

Vendor

Est. GPU Scale

Chip Type

(H100) Equiv. (H100) 关键说明 / Key Remarks 🥇 1 Google 1,000,000+ TPU v5p/v6 ~1,500,000

TPUs don’t directly benchmark to H100; largest single AI compute cluster. 🥈 2 Meta 600,000 H100 600,000

Two 24,576 H100 clusters, completed by late 2024. 🥉 3 Microsoft 400,000-500,000 H100/B200 ~450,000 Azure AI 基础设施;Stargate 项目

Azure AI infra; Stargate project targets $100B+. Amazon 200,000+ Trainium2 ~200,000

Custom chip route; primary compute backbone for Anthropic. xAI 200,000 H100/B200 ~200,000 Memphis Colossus 集群;100K

Memphis Colossus cluster; 100K H100 built in just 122 days. OpenAI 100,000+ H100/B200 ~100,000 依赖 Microsoft Azure;Stargate

Relies on MS Azure; Stargate 2028 long-term goal far exceeds this. ByteDance 100,000+ H100/H800/ B200 ~100,000

Largest AI compute buyer in China; heavily affected by export controls. Anthropic 50,000-100,000 H100/ Trainium2 ~75,000

AWS + GCP dual-cloud; raised over $8B+ in funding. Alibaba 50,000-100,000 H800/A100 ~50,000

Domestic chip substitution is progressing in parallel. Huawei 50,000+ Ascend 910B/ C ~40,000

Fully autonomous and self-controlled AI chip ecosystem. Tencent 50,000+ H800 ~50,000

Powers the Tencent Hunyuan LLM matrix ecosystem. 大模型厂商产能与接入时间表 / LLM Vendor Capacity & Timeline Breakdown

Rank

Vendor

Est. GPU Scale

Chip Type

(H100) Equiv. (H100) 关键说明 / Key Remarks DeepSeek 30,000-50,000 H800/H100 ~40,000

Extremely high efficiency; trained V3 for only $5.6M. SenseTime 30,000-50,000 A100/H800 ~35,000

SenseCore AI DC; restricted by US Entity List. Baidu 30,000-50,000 A100/H800 ~35,000

In-house Kunlun chips are being deployed in parallel. Moonshot AI 10,000-30,000 H800 ~20,000

$1B+ Mainly hosted on Alibaba Cloud; raised over $1B+ in total. Zhipu AI 10,000-20,000 — ~15,000

Tsinghua spinoff; receives strong strategic government support. MiniMax 10,000-20,000 — ~15,000

Focuses on multimodal/video generation; high GPU burn rate. 01.AI 5,000-10,000 — ~7,000

Founded by Dr. Kai-Fu Lee. StepFun 5,000-10,000 — ~7,000

Founded by a team of former Microsoft executives. iFlytek 10,000+

Compute infrastructure is deeply integrated with Huawei.

III. API Integration Timeline 时间 / Timeline 里程碑事件 / Milestone Event 接入方式 / Access Method 2024-03 AI21 Jamba 发布 AI21 Jamba released

API + Open Source 2024-04 Cohere Command R+ 发布 Cohere Command R+ released API 大模型厂商产能与接入时间表 / LLM Vendor Capacity & Timeline Breakdown 时间 / Timeline 里程碑事件 / Milestone Event 接入方式 / Access Method 2024-05

GPT-4o & ByteDance Doubao API launched API 2024-07 Mistral Large 2 发布 Mistral Large 2 released

API + Open Source 2024-09

Qwen 2.5 series fully open-sourced

Open Source 2024-12 DeepSeek V3 发布(每百万 token 仅 $0.27) DeepSeek V3 released ($0.27 per million tokens)

API + Open Source 2025-01

DeepSeek R1 reasoning model launched

API + Open Source 2025-02

Grok 3 formally released X Premium+ / API 2025-03 Gemini 2.5 Pro 达到 GA 阶段 Gemini 2.5 Pro reaches GA stage Google AI Studio 2025-05 Claude Opus 4 达到 GA 阶段 Claude Opus 4 reaches GA stage API 2025-06

Moonshot AI Kimi K2 released API 2025-10

GPT-5.1 series officially launched API 2025-11 Claude Opus 4.5 发布 Claude Opus 4.5 released API 2025-12 DeepSeek V4 Pro / Qwen 3.7 Max / GPT-5.2 集中发布 DeepSeek V4 Pro / Qwen 3.7 Max / GPT-5.2 launched API 2026-01 Grok 4.3 及 Gemini 3 Pro 发布 Grok 4.3 & Gemini 3 Pro released API 2026-04 Claude Mythos 预览版上线 / Kimi K2.6 开源 Claude Mythos Preview / Kimi K2.6 open-sourced

API + Open Source

IV. Panorama of China’s “War of a Hundred Models”

China’s vendor ecosystem extracted from the LLMs-In-China repository (key core players): Tier 1 — 算力 > 50K GPU 等效 / Compute > 50K GPU Equivalent

ByteDance (Doubao/Volcengine) — China’s largest investor in AI compute infrastructure.

Alibaba (Tongyi Qianwen/Qwen) — Maintains the most capable open-source model series domestically. 大模型厂商产能与接入时间表 / LLM Vendor Capacity & Timeline Breakdown

Huawei (Pangu/Ascend Ecosystem) — Full-stack ecosystem with fully autonomous, in-house chip design.

Tencent (Hunyuan) — Backed by heavy integration across cloud services and social ecosystem. Tier 2 — 算力 10K-50K / Compute 10K-50K

DeepSeek — Shook Silicon Valley with extreme architectural efficiency and cost performance.

Baidu (Wenxin Yiyan/ERNIE) — Veteran legacy player with the earliest commercial footprint.

SenseTime (SenseNova) — Computer Vision legacy pioneer fully transitioning into the LLM era.

Zhipu AI (GLM/ChatGLM) — Tsinghua academic roots; a pillar player with strong state support.

Moonshot AI (Kimi) — Pioneer of long-context processing; K2.6 topped open-source charts.

MiniMax (Hailuo AI) — Deep focus on multimodal systems and high-quality video generation.

iFlytek (Spark) — Collaborates deeply with Huawei Ascend for speech AI and vertical enterprise solutions. Tier 3 — 算力 5K-10K / Compute 5K-10K

01.AI (Yi, founded by Kai-Fu Lee) — Highly innovative team with a global market vision.

StepFun (founded by ex-Microsoft execs) — Strong technical DNA in multimodal foundation models.

Baichuan (founded by Wang Xiaochuan) — Deep focus on healthcare and vertical LLM deployment.

ModelBest (OpenBMB) — Specializes in high-efficiency edge-side small language models (MiniCPM).

Unisound (UniGPT) & Xiaomi (MiLM, focused on edge computing embedded in smart hardware devices). Tier 4 — 特定领域 / Tier 4 — Domain-Specific

Meituan, Didi, Pinduoduo — Deeply embedded in internal core business scenarios; no public APIs.

CAS (Zidong Taichu), Shanghai AI Lab (InternLM) — Highly respected academic and open-source research platforms.

Kunlun Tech, Zhihu (Zhihai-Tu), NetEase — In-house models built for niche apps without broad public APIs. 大模型厂商产能与接入时间表 / LLM Vendor Capacity & Timeline Breakdown

V. Summary of Key Trends

  1. The “300” Count Reality: If counting every fine-tuned checkpoint, quantized variant, and niche vertical model, the global ecosystem exceeds 300+. However, only about 30 major platforms provide robust, commercial-grade standard API access.

  2. Extreme Capacity Concentration: The top 5 hyperscalers (Google, Meta, Microsoft, Amazon, xAI) command ~70% of premium global AI training compute. The top 20 represent over 95% of total capacity.

  3. Export Control Dynamics in China: With the H100 export ban, Huawei Ascend 910B/C emerged as the domestic cornerstone. DeepSeek achieving world-class metrics on restricted H800 chips proves that architectural innovation and algorithmic efficiency outweigh raw hardware brute-force.

  4. Crushing of Integration Barriers: Since 2025, every leading player has fully commoditized access via API. Pricing ranges from $15/ M tokens for top-tier reasoning variants down to zero for open-source self-hosting, igniting a brutal market price war.

  5. Unstoppable Rise of Open Weights: The open-sourcing of premier base models (Kimi K2.6, DeepSeek V4, Qwen 3.7, Llama 4) marks a profound shift—the definition of “capacity” has transitioned from raw GPU counts to community deployment speed and ecosystem integration. 大模型厂商产能与接入时间表 / LLM Vendor Capacity & Timeline Breakdown