GPUs & AI Accelerators

Merchant processors for training models and serving their outputs

Heat (est.) Spend (est.)

The single largest line of AI spend, allocated by the vendor rather than bought freely, and the reference point for the whole market.

Merchant AI processors execute the matrix and attention operations behind training and inference. GPUs dominate flexible deployment, while regional suppliers and specialized architectures compete on software access, memory use, response time and power consumption.

Why now

Rubin and MI455X move the GPU roadmap forward while Nvidia commercializes Groq-derived LPUs. Inference specialists raise new capital, and Enflame joins the wave of listed Chinese accelerators. Announced capacity and preproduction designs are not equivalent to installed supply.

Merchant Training Accelerators

Programmable GPUs and tensor processors sold to AI developers for training and mixed training-inference workloads. Software compatibility, usable memory and deployment scale determine adoption.

Heat (est.)
  1. Blackwell Ultra and Rubin GPUs supply CUDA training; Vera Rubin NVL72 is the current Rubin platform

    leader
    NVDA · USExposure2026 AI data-center accelerator revenue: ~81% (est.)
  2. MI350-series GPUs and July-launched MI455X serve training; Helios volume deployments are planned for 2H 2026

    challenger
    AMD · USExposure
  3. Ascend 910C runs Chinese training workloads; 950DT is scheduled for Q4 2026, after the inference-focused 950PR

    major
    Private · CNExposure
  4. Habana Gaudi 3 remains its merchant training ASIC; Crescent Island is an inference GPU, not its replacement

    challenger
    INTC · USExposure
  5. WSE-3 wafer-scale processor powers CS-3 training without partitioning a model across conventional GPU dies

    challenger
    CBRS · USExposure
  6. Siyuan MLU cloud processors and NeuWare distributed-training libraries support Chinese model-training clusters

    challenger
    688256.SS · SSE · CNExposure
  7. DCU accelerator cards use DTK's HIP, math and collective libraries to run distributed training in Chinese clusters

    challenger
    688041.SS · SSE · CNExposure
  8. MTT S5000 PH100 GPUs train Qwen models through MUSA and FlagOS, with published multi-node training validation

    challenger
    688795.SS · SSE · CNExposure
  9. MXC600 training-inference chips power C600 modules with MXMACA software and scalable MetaXLink interconnect

    challenger
    688802.SS · SSE · CNExposure
  10. Biren166M/L GPU modules support distributed model training; 166L powers the deployed Guangyue 128-card supernode

    challenger
    6082.HK · HKEX · CNExposure
  11. CloudBlazer DTU accelerators and TopsRider software implement distributed cloud training for domestic AI customers

    challenger
    688801.SS · SSE · CNExposure
  12. TianGai 150/300 GPU families supply programmable training compute; ZhiKai remains the distinct inference line

    challenger
    9903.HK · HKEX · CNExposure
  13. T-Head's Zhenwu M890 supports FP32-to-FP4 training and inference and is supplied to external enterprise customers

    challenger
    BABA · CNExposure
  14. SN40L reconfigurable dataflow chips support model customization; SN50 shifts emphasis toward agentic inference

    niche
    Private · USExposure
  15. Blackhole and Wormhole Tensix processors use programmable tensor tiles for open-stack AI compute

    challenger
    Private · USExposure
  16. MN-Core 2 chips are sold in developer systems and boards for AI training and scientific compute

    niche
    Private · JPExposure
  17. TX81 reconfigurable modules target model training and long-context inference with Mesh/Torus cluster links

    emerging
    Private · CNExposure
  18. Graphcore's documented Bow-2000 and Poplar train neural models; continued volume availability is not independently confirmed

    niche
    9984.T · TSE · JPExposure

Sources: reuters.com, mthreads.com, metax-tech.com, birentech.com, tsingmicro.com, amd.com

China AI Accelerators

Domestic AI processors compete for Chinese cloud and enterprise deployments as import rules and local procurement policies change. Compiler maturity and access to usable memory constrain substitution.

Heat (est.)
  1. Ascend 950PR powers Atlas 350 inference; the 910C and planned 950DT address training and decode

    leader
    Private · CNExposure
  2. Siyuan MLU processors use NeuWare and MagicMind for domestic cloud AI; no unverified MLU690 volume claim is assumed

    major
    688256.SS · SSE · CNExposure
  3. DCU products such as K100_AI and BW1000 use the DTK stack; Hygon remains independent after the terminated Sugon merger

    major
    688041.SS · SSE · CNExposure
  4. PH100-based MTT S5000 supports FP8 through FP64 and MUSA software for training and inference

    challenger
    688795.SS · SSE · CNExposure
  5. C600 OAM GPUs use MXMACA software and MetaXLink; C500 remains a PCIe training-inference option

    challenger
    688802.SS · SSE · CNExposure
  6. Biren166M/L training-inference modules and 166C inference cards extend its BR100 GPU lineage

    challenger
    6082.HK · HKEX · CNExposure
  7. TianGai cloud training GPUs and ZhiKai inference GPUs form its domestic software-compatible compute portfolio

    challenger
    9903.HK · HKEX · CNExposure
  8. CloudBlazer DTU accelerators and TopsRider serve training and inference; Enflame completed its September 11 STAR listing

    challenger
    688801.SS · SSE · CNExposure
  9. Kunlunxin sells AI accelerators externally; its announced M100 inference and M300 training roadmap targets 2026 and 2027

    major
    BIDU · CNExposure
  10. Zhenwu M890 and SAIL support FP32-through-FP4 training and inference, with external enterprise shipments disclosed

    challenger
    BABA · CNExposure
  11. VA10 inference accelerators and VastStream software serve data-center neural and video workloads

    emerging
    Private · CNExposure
  12. Goldwasser GPU+ chips and UL modules target cloud-edge inference with its own compiler and runtime

    niche
    Private · CNExposure
  13. SOPHON BM1684X tensor processors support inference cards and multi-chip video-analysis deployments

    niche
    Private · CNExposure
  14. PACE 2 supplies domestic optical/electronic tensor computation; its Photowave interconnect business is outside this segment

    niche
    1879.HK · HKEX · CNExposure
  15. JM and Jinghong GPUs support DeepSeek models; Chengheng's newly launched CH37 extends compute to edge AI

    challenger
    300474.SZ · SZSE · CNExposure
  16. TX81 cloud modules and mass-produced TX5 edge chips use the company's reconfigurable processing architecture

    challenger
    Private · CNExposure
  17. Antoum sparse-compute chips power S4/S10 inference cards; sparsity-dependent throughput is workload-specific

    niche
    Private · CNExposure
  18. LX PRO TrueGPU cards provide 24GB GDDR6 and OpenCL compute for AI PCs and workstations, alongside graphics

    emerging
    Private · CNExposure
  19. Fenghua GPUs combine graphics with OpenCL AI compute; the published Fenghua 1 platform supports neural frameworks

    niche
    Private · CNExposure

Sources: jingjiamicro.com, tsingmicro.com, kisacoresearch.com, lisuantech.com, reuters.com, metax-tech.com

LLM Inference Processors

Inference chips execute model prompts and generate tokens for deployed services. Buyers trade model coverage, response latency and memory capacity against software migration and electricity costs.

Heat (est.)
  1. Groq-derived LP30 LPUs entered full production in August 2026; Rubin GPUs handle complementary inference stages

    leader
    NVDA · USExposure
  2. MI355X and MI455X target large-model serving; OpenAI and Meta each contracted up to 6GW across generations

    challenger
    AMD · USExposure
  3. AI200 targets 2026 inference deployments under HUMAIN's 200MW plan; AI250 commercial sampling targets mid-2027

    challenger
    QCOM · USExposure
  4. Crescent Island Xe3P inference GPU targets 2H 2026 sampling with 160GB LPDDR5X in Intel's reference design

    emerging
    INTC · USExposure
  5. WSE-3 powers low-latency token generation; OpenAI agreed to deploy 750MW of Cerebras inference capacity

    challenger
    CBRS · USExposure
  6. SN50 RDUs target token decode alongside Intel Xeon hosts; SoftBank Corp. is the initial deployment partner

    challenger
    Private · USExposure
  7. Corsair digital in-memory accelerators entered production in June 2026 for hyperscalers and frontier labs

    challenger
    Private · USExposure
  8. RNGD tensor-contraction processors run LG AI Research's EXAONE; LG CNS is extending enterprise deployment

    challenger
    Private · KRExposure
  9. REBEL-Quad combines four AI chiplets and 144GB HBM3E for large-model prefill and decode

    challenger
    Private · KRExposure
  10. Blackhole Tensix chips and TT-Metal software execute local and server LLM inference on programmable tiles

    niche
    Private · USExposure
  11. Atlas FPGA inference is deployed at Oracle; the Asimov custom ASIC targets production in 2H 2027

    emerging
    Private · USExposure
  12. Company-disclosed low-voltage prefill hardware and cluster-scale decode memory underpin early Jane Street deployments

    emerging
    Private · USExposure
  13. Bertha LPU targets token decode with 192GB LPDDR5X; Forte FPGA chips underpin the earlier HX platform

    emerging
    Private · KRExposure
  14. Spyre's 32 AI cores provide PCIe-attached LLM inference for IBM z17, generally available since October 2025

    niche
    IBM · USExposure
  15. Retains LPU inference technology licensed non-exclusively to Nvidia; independent GroqCloud continues operating

    niche
    Private · USExposure
  16. HC1 hardwires Llama 3.1 8B into model-specific silicon; HC2 remains a roadmap and the AMD acquisition is pending

    emerging
    Private · CAExposure

Sources: reuters.com, nvidianews.nvidia.com, qualcomm.com, intelcapital.com, furiosa.ai, rebellions.ai

Memory-Centric AI Compute

Architectures put model weights close to arithmetic, using SRAM, analog arrays or tightly coupled DRAM. They target memory-bound inference, but several designs remain preproduction or limited to edge models.

Heat (est.)
  1. Groq 3 LP30 uses on-chip SRAM for low-latency decode, complementing Rubin GPU memory capacity

    major
    NVDA · USExposure
  2. WSE-3 distributes 44GB of SRAM beside wafer-scale compute cores, reducing movement of activations and working data

    major
    CBRS · USExposure
  3. Corsair uses digital in-memory compute and on-chip SRAM to reduce LLM decode data movement

    leader
    Private · USExposure
  4. Licensed LPU architecture keeps model data in on-chip SRAM under compiler control; GroqCloud retains this compute lineage

    niche
    Private · USExposure
  5. AI250's High Bandwidth Compute architecture targets near-memory inference; commercial sampling is planned for mid-2027

    emerging
    QCOM · USExposure
  6. MN-Core L1000 stacks DRAM above logic for generative inference; the company still describes it as under development

    emerging
    Private · JPExposure
  7. Interleaved memory and compute targets frontier-model inference; first commercial chips are expected in 2027

    emerging
    Private · GBExposure
  8. The MatX One keeps weights in SRAM and KV caches in HBM to combine fast decoding with long-context support

    emerging
    Private · USExposure
  9. Asimov's memory-focused ASIC targets large-model inference; Atlas currently implements the approach on FPGAs

    emerging
    Private · USExposure
  10. Jotunn8 inference processor targets the memory wall; July funding supports commercial deployment

    emerging
    Private · FRExposure
  11. CRAFTWERK pairs programmable inference ASICs with processor-memory co-design; announced CWS systems target 2028

    emerging
    Private · NLExposure
  12. M1 and M1076 analog processors compute against flash-resident weights for compact inference workloads

    niche
    Private · USExposure
  13. EN100 analog in-memory accelerator targets low-power client and edge AI with more than 200 TOPS

    emerging
    Private · USExposure
  14. Planned MX4 connects compute tiles directly to 3D memory; 2026 test silicon precedes 2027 sampling and 2028 production

    emerging
    Private · USExposure
  15. HC1 unifies model storage and computation on-chip, avoiding external HBM while retaining LoRA-based fine-tuning

    emerging
    Private · CAExposure
  16. Z1 stores and processes probabilistic states in place; taped-out thermodynamic silicon targets 2027 early access

    emerging
    Private · USExposure

Sources: d-matrix.ai, qualcomm.com, preferred.jp, fractile.ai, nvidia.com, cerebras.ai

Spatial & Specialized Compute

Wafer-scale, spatial, model-specific, photonic and thermodynamic processors change how AI arithmetic and data movement execute. Compiler maturity and demonstrated workload support separate working niche products from experimental silicon.

Heat (est.)
  1. WSE-3 retains an entire wafer as a connected compute fabric, removing conventional inter-GPU partition boundaries

    leader
    CBRS · USExposure
  2. Groq-derived LP30 adds deterministic, statically scheduled dataflow for token generation alongside general Rubin GPUs

    major
    NVDA · USExposure
  3. Blackhole Tensix tiles combine tensor arithmetic with RISC-V control and explicit software-managed data movement

    challenger
    Private · USExposure
  4. SN50 reconfigurable dataflow units map model operations onto a spatial fabric for high-throughput decode

    major
    Private · USExposure
  5. Licensed LPU design statically schedules a programmable assembly line across compute and inter-chip data movement

    niche
    Private · USExposure
  6. Maverick-2 dynamically maps workloads onto configurable compute; Sandia's Spectra validates its HPC deployment

    niche
    Private · ILExposure
  7. MN-Core 2 uses a compiler-controlled hierarchy for dense matrix workloads instead of a general GPU execution model

    niche
    Private · JPExposure
  8. Graphcore's Bow IPU implements distributed SRAM and independently scheduled tiles; this describes its documented product lineage

    niche
    9984.T · TSE · JPExposure
  9. The MatX One specializes its systolic compute and memory hierarchy for large language models rather than graphics

    emerging
    Private · USExposure
  10. Sohu's transformer specialization has evolved into separate prefill compute and decode-memory designs in early deployments

    emerging
    Private · USExposure
  11. DX-1 decode accelerator uses the X-1 optical-linked architecture; initial customer deliveries target 2H 2027

    emerging
    Private · GBExposure
  12. BER10 introduces a European RISC-V accelerator family for AI and HPC, with a full software stack under development

    emerging
    Private · ESExposure
  13. CN101 experimental thermodynamic ASIC demonstrates on-chip generative workloads; commercial successors remain a roadmap

    emerging
    Private · USExposure
  14. HC1 fixes a model into dedicated silicon, trading general programmability for specialized low-latency inference

    emerging
    Private · CAExposure
  15. X0 demonstrates probabilistic circuits; Z1 thermodynamic sampling silicon has taped out for 2027 early access

    emerging
    Private · USExposure
  16. PACE 2 combines a 128x128 optical matrix multiplier with electronic tensor compute for programmable AI/HPC arithmetic

    niche
    1879.HK · HKEX · CNExposure
  17. Gen 2 Native Processing Unit performs nonlinear photonic arithmetic; LRZ evaluates it and IONOS is a commercial customer

    niche
    Private · DEExposure
  18. TX81 maps AI graphs onto reconfigurable compute and scales modules through direct Mesh/Torus connections

    challenger
    Private · CNExposure
  19. Antoum sparse processing units skip zero-valued work; S4 hardware supports sparse neural inference

    niche
    Private · CNExposure
  20. Envise demonstrates photonic matrix arithmetic on ResNet and BERT; current commercial focus is Passage interconnect

    emerging
    Private · USExposure
  21. Open ET-SoC-1 RTL preserves Esperanto's manycore tensor architecture; new MRAM-based silicon remains in development

    emerging
    Private · USExposure

Sources: lightmatter.co, pubmed.ncbi.nlm.nih.gov, aifoundry.org, tsingmicro.com, cerebras.ai, taalas.com

Glossary

GPU
A parallel processor originally designed for graphics and now widely used for AI matrix operations.
Prefill
The inference stage that processes a prompt and builds the context needed to generate an answer.
Decode
The inference stage that produces output tokens, often constrained by moving model weights from memory.
Dataflow
An execution model that maps operations and their data movement directly onto a connected compute fabric.
In-memory compute
Arithmetic performed inside or next to memory arrays to reduce data movement and its energy cost.

Heat, spend, exposure and shares are editorial estimates. How to read the map ↗