5 segments · 52 companies · as of Sep 18, 2026
GPUs & AI Accelerators
Merchant processors for training models and serving their outputs
Where it sits
Fed by: Chip ManufacturingFeeds: Servers & Systems
Merchant AI processors execute the matrix and attention operations behind training and inference. GPUs dominate flexible deployment, while regional suppliers and specialized architectures compete on software access, memory use, response time and power consumption.
Why now
Rubin and MI455X move the GPU roadmap forward while Nvidia commercializes Groq-derived LPUs. Inference specialists raise new capital, and Enflame joins the wave of listed Chinese accelerators. Announced capacity and preproduction designs are not equivalent to installed supply.
Merchant Training Accelerators
Programmable GPUs and tensor processors sold to AI developers for training and mixed training-inference workloads. Software compatibility, usable memory and deployment scale determine adoption.
Blackwell Ultra and Rubin GPUs supply CUDA training; Vera Rubin NVL72 is the current Rubin platform
leaderMI350-series GPUs and July-launched MI455X serve training; Helios volume deployments are planned for 2H 2026
challengerAscend 910C runs Chinese training workloads; 950DT is scheduled for Q4 2026, after the inference-focused 950PR
majorHabana Gaudi 3 remains its merchant training ASIC; Crescent Island is an inference GPU, not its replacement
challengerWSE-3 wafer-scale processor powers CS-3 training without partitioning a model across conventional GPU dies
challengerSiyuan MLU cloud processors and NeuWare distributed-training libraries support Chinese model-training clusters
challengerDCU accelerator cards use DTK's HIP, math and collective libraries to run distributed training in Chinese clusters
challengerMTT S5000 PH100 GPUs train Qwen models through MUSA and FlagOS, with published multi-node training validation
challengerMXC600 training-inference chips power C600 modules with MXMACA software and scalable MetaXLink interconnect
challengerBiren166M/L GPU modules support distributed model training; 166L powers the deployed Guangyue 128-card supernode
challengerCloudBlazer DTU accelerators and TopsRider software implement distributed cloud training for domestic AI customers
challengerTianGai 150/300 GPU families supply programmable training compute; ZhiKai remains the distinct inference line
challengerT-Head's Zhenwu M890 supports FP32-to-FP4 training and inference and is supplied to external enterprise customers
challengerSN40L reconfigurable dataflow chips support model customization; SN50 shifts emphasis toward agentic inference
nicheBlackhole and Wormhole Tensix processors use programmable tensor tiles for open-stack AI compute
challengerMN-Core 2 chips are sold in developer systems and boards for AI training and scientific compute
nicheTX81 reconfigurable modules target model training and long-context inference with Mesh/Torus cluster links
emergingGraphcore's documented Bow-2000 and Poplar train neural models; continued volume availability is not independently confirmed
niche
Sources: reuters.com, mthreads.com, metax-tech.com, birentech.com, tsingmicro.com, amd.com
China AI Accelerators
Domestic AI processors compete for Chinese cloud and enterprise deployments as import rules and local procurement policies change. Compiler maturity and access to usable memory constrain substitution.
Ascend 950PR powers Atlas 350 inference; the 910C and planned 950DT address training and decode
leaderSiyuan MLU processors use NeuWare and MagicMind for domestic cloud AI; no unverified MLU690 volume claim is assumed
majorDCU products such as K100_AI and BW1000 use the DTK stack; Hygon remains independent after the terminated Sugon merger
majorPH100-based MTT S5000 supports FP8 through FP64 and MUSA software for training and inference
challengerC600 OAM GPUs use MXMACA software and MetaXLink; C500 remains a PCIe training-inference option
challengerBiren166M/L training-inference modules and 166C inference cards extend its BR100 GPU lineage
challengerTianGai cloud training GPUs and ZhiKai inference GPUs form its domestic software-compatible compute portfolio
challengerCloudBlazer DTU accelerators and TopsRider serve training and inference; Enflame completed its September 11 STAR listing
challengerKunlunxin sells AI accelerators externally; its announced M100 inference and M300 training roadmap targets 2026 and 2027
majorZhenwu M890 and SAIL support FP32-through-FP4 training and inference, with external enterprise shipments disclosed
challengerVA10 inference accelerators and VastStream software serve data-center neural and video workloads
emergingGoldwasser GPU+ chips and UL modules target cloud-edge inference with its own compiler and runtime
nicheSOPHON BM1684X tensor processors support inference cards and multi-chip video-analysis deployments
nichePACE 2 supplies domestic optical/electronic tensor computation; its Photowave interconnect business is outside this segment
nicheJM and Jinghong GPUs support DeepSeek models; Chengheng's newly launched CH37 extends compute to edge AI
challengerTX81 cloud modules and mass-produced TX5 edge chips use the company's reconfigurable processing architecture
challengerAntoum sparse-compute chips power S4/S10 inference cards; sparsity-dependent throughput is workload-specific
nicheLX PRO TrueGPU cards provide 24GB GDDR6 and OpenCL compute for AI PCs and workstations, alongside graphics
emergingFenghua GPUs combine graphics with OpenCL AI compute; the published Fenghua 1 platform supports neural frameworks
niche
Sources: jingjiamicro.com, tsingmicro.com, kisacoresearch.com, lisuantech.com, reuters.com, metax-tech.com
LLM Inference Processors
Inference chips execute model prompts and generate tokens for deployed services. Buyers trade model coverage, response latency and memory capacity against software migration and electricity costs.
Groq-derived LP30 LPUs entered full production in August 2026; Rubin GPUs handle complementary inference stages
leaderMI355X and MI455X target large-model serving; OpenAI and Meta each contracted up to 6GW across generations
challengerAI200 targets 2026 inference deployments under HUMAIN's 200MW plan; AI250 commercial sampling targets mid-2027
challengerCrescent Island Xe3P inference GPU targets 2H 2026 sampling with 160GB LPDDR5X in Intel's reference design
emergingWSE-3 powers low-latency token generation; OpenAI agreed to deploy 750MW of Cerebras inference capacity
challengerSN50 RDUs target token decode alongside Intel Xeon hosts; SoftBank Corp. is the initial deployment partner
challengerCorsair digital in-memory accelerators entered production in June 2026 for hyperscalers and frontier labs
challengerRNGD tensor-contraction processors run LG AI Research's EXAONE; LG CNS is extending enterprise deployment
challengerREBEL-Quad combines four AI chiplets and 144GB HBM3E for large-model prefill and decode
challengerBlackhole Tensix chips and TT-Metal software execute local and server LLM inference on programmable tiles
nicheAtlas FPGA inference is deployed at Oracle; the Asimov custom ASIC targets production in 2H 2027
emergingCompany-disclosed low-voltage prefill hardware and cluster-scale decode memory underpin early Jane Street deployments
emergingBertha LPU targets token decode with 192GB LPDDR5X; Forte FPGA chips underpin the earlier HX platform
emergingSpyre's 32 AI cores provide PCIe-attached LLM inference for IBM z17, generally available since October 2025
nicheRetains LPU inference technology licensed non-exclusively to Nvidia; independent GroqCloud continues operating
nicheHC1 hardwires Llama 3.1 8B into model-specific silicon; HC2 remains a roadmap and the AMD acquisition is pending
emerging
Sources: reuters.com, nvidianews.nvidia.com, qualcomm.com, intelcapital.com, furiosa.ai, rebellions.ai
Memory-Centric AI Compute
Architectures put model weights close to arithmetic, using SRAM, analog arrays or tightly coupled DRAM. They target memory-bound inference, but several designs remain preproduction or limited to edge models.
Groq 3 LP30 uses on-chip SRAM for low-latency decode, complementing Rubin GPU memory capacity
majorWSE-3 distributes 44GB of SRAM beside wafer-scale compute cores, reducing movement of activations and working data
majorCorsair uses digital in-memory compute and on-chip SRAM to reduce LLM decode data movement
leaderLicensed LPU architecture keeps model data in on-chip SRAM under compiler control; GroqCloud retains this compute lineage
nicheAI250's High Bandwidth Compute architecture targets near-memory inference; commercial sampling is planned for mid-2027
emergingMN-Core L1000 stacks DRAM above logic for generative inference; the company still describes it as under development
emergingInterleaved memory and compute targets frontier-model inference; first commercial chips are expected in 2027
emergingThe MatX One keeps weights in SRAM and KV caches in HBM to combine fast decoding with long-context support
emergingAsimov's memory-focused ASIC targets large-model inference; Atlas currently implements the approach on FPGAs
emergingJotunn8 inference processor targets the memory wall; July funding supports commercial deployment
emergingCRAFTWERK pairs programmable inference ASICs with processor-memory co-design; announced CWS systems target 2028
emergingM1 and M1076 analog processors compute against flash-resident weights for compact inference workloads
nicheEN100 analog in-memory accelerator targets low-power client and edge AI with more than 200 TOPS
emergingPlanned MX4 connects compute tiles directly to 3D memory; 2026 test silicon precedes 2027 sampling and 2028 production
emergingHC1 unifies model storage and computation on-chip, avoiding external HBM while retaining LoRA-based fine-tuning
emergingZ1 stores and processes probabilistic states in place; taped-out thermodynamic silicon targets 2027 early access
emerging
Sources: d-matrix.ai, qualcomm.com, preferred.jp, fractile.ai, nvidia.com, cerebras.ai
Spatial & Specialized Compute
Wafer-scale, spatial, model-specific, photonic and thermodynamic processors change how AI arithmetic and data movement execute. Compiler maturity and demonstrated workload support separate working niche products from experimental silicon.
WSE-3 retains an entire wafer as a connected compute fabric, removing conventional inter-GPU partition boundaries
leaderGroq-derived LP30 adds deterministic, statically scheduled dataflow for token generation alongside general Rubin GPUs
majorBlackhole Tensix tiles combine tensor arithmetic with RISC-V control and explicit software-managed data movement
challengerSN50 reconfigurable dataflow units map model operations onto a spatial fabric for high-throughput decode
majorLicensed LPU design statically schedules a programmable assembly line across compute and inter-chip data movement
nicheMaverick-2 dynamically maps workloads onto configurable compute; Sandia's Spectra validates its HPC deployment
nicheMN-Core 2 uses a compiler-controlled hierarchy for dense matrix workloads instead of a general GPU execution model
nicheGraphcore's Bow IPU implements distributed SRAM and independently scheduled tiles; this describes its documented product lineage
nicheThe MatX One specializes its systolic compute and memory hierarchy for large language models rather than graphics
emergingSohu's transformer specialization has evolved into separate prefill compute and decode-memory designs in early deployments
emergingDX-1 decode accelerator uses the X-1 optical-linked architecture; initial customer deliveries target 2H 2027
emergingBER10 introduces a European RISC-V accelerator family for AI and HPC, with a full software stack under development
emergingCN101 experimental thermodynamic ASIC demonstrates on-chip generative workloads; commercial successors remain a roadmap
emergingHC1 fixes a model into dedicated silicon, trading general programmability for specialized low-latency inference
emergingX0 demonstrates probabilistic circuits; Z1 thermodynamic sampling silicon has taped out for 2027 early access
emergingPACE 2 combines a 128x128 optical matrix multiplier with electronic tensor compute for programmable AI/HPC arithmetic
nicheGen 2 Native Processing Unit performs nonlinear photonic arithmetic; LRZ evaluates it and IONOS is a commercial customer
nicheTX81 maps AI graphs onto reconfigurable compute and scales modules through direct Mesh/Torus connections
challengerAntoum sparse processing units skip zero-valued work; S4 hardware supports sparse neural inference
nicheEnvise demonstrates photonic matrix arithmetic on ResNet and BERT; current commercial focus is Passage interconnect
emergingOpen ET-SoC-1 RTL preserves Esperanto's manycore tensor architecture; new MRAM-based silicon remains in development
emerging
Sources: lightmatter.co, pubmed.ncbi.nlm.nih.gov, aifoundry.org, tsingmicro.com, cerebras.ai, taalas.com
Glossary
- GPU
- A parallel processor originally designed for graphics and now widely used for AI matrix operations.
- Prefill
- The inference stage that processes a prompt and builds the context needed to generate an answer.
- Decode
- The inference stage that produces output tokens, often constrained by moving model weights from memory.
- Dataflow
- An execution model that maps operations and their data movement directly onto a connected compute fabric.
- In-memory compute
- Arithmetic performed inside or next to memory arrays to reduce data movement and its energy cost.
Heat, spend, exposure and shares are editorial estimates. How to read the map ↗
