Who Supplies Google's AI Infrastructure? A TPU Vendor Map

Who Supplies Google's AI Infrastructure? A TPU Vendor Map

A five-layer vendor-by-vendor map of Google's TPU program, covering 15+ suppliers from silicon design to final assembly and analyzing which positions are structurally defensible versus commoditized.

Google’s TPU program is often described as an in-house chip story. For supply chain analysis, that framing is too clean. The current TPU buildout sits on a five-layer vendor stack: silicon design partners, wafer fabrication and advanced packaging, HBM memory, optical interconnects, and final assembly. The economics change sharply from one layer to the next. Some vendors are tied to scarce capacity or architecture lock-in; others are doing necessary work that can be competed away once Google qualifies alternatives.

That distinction matters because Google’s AI infrastructure plan is no longer a lab-scale sourcing exercise. Alphabet’s 2026 capex requirement is estimated at $175 billion to $185 billion, above operating cash flow, and the company conducted an $80 billion equity raise, its first since the 2004 IPO, to help fund the buildout. Google also needs to double AI serving capacity every six months, which turns supplier position into a delivery-risk question rather than a branding question.[1]

Five-layer schematic of Google's TPU supply chain with TSMC CoWoS shown as the central bottleneck

The Five-Layer TPU Vendor Map

A useful Google AI supply chain stock analysis should start with role quality, not ticker familiarity. The same TPU shipment can pull through a high-margin optical switch, a constrained advanced-packaging slot, a large but more substitutable memory allocation, and commodity assembly hours. Data Gravity’s vendor map is the best organizing surface because it shows the stack as a set of dependencies rather than a flat list of suppliers.[2]

LayerVendors and rolesWhat makes the position defensible or exposed
Silicon design and platform functionsBroadcom for Sunfish training chip; MediaTek for Zebrafish inference chip; Intel for Xeon and IPU periphery; Marvell potentially for a memory processing unit if negotiations concludeDefensibility depends on design-in depth, contract duration, and how hard Google makes it to swap a function once qualified
Fabrication and advanced packagingTSMC wafer fabrication and CoWoS advanced packagingThe central bottleneck; shipment estimates widen or narrow around CoWoS allocation
HBM memorySamsung majority HBM3E supply for Ironwood; SK Hynix remainder; Micron limited as of mid-2026Large revenue opportunity, but memory capacity is more exposed to qualification shifts and broader commoditization than architecture-locked components
Optical interconnectsLumentum Apollo optical circuit switch position, with related optical and networking suppliers in the interconnect layerThe clearest premium-margin position where architecture lock-in and limited substitution appear strongest
Final assembly and systems integrationODM and assembly partners including server, rack, and integration suppliersOperationally important, but generally weaker pricing power because Google can qualify multiple manufacturers over time

The table is deliberately uneven. A supplier that receives a purchase order does not automatically receive economic leverage. In this chain, leverage concentrates where Google either cannot easily re-route around a bottleneck or has already committed the architecture around a specialized component.

Silicon Partners: Google Is Designing, But It Is Not Alone

Google’s internal TPU design capability is real, but the external partner map shows a more distributed chip strategy. TNW reports that Broadcom builds the Sunfish training chip, MediaTek builds the Zebrafish inference chip, Intel supplies Xeon and IPU functions around the TPU platform, and Marvell is in talks for a memory processing unit. The Marvell point should stay conditional: as of mid-2026, the brief supports negotiations, not a confirmed signed deal.[3]

Broadcom has the strongest disclosed commercial anchor in this layer. Its long-term agreement with Google runs through 2031, and TNW reports a $73 billion AI backlog connected to its AI business. That does not make Broadcom interchangeable with every other chip-services supplier in the map. A multiyear design and supply relationship around training silicon gives it a different kind of visibility than vendors that compete for assembly or memory allocation cycle by cycle.[3]

MediaTek’s role is different. TNW describes Zebrafish as an inference chip that is 20% to 30% cheaper. That cost-down function matters because inference volume is where serving-capacity pressure becomes continuous procurement pressure. If Google has to double AI serving capacity every six months, a lower-cost inference path is not a side experiment; it is one way to keep the serving footprint from turning into a runaway bill.[1][3]

Intel’s position is more peripheral to the TPU itself, but still relevant to the data center platform. Xeon and IPU functions help surround the accelerator layer with compute, control, and infrastructure handling. Those roles can be large and necessary without having the same lock-in profile as a training-chip partner or a sole-source optical architecture supplier.

TSMC CoWoS Is the Constraint That Reprices the Whole Map

If there is one layer that should be circled in red, it is TSMC’s CoWoS advanced packaging. Wafer fabrication matters, but the shipment debate around Google TPUs is specifically tied to advanced-packaging allocation. Data Gravity’s map and shipment discussion point to CoWoS as the binding constraint across the chain, not chip design demand.[2]

The shipment range shows the constraint more clearly than a single forecast would. Google expects 4.3 million TPU shipments in 2026, while DigiTimes estimates 3.3 million and BofA estimates 4.6 million. The gap is attributed to CoWoS allocation assumptions rather than lack of end demand.[2][3]

TSMC’s CoWoS capacity target is 140,000 wafers per month by the end of 2026, roughly four times early-2024 levels. That expansion is large, but the relevant procurement question is not whether capacity grows. It is who receives the next increment when multiple AI programs want the same packaging steps at the same time.[1][2]

This is why CoWoS sits across the entire vendor map rather than inside one box. Broadcom and MediaTek cannot turn design wins into volume without packaging. HBM suppliers cannot ship into a TPU system if the accelerator package is not scheduled. Optical suppliers can be designed in and still wait behind accelerator availability. Assembly partners can add shifts and racks, but they cannot assemble around missing packaged silicon.

The same pattern appears more broadly in AI hardware markets: physical bottlenecks can move revenue timing and market risk even when demand is intact. That is the useful connection to AI chip supply chain bottlenecks. For Google’s TPU program, the bottleneck is not an abstract shortage story. It is the packaging allocation that separates the low end and high end of 2026 shipment estimates.

HBM: Large Allocation, Less Structural Control

HBM deserves attention, but not the same kind of attention as CoWoS or optical switching. For Ironwood, Samsung holds about 60% of HBM3E supply, SK Hynix supplies the remainder, and Micron is not a meaningful factor as of mid-2026, according to Data Gravity’s summary of the HBM split.[2]

That split tells a procurement story. Google is not relying on a single memory vendor, but neither is the allocation evenly distributed. Samsung’s majority position may bring volume, schedule influence, and qualification value. SK Hynix remains relevant because a second source reduces single-supplier exposure. Micron’s limited role means it should not be treated as a central Google TPU beneficiary on the evidence available here.

The economic defensibility is more mixed. HBM is hard to make and essential to accelerator performance, but memory suppliers still operate in a market where qualification, capacity timing, and broader pricing cycles matter. A supplier can have a large allocation and still face more substitution risk than a component locked into the physical architecture of Google’s switching fabric.

Optical Interconnects: Where Lock-In Looks Most Visible

Lumentum’s Apollo optical circuit switch position is the cleanest example of premium economics in the current map. Data Gravity identifies Lumentum’s Apollo OCS as a sole-source-style position and estimates average selling price at roughly $150,000 per unit.[2]

That is a different kind of supplier role from selling a replaceable module into a broad rack bill of materials. Optical circuit switching becomes part of how the cluster is architected. Once the interconnect design, control assumptions, and qualification work are built around a specific switching approach, the switching supplier is not immune to second-sourcing forever, but it is better insulated than an assembly vendor bidding on the next tranche of systems.

The margin signal matters because it reflects more than volume. A $150,000 estimated unit ASP in a sole-source-style role points to a supplier capturing value from scarcity and integration complexity, not just from the number of data centers Google builds. That is the kind of position procurement teams watch carefully: it can become a cost driver, a schedule dependency, and a negotiation problem at the same time.[2]

Gradient spectrum showing AI supply chain vendor defensibility from sole-source lock-in to commoditized assembly and memory supply

Assembly ODMs: Necessary Work, Weaker Pricing Power

Final assembly is where many AI infrastructure maps become misleading. Server, rack, and systems-integration partners are essential to delivery. They absorb schedule pressure, labor planning, factory loading, and late-stage configuration changes. But essential does not automatically mean defensible.

ODMs generally sit in a more substitutable part of the chain than CoWoS packaging or architecture-specific optical switching. Google can qualify multiple manufacturers, divide work by region or product generation, and use competitive pressure once designs stabilize. The assembly layer can still benefit from enormous TPU volume, but its bargaining power is usually capped by the buyer’s ability to move work after qualification.

This is where a procurement view diverges from a headline supplier list. A company may appear in the TPU supply chain and still have commodity-margin exposure. That is not a criticism of execution. It is a recognition that the margin pool is not distributed evenly across the bill of materials.

Shipment Forecasts Are Really Packaging Scenarios

The 2026 TPU shipment estimates should be read as scenarios around capacity allocation. A 3.3 million estimate and a 4.6 million estimate imply very different pull-through for memory, optical components, boards, racks, and assembly. But the brief attributes the difference to CoWoS allocation, not to a dispute over whether Google wants the chips.[2][3]

2026 TPU shipment viewEstimateInterpretation
Lower estimate3.3 millionA tighter CoWoS allocation case; downstream suppliers receive less volume despite demand
Google expectation4.3 millionThe central planning signal reported in the brief
Higher estimate4.6 millionA more favorable packaging-allocation case, not necessarily a stronger demand case

This distinction prevents a common mistake in Google AI supply chain stock analysis: treating every higher shipment number as proof of stronger end-market demand and every lower number as demand weakness. In this case, the more precise reading is operational. The spread measures how much packaged accelerator supply Google can secure.

The Financial Signal Is About Supplier Longevity

Alphabet’s financing and capex signals belong in this map because they affect supplier planning. An $80 billion equity raise and a 2026 capex requirement estimated at $175 billion to $185 billion indicate that AI infrastructure demand is being funded at a scale that can justify long-term agreements, capacity reservations, and design-specific component commitments.[1]

They do not settle the sustainability debate. Large AI infrastructure spending can still face questions about returns, utilization, and timing. For the vendor map, the more practical point is that suppliers are not planning around a small experimental run. They are deciding whether to allocate scarce capacity, accept customer-specific design work, or build inventory buffers around a buyer whose compute requirement is expanding aggressively.

That is also why hyperscaler infrastructure risk cannot be reduced to one customer or one stock reaction. The same issue appears in adjacent cases such as CoreWeave’s infrastructure supply chain risk: when compute demand is real but physical inputs are constrained, the timing of capacity becomes part of the business model.

A Structural Heatmap of Defensibility

The vendor map is best read as a heatmap, not a winner list. The strongest positions combine scarcity, design-in depth, and limited substitution. The weakest positions can still grow with Google’s volume, but they carry more competitive pressure.

Defensibility zoneVendor positionsReason
HighestTSMC CoWoS; Lumentum Apollo OCSCoWoS controls the binding packaging constraint; Apollo OCS shows premium, sole-source-style architecture lock-in
High to moderateBroadcom training-chip relationshipLong-term agreement through 2031 and deep design role create visibility, though still tied to Google’s platform decisions
ModerateMediaTek inference chip; Intel periphery; possible Marvell memory processing unit if confirmedImportant functional roles, with defensibility depending on qualification depth and whether Google preserves multiple options
Moderate to lowerSamsung and SK Hynix HBM supplyEssential and capacity-sensitive, but more exposed to memory-cycle economics and multi-sourcing
LowerAssembly ODMs and rack integration partnersOperationally critical but more substitutable after qualification, with weaker structural pricing power

This is structural analysis, not investment advice. A defensible supply chain position can still be a poor investment at the wrong price, and a commodity-margin role can still generate revenue growth during a capacity boom. The useful judgment is narrower: where would Google have the hardest time replacing a supplier without schedule risk, requalification work, or architecture disruption?

On that question, CoWoS and optical switching sit at the hot end of the map. Broadcom’s training-chip role has unusual contractual depth. MediaTek, Intel, and a possible Marvell role show how Google spreads chip-function dependencies without handing every function the same economic quality. HBM suppliers carry meaningful allocation value but less architectural control. Assembly partners turn the program into physical infrastructure, but they are more exposed to second-sourcing and margin pressure.

For readers looking at procurement strategy rather than the vendor atlas itself, the companion piece on lessons from Google’s AI chip supply chain strategy covers the multi-sourcing and bottleneck-management playbook. This map answers the prior question: which suppliers sit where, and which positions are hardest for Google to route around.

References

  1. Alphabet resets the bar for AI infrastructure spending, CNBC
  2. Google's TPU Supply Chain, Data Gravity
  3. Google assembles four-partner chip supply chain, TNW

Comments

Join the discussion with an anonymous comment.

Loading comments...
Blogarama - Blog Directory