Thursday, July 23, 2026

AI/HPC HW: AMD vs NVIDIA

Interesting a subtle "shift" in talking about
"High Performance Computing" and "Accelerated Computing"
instead of AI-centric GPU and TPU and NPU

Technically those integrated rack-size systems are more efficient and relatively easier to deploy
than a large number of relatively small processors, so that is one driver of modern Moore's law.

But the real reason may be a need to diversify out of AI/LLM-only tasks,
since profits from them are not materializing (yet), not even close, 
and the "industry" is likely looking for other use-cases to make new "promises"
and to keep justifying huge investments.

If nothing else, having some HW competition is good.
But the "game" they play with those mega-investments is dangerous.


 AMD just landed its biggest AI deal yet - TheStreet

  • Major Microsoft Partnership: Microsoft agreed to deploy AMD’s Helios rack-scale AI systems across its Azure data centers for AI inference, joining Meta, OpenAI, and Oracle as early platform adopters.

  • Microsoft’s Three-Vendor Strategy: Microsoft is diversifying its AI hardware stack across three suppliers: Nvidia GPUs, its own custom Maia chips, and AMD’s Helios platform.

  • Full-Stack Architecture: The Helios system integrates AMD’s Instinct MI455X GPUs, EPYC Venice CPUs, Pensando networking hardware, and ROCm software into a unified rack setup.

  • Value vs. Sticker Price: Estimated at $5.0M–$5.5M per rack (compared to $3.5M–$4.0M for Nvidia’s Vera Rubin system), AMD is pitching Helios based on cost per completed AI task rather than initial hardware cost.

  • Market Reaction: AMD shares jumped up to 5%, with Wall Street analysts raising price targets (e.g., UBS to $700, Rosenblatt to $655). Analysts project AMD’s data center revenue to reach ~$31.2 billion this year.



This expansion strengthens Microsoft's heterogeneous cloud strategy, providing developers and enterprises with greater choice, compute efficiency, and open-architecture scalability beyond traditional single-vendor GPU deployments.




  • Massive Capital Expenditures vs. Revenue Gap: AI companies are spending tens of billions on infrastructure, high-end compute, and data center buildouts, but software revenues are failing to cover these astronomical operational costs.

  • Off-Balance-Sheet Debt (SPVs): To avoid diluting equity or signaling financial strain, tech firms and AI infrastructure players rely on Special Purpose Vehicles (SPVs) and complex private-credit arrangements. This structure keeps massive liabilities hidden off standard corporate balance sheets.

  • Circular Financing Risks: Tech giants fund special entities that purchase hardware and compute, creating a web of interlocked counterparty risk reminiscent of 2008-style financial instruments.

  • The "Brokenomics" of Scale: The core argument asserts that current AI business models rely on cheap capital to subsidize compute; without continuous bailouts or lower infrastructure costs, the current debt structure is unsustainable long-term.


the debt that is "off balance sheets" is now higher than one "on balance sheets".
this can have very dangerous effects if demand is not as strong as expected.

Some companies with most dept, like Oracle, may even not be able to survive downturn without major restructuring.



  • Pushing Back on "AI Doomers":

    • Huang rejects predictions of catastrophic, mass job loss (such as claims that AI will eliminate half of all jobs) as "ridiculous" and overly sensationalized "science fiction."

    • He argues that automating specific tasks does not eliminate entire roles. Instead, he believes AI will increase overall human productivity and lead to long-term job creation rather than mass unemployment.

    • He warns that excessive fearmongering creates an unnecessary backlash and risks deterring talent from entering critical fields like software engineering.

  • U.S.–China AI Dynamics & Cooperation:

    • Allowing Chinese AI Models: Huang advocates that American companies should "absolutely" be permitted to use open Chinese AI models, dismissing fears about inherent hidden "backdoors" as misconceptions, provided proper safety guardrails are applied.

    • Open vs. Closed Models: He stresses that the global ecosystem requires both open-source and proprietary models. Over-reliance on single closed ecosystems creates single points of failure.

    • Export Restrictions & Competition: Huang points out that restrictive trade policies have not fully stopped Chinese technological advancement (e.g., through hardware stacking and energy efficiency) and urges Washington to engage in research-level dialogue with China to establish global AI safety norms.