White and blue abstract network on server room data center 3D rendering
Ecosystem Home Page Technical/How-To Technology Partners VCF Performance VCF Private AI Services

Deploy VCF on Supermicro HGX Servers with NVIDIA GPUs for AI Workloads

By Yuankun Fu, Agustin Malanco, Ramesh Radhakrishnan, Vrushal Dongre
Supermicro: Sil Angelescu, Duyen Truong

Executive Summary

Specialized AI computing infrastructure—combining high-performance hardware, optimized software, and data pipelines—is essential to continuously ingest raw data and process it into artificial intelligence models and inference at scale.

This reference architecture shows how to deploy a pool of GPU-accelerated infrastructure, wired with a lossless high-speed fabric and operated as a single multi-tenant system. This ensures that training, fine-tuning, and inference workloads — and increasingly, autonomous agentic AI pipelines — can be provisioned, scheduled, and governed with the same rigor as any other enterprise cloud service. Furthermore, this AI environment lives inside the enterprise’s own security boundary, keeping model weights, training data, and inference traffic on infrastructure the organization already controls.

Building one means the underlying infrastructure has to deliver three things at once: compute density, a lossless fabric that keeps distributed training from stalling, and multi-tenant cloud orchestration that a platform team can actually operate. Falling short on any one of the three turns expensive GPU resources into either an underutilized GPU island or an unmanageable one-off.

This blog describes an HGX Reference Architecture for deploying Supermicro SuperCloud hardware management with VMware Cloud Foundation (VCF) 9.1 and vSphere Kubernetes Service (VKS) to deliver exactly that, on certified Supermicro servers using NVIDIA HGX GPUs. The compute layer consists of two Supermicro servers equipped with NVIDIA HGX B200 GPUs— one direct-to-chip liquid-cooled, one air-cooled — allowing the same workload to be compared across both thermal designs. The NVIDIA GPU backend fabric utilizes a RoCEv2 Ethernet network built on a Supermicro switch using Broadcom Tomahawk 4 chip, connecting each node by eight NVIDIA ConnectX-7 adapters running at 400G over OSFP112 passive copper, for 3.2 Tb/s of aggregate GPU fabric bandwidth per node. A four-host AMD EPYC management domain runs the VCF control plane on VMware vSAN ESA.

The result is cloud-like operational agility — self-service Kubernetes clusters, policy-driven storage, automated network isolation — without giving up the near-bare-metal GPU performance that AI training and inference demand.

Reference Architecture Overview

The environment pairs a fully operational VCF 9.1 management domain with a high-performance GPU workload cluster, sitting on a Supermicro SuperCloud hardware-management layer underneath. Each layer has a clear job.

AI solutions. Enterprise AI applications—such as large language model services and autonomous agentic workflows—sit at the top of the stack. These solutions are powered by high-performance inference runtimes and caching engines, including vLLM, NVIDIA NIM, and LMCache. They consume the infrastructure layers below through standard Kubernetes and compatible APIs, ensuring application teams never interact with the underlying physical infrastructure directly.

GPUaaS and workload orchestration. GPU capacity is published as an on-demand service catalog. This brief used NVIDIA Run:ai at this layer, covering interactive workspaces, training jobs, and inference services — see Deploying NVIDIA Run:ai on VMware Cloud Foundation. The SuperCloud Developer Console (SDX) is an additional option offering AI workload execution and acceleration platform that enables efficient GPU sharing across multi-tenant environments,

Kubernetes platform and GPU runtime. VKS is the Kubernetes runtime that ships natively with VCF. Clusters are provisioned directly from vSphere through the Supervisor, a Kubernetes control plane that runs on the ESXi hypervisor layer and serves as the management plane for all VKS workload clusters, vSphere Namespaces, and Supervisor Services. There is no separate Kubernetes distribution to install, no external control plane to operate, and no extra licensing. Worker nodes are vSphere VMs, so GPUs attached with NVIDIA vGPU or with Dynamic DirectPath I/O are consumed by Kubernetes through the same vSphere scheduler, with the vMotion and DRS behavior that operators already understand. The Kubernetes layer is not a new silo — it is part of the same cluster the platform team already manages.

VMware Cloud Foundation (VCF) 9.1, the leading private cloud platform.  VCF provides a flexible and simplified private cloud platform with public cloud extensibility that integrates leading products including vSphere (compute), vSAN (storage), NSX (networking), vSphere Kubernetes Service (VKS), VCF Operations, VCF Automation and Private AI services into a single solution. VMware Cloud Foundation is a platform that enables you to modernize infrastructure, accelerate developer productivity and provide greater resilience and security. 

Rack management. Beneath VCF sits a hardware-level management plane from the Supermicro SuperCloud Software Suite: SuperCloud Director (SCD) and SuperCloud Automation Center (SCAC) provide hardware-based multi-tenant management of the physical servers and network, and SuperCloud Composer (SCC) handles rack- and datacenter-scale power, cooling, and firmware. Because this layer operates below and independently of VCF, the same racks can carry VCF-orchestrated GPUaaS tenants alongside directly rented bare-metal or rack tenants, with SuperCloud Developer Console (SDX) as the self-service front end across both models.

Physical layer. GPU and CPU servers, NVIDIA GPUs, and a high-speed RoCEv2 or InfiniBand fabric. Broadcom’s VM DirectPath I/O for General GPU compatibility list certifies a broad set of Supermicro servers for NVIDIA GPU passthrough. The table below is limited to Supermicro platforms certified with Hopper-generation or later GPUs (Hopper, Ada Lovelace, Blackwell) — Ampere-class (A100/A40/A30) and Tesla (T4) platforms are excluded as outside the scope of this blog:

Certified Compute & GPU NodesForm factorCertified GPUs (Hopper+)
Supermicro AS-4126GS-NBR-LCC4U, direct-to-chip liquid cooled, 8-GPU HGXNVIDIA B200 (HGX B200 NVL)
Supermicro SYS-A22GA-NBRT10U, air cooled, 8-GPU HGXNVIDIA B200 (HGX B200 NVL)
Supermicro AS-5126GS-TNRT25U, air cooled, PCIe GPUNVIDIA RTX PRO 6000 Blackwell Server Edition
Supermicro SYS-521GE-TNRT5U, air cooled, PCIe GPUNVIDIA H100, H100 NVL, L40S
Supermicro SYS-221H-TNR2U, air cooled, PCIe GPUNVIDIA H100, L40S
Supermicro SYS-821GE-TNHR8U, air cooled, PCIe GPUNVIDIA H200 SXM5

This is a representative subset, not the complete list — the full Broadcom compatibility list also includes other Supermicro Hopper/Ada/Blackwell platforms (for example AS-4125GS-TNRT2, SYS-221GE-NR, SYS-421GE-TNRT) beyond the servers and GPUs used in this brief (HGX B200 NVL, HGX B300 NVL8, RTX PRO 6000 Blackwell Server Edition, L40S). See the Broadcom Compatibility Guide — filter the “VM Direct Path IO for General GPU” program by partner Supermicro Computer, Inc. and by GPU partner “NVIDIA” — to confirm certification status for your exact configuration. The deployment in this brief uses the liquid-cooled AS-4126GS-NBR-LCC and its air-cooled counterpart, AS-A4126GS-TNBR; Section 2 details the exact configuration.

Physical Infrastructure in Deployment

This brief validates two configurations from the certified list above. Apart from cooling, the two GPU nodes are identically configured. The VCF management domain runs on separate, AMD EPYC-based Supermicro servers, so a management operation never competes with a training job for compute or PCIe bandwidth. Two physically separate networks connect everything: a lossless RoCEv2 fabric for GPU-to-GPU traffic, and a standard Ethernet network for ESXi management, vMotion, vSAN, and NSX overlay.

Physical InfrastructureConfiguration
GPU Node 1Supermicro AS-4126GS-NBR-LCC (NVDA-B200-n01), liquid-cooled — 8× NVIDIA B200, 4× Intel X710 mgmt NICs, 8× Mellanox ConnectX-7 (CX-7) RDMA NICs
GPU Node 2Supermicro AS-A4126GS-TNBR (NVDA-B200-n02), air-cooled — same GPU, NIC, and CX-7 configuration as Node 1
Management servers4x Supermicro servers with AMD EPYC CPUs, running the VCF control plane (vCenter, NSX manager, SDDC Manager) on vSAN ESA
RDMA backend networkSupermicro switch SSE-T8032S with Broadcom Tomahawk 4 switch connecting the CX-7 adapters on both GPU nodes over passive DAC cabling, running RoCEv2 with PFC and ECN for a lossless fabric
Management networkSupermicro SSE-F3548S(R) Standard Ethernet switching carrying ESXi management, vMotion, vSAN, and NSX overlay traffic — physically isolated from the RDMA backend network

Single-Node MLPerf Inference Results

Before measuring anything across the fabric, it is worth establishing what a single virtualized node can do. We reuse previously published results from The New Paradigm: MLPerf Inference 5.1 Confirms VCF is the Future of AI/ML Performance, measured with DirectPath I/O passthrough of 8× NVIDIA Blackwell B200 GPUs on the Supermicro liquid-cooled server. The single-node result matters as a baseline: virtualization overhead at the GPU level is already accounted for.

Serving GLM-5.2 with NVIDIA NIM on a Single Supermicro HGX B200 Liquid Server

To validate the inference capabilities of this architecture, we tested the GLM-5.2 model using NVIDIA NIM (sglang backend) on a single AS-4126GS-NBR-LCC 8× NVIDIA B200 HGX server. Leveraging our previously established SPOC-style capacity discovery and NSGA-II tuning methodology, we drove the evaluation using a multi-turn coding-agent dataset by NVIDIA AIPerf benchmark tool.

Our preliminary findings demonstrate highly stable and scalable performance (SLA: p95 TTFT < 10000 ms, p95 ITL < 200 ms):

  • Heavyweight Agentic Workload: The evaluation utilized a realistic multi-turn coding-agent dataset, mirroring production distributions rather than simplistic synthetic benchmarks. Each concurrent session is highly stateful, comprising an average of 20.7 continuous conversational turns. To rigorously test the system’s limits, the workload was categorized into distinct depth tiers, demonstrating how massive the per-turn token lengths actually are:
    • T0 (Baseline): 200K context cap, averaging 83K input tokens (ISL) and 310 output tokens (OSL) per turn.
    • T1 (Medium): 262K context cap, averaging 132K input tokens (ISL) and 287 output tokens (OSL) per turn.
    • T2 (Deep): 524K context cap, averaging 229K input tokens (ISL) and 256 output tokens (OSL) per turn.
  • 1M Context Validation: The server successfully handled up to 1,033,930 tokens, proving the advertised 1M context window is real and gracefully enforced without crashing.
  • Maximum Feasible Capacity: Through rigorous capacity discovery at the T0 baseline, the system sustained strict SLAs up to a peak concurrency of 26. This proves exceptional compute and memory headroom for deep-context agentic workloads on a single node.
  • Parameter Optimization (Production Load): While the node can push 26 concurrent heavyweight sessions at T0 depth, enterprise deployments require safety margins. We locked the load at 18 concurrency (roughly 70% of peak capacity) to conduct exhaustive NSGA-II parameter tuning specifically optimized for the T0 workload.
  • Optimal Throughput: As illustrated in the Pareto front chart, all optimized configurations at this sustained production load easily cleared the 200 ms ITL SLA. During the parameter search phase executing the T0 profile, the winning configuration achieved a peak throughput of 356.3 output tokens per second, with a TTFT p95 of 3483 ms and an ITL p95 of 46.3 ms.
  • Sustained Stability (Golden Run): When subjected to a rigorous 40-minute soak test running the T0 baseline profile, this winning configuration demonstrated exceptional sustained stability. It delivered 338.5 tokens per second with an even tighter TTFT p95 of 2887 ms and an ITL p95 of 41.9 ms, maintaining a 99.89% success rate with zero GPU Xid faults.
  • Speculative Decoding: Enabling EAGLE speculative decoding yielded an additional 8% throughput boost, though with a slight trade-off in TTFT tail latency.

A comprehensive deep dive into these tuning results and the cross-node performance will be published in our upcoming full technical reference architecture paper.

Conclusion and Next Steps

This architecture demonstrates that an AI infrastructure does not require a bespoke, hand-built GPU silo. The same VCF instance that runs enterprise workloads can provision GPU-accelerated Kubernetes clusters, attach a lossless RoCEv2 fabric, and serve production LLM inference — managed with the tools, skills, and operational model a platform team already has.  We are targeting to publish the full solution architecture technical paper soon.

Appendix

Table 1. Supermicro rack management software

Rack Management Software (SKU)Description
Supermicro SFT-CLD-DIR-NODE-[1|3|5]YSuperCloud Director (SCD), 1x per managed node, [1|3|5] Year
Supermicro SFT-AICON-NODE-[1|3|5]YSuperCloud Dev Console (SDX), 1x per managed node, [1|3|5] Year
Supermicro SFT-AUTOMATION-NODESuperCloud Automation Center (SCAC),1x per managed node, Perpetual
Supermicro SFT-DAILY-PRO-SERVE-ROWSuperCloud GPU Operations Center (GOC) Services, 5x per managed node, On-Site / Remote Software Professional Service.
Supermicro SFT-SDLCR-NODE

(SDLCR: Supermicro BMC Direct Liquid Cooling Release)
SuperCloud Composer (SCC) – Liquid Cooling Feature, 1x per liquid-cooled node, Perpetual for Supermicro-BMC servers to enable liquid cooling infrastructure monitoring (e.g., leak detection policy, power strike, etc).
Note: Use SFT-ODLCR-NODE for Open BMC liquid-cooled servers.
Supermicro SFT-ENTMGMT-7YSuperCloud Composer (SCC) – Air Cooling / Enterprise Management, 1x per air-cooled managed node and 1x per management node, Enterprise Management and Workflows for Thermal, Power and Security, 7-Year Term (Server Lifetime).

Note: All management software components listed above, with the exception of the SuperCloud Dev Console (SDX), utilize the out-of-band management network.

Acknowledgments
The authors thank Roger Fortier and Shobhit Bhutani from Broadcom’s VMware Cloud Foundation division for reviewing and improving the paper.


Discover more from VMware Cloud Foundation (VCF) Blog

Subscribe to get the latest posts sent to your email.