Introduction
While artificial intelligence offers transformative potential, enterprise deployment remains constrained. Privacy, IP protection, and compliance present severe existential threats. Organizations must enforce granular access boundaries and guard proprietary data. The rapid shift toward Agentic AI dramatically amplifies these foundational friction points, introducing critical operational, cost, and governance challenges. Unlike traditional single-call inference models, autonomous agents operate in dynamic loops that drive unpredictable, non-linear token cost escalation and overwhelm legacy IT infrastructure.
New Challenges Emerge in the Agentic AI Era
In the Agentic AI era, enterprises are faced with additional challenges to work through.
- Operational complexity: Agentic workloads do not behave like traditional inference requests. Agents chain together dozens of model calls, tool invocations, and reasoning loops. Traditional infrastructure was simply not built for this kind of sustained, dynamic demand. The deployment challenges include:
- Day 0 deployment: Compute, Networking, Storage, Kubernetes, & networking infrastructure.
- Day 2 operations: Lifecycle management including patching, upgrades, model versioning, and framework migrations all become continuous, overlapping cycles rather than discrete projects.
- Token cost escalation: AI agents operate in iterative loops, hence token consumption scales non-linearly. In fact, token cost is projected to grow up to 24X by 2030. Context-bloating, multi-agent communication, error correction, and reflection, and internal thinking prompts are examples of why token costs are escalating unpredictably. Public cloud-based LLMs being per token based charges, have seen unexpected cost increases.
- Governance gaps: Engineering teams are building applications and deploying autonomous agents faster than privacy, security, governance and operational processes can keep up. This has created a fractured environment where leaders may not have a centralized view of which models, tools and agents are being used and how. Infrastructure must enforce fine-grained privacy, infrastructure and data access controls that limit what each agent can access and modify, even when the underlying model is trusted.
VMware Private AI Cloud

To address these challenges Broadcom has introduced VMware Private AI Cloud, enabling enterprises to– Scale AI Cost-Effectively, Operate More Securely, and Innovate Rapidly. Built on Broadcom’s advanced software capabilities, VMware Private AI Cloud gives organizations a production-ready path to securely building, running, and governing inference workloads, agentic applications, and traditional enterprise workloads together on VCF, Offering diverse hardware, model, and accelerator choices
- Scale AI Cost-Effectively: VMware Private AI Cloud addresses the three core AI cost drivers: hardware CapEx, operational complexity, and token economics (tokenomics). VCF supports GPUs, CPUs, and accelerators from leading vendors, along with server hardware from major OEM and ODM vendors, allowing customers to run heterogeneous clusters cost-effectively. To optimize tokenomics and resource usage, it features token monitoring, multi-tenant model sharing, enhanced GPU/vGPU tracking, and an AI metrics observability dashboard. One of the key innovations of VMware Private AI Cloud is VMware AI Factory, which provides a fast path to go from bare metal deployment to the first model deployment.
- Operate More Securely: Designed with a defense-in-depth approach aligned to NIST CSF 2.0, VCF protects against AI-accelerated threats by minimizing the attack surface and enabling continuous compliance. Automated, non-disruptive updates keep systems current, while VMware vDefend uses virtual patching and hypervisor-level lateral security with micro-segmentation to enforce Zero Trust and block exploits. Furthermore, vDefend’s multi-layer threat defense and VMware Avi Load Balancer’s web application firewall and API protection prevent sophisticated attacks.
- Innovate Rapidly for the Agentic AI Era: Unlike traditional apps, autonomous AI agents can act unchecked, exceed scope, or misinterpret instructions. Consequently, trust depends on robust controls and data integrity. VMware Tanzu Platform, with VMware vDefend, provides a foundation for trustworthy enterprise agents via a deny-by-default architecture, a prebuilt harness, and a curated marketplace
Let’s get into the details of VMware AI Factory, the key innovation in VMware Private AI Cloud.
VMware AI Factory: Fueling VMware Private AI Cloud

VMware AI Factory is the software-defined foundation of VMware Private AI Cloud. It provides customers a simplified path to production AI with new automation innovations for deploying AI-ready infrastructure and supporting Day 2 operations. With VMware AI Factory, customers can achieve faster time to first model deployment and better manage AI tokenomics.
VMware AI Factory brings AI applications directly to enterprise private data within a secure private cloud environment. VCF’s unique infrastructure automation capabilities can reduce the time from bare metal server deployment to serving the first AI model from weeks to a matter of hours. VMware AI Factory streamlines AI infrastructure management by fully automating hardware provisioning, software stack enablement, and end-to-end lifecycle management. By integrating hardware and software operations into a unified, automated solution, organizations can rapidly scale AI workloads while minimizing operational complexity.
Partnerships Enabling VMware AI Factory

VMware AI Factory combines VCF with certified Dell PowerEdge servers and VCF AI ReadyNodes from Cisco, Lenovo, Supermicro and others, and customers’ preferred AI software and accelerator architectures.
Broadcom and AMD are collaborating to deliver a VMware AI Factory that pairs VCF with AMD Instinct GPUs and the open AMD ROCm software ecosystem. Zero-touch provisioning will orchestrate the end-to-end deployment of the entire stack, from vSphere and vSAN through Kubernetes and the AMD GPU operator, and the AMD DVX driver can attach GPUs to large VMs consumed by a VMware vSphere Kubernetes Service cluster.
To further streamline VCF AI Factory deployment, Broadcom has announced a new partnership with MetalSoft to deliver integrated heterogeneous bare metal automation for VCF that drops bare-metal provisioning time from weeks to minutes. The integration will help IT provision or repave physical servers from multiple vendors directly through the VCF management console, unifying the software and hardware lifecycle into a single operational model and eliminating the need for vendor-specific tools for hardware and firmware management.
New Capabilities delivered to VMware AI Factory
VMware Cloud Foundation (VCF) Private AI Services help make AI operational, governable, and cost-effective and are included as part of VCF. We will continue to expand the number of services offered through VCF Private AI Services. Let’s explore these capabilities.
Multi-tenant Model Sharing

Model Runtime will be enhanced to allow the secure sharing of AI models between tenants or separate lines of business in their namespaces while maintaining full data privacy for each. Practically, this means enterprises and cloud service providers can have one Model Runtime service running and scaling models for their entire organization, and then each team or division can have separate private and secure namespaces for their private data. This capability eliminates the need for organizations to deploy redundant copies of the model, wasting GPU allocation and infrastructure resources. It also maintains data privacy and isolation while optimizing TCO.
Future Release Capabilities
AI Gateway

To balance the needs today of token cost optimization, use cases and performance enterprises need both cloud and on-premises deployed models. Cloud LLMs can have high token usage hence governance is paramount. AI Gateway capability will greatly assist with solving these competing tradeoffs. Let’s get into the specifics:
- Intelligent Prompt Routing: Dynamically maps and distributes incoming requests to the most suitable on-premises or cloud models based on factors such as use cases, domain specialization, and token costs to optimize performance.
- Usage and Token Limiting: User-level usage and token limiting helps minimize token usage.
- Application Authorization: Identifies the requesting application before routing the prompt. The AI Gateway uses OpenID Connect token-based authorization to enforce access controls.
Secure Agent Framework

Autonomous AI agents generate and execute code dynamically, creating significant security risks without strict controls. An unchecked agent—without deterministic limits on what an agent can access—could accidentally run catastrophic commands, such as deleting VMs or wiping on-premises databases, as demonstrated by several cloud-based AI providers’ test agents running rogue.
We will release 2 related capabilities to provide control and safeguards for AI Agents:
- Sandboxing: This will establish a secure virtualized container space where dynamic, agent- generated code is isolated and executed without affecting the broader environment.
- Agent Harness: This will establish the control layer that defines how agents are invoked, what tools they can access, how they communicate with each other, and how their outputs are validated before being acted upon.
Model Autoscaling

AI workloads cannot react dynamically to sudden spikes in requests or agentic loops. To address this, we will introduce Model Autoscaling. With this capability, the admin or AI operator will be able to set a threshold for latency and sessions, and the AI model will be automatically scaled when those thresholds are reached so that latency and performance SLAs are maintained. When usage drops below a threshold, the workload scales down.
With this capability, enterprises will gain performance improvements and lower TCO through event-driven scaling that keeps token latency low, and avoid over-provisioning of GPU resources. Enterprises can further optimize their AI investments through more efficient sharing of expensive GPU resources between different workloads.
Additional AI innovations and announcements from Explore
Broadcom’s Support for a Wide Variety of Models delivered through VCF
The rapidly evolving AI space requires the availability of different models. The need for a variety of AI models stems from several competing factors like privacy, security, governance, token costs and specialization and domain expertise.
Broadcom is committed to helping enterprises overcome these competing factors by supporting open-weight models and strictly commercial solutions. VMware AI Factory gives enterprises a production-ready path to running leading AI models on-premises. Leveraging vLLM as the default model runtime gives customers the ability to run more than 150 open source models performance-optimized on VCF. Today, Broadcom is announcing that the following models have been tested to run on VCF:
- Nemotron 3: The NVIDIA Nemotron 3 family of open, multimodal models delivers leading accuracy and efficiency to help agents complete tasks faster. Combining hybrid Mamba-Transformer MoE architecture, one-million-token context window and multi-environment reinforcement learning, Nemotron 3 enables scalable, long-running agentic workflows across enterprise applications.
- Gemma 4: Google DeepMind’s latest open-source, open-weight multimodal model family, purpose-built for developers and the research community for bringing local execution, and enabling enterprises to build and deploy autonomous AI agents.
- cotomi: NEC’s proprietary AI model optimized for Japanese language, trained on curated, highly reliable datasets. It empowers enterprises by seamlessly combining high-speed processing with a 40% improvement in token efficiency.
- Qwen3.8-27B: Alibaba’s Qwen3.8-27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal interaction, and long-running agent tasks, with flexible thinking that can be enabled or disabled. It is a native multimodal dense model with 27 billion parameters, designed for efficient local deployment and commercial use.
- GLM 5.2: Z.ai (formerly Zhipu AI)’s open-source General Language Model enables enterprises to deploy coding and reasoning agents locally for multi-step autonomous workflows with data sovereignty and optimal hardware performance.
VCF Achieves NVIDIA Certification: AI Workload Performance Validated to Operate at Near Bare Metal Performance
Recently NVIDIA launched the NVIDIA-Certified Hypervisors program. This program certifies that hypervisors in this program operate at near bare metal performance for AI and HPC workloads.
We are happy to announce that VMware vSphere 9.1 (and all future vSphere 9 releases) achieved NVIDIA-Certified Hypervisor certification. AI and HPC workloads on VCF are now certified to operate at near bare metal performance. This certification enables customers to confidently VCF as the performance-optimized private cloud platform, with support across major ecosystem partners, to power AI and accelerated computing applications in enterprise data centers. For further details, please read this recently published blog.
Partners Added to the VCF Private AI Services Ecosystem
We’re continuing to enhance a stronger AI ecosystem for the enterprise. Beyond core features, VCF Private AI Services is expanding its reach through new strategic alliances. Please join us in welcoming our latest partners:
- Appian: Appian provides AI automation for mission-critical work, automating complex processes in large enterprises and governments. The Appian Platform is known for its unique reliability and scale, backed by over 25 years of deep expertise in enterprise operations. For installation documentation of Appian on VCF, visit this page.
- ClearML: ClearML provides an AI orchestration layer that governs GPU access, models, and agents across VMware Cloud Foundation. This increases GPU utilization and lowers the cost of running AI workloads, giving enterprises an inbuilt path to AI-as-a-Service. Learn more about ClearML here.
- Eve Security: Eve provides the governance, observability, and runtime control enterprises need to safely scale AI agents. With its agent in the loop, it automatically interrogates high-risk or anomalous activity, enriches decisions with context from identity, DLP, and underlying systems, and applies real-time controls when needed. Learn more about Eve Security, here.
- Solo.io: Solo.io builds open source agentic infrastructure for enterprises running AI in production. Its kagent runtime and agentgateway data plane give platform teams a way to deploy AI agents and govern every model, tool, and agent-to-agent call, on infrastructure they own and operate. For additional information about Solo.io and Broadcom, read this blog.
- TrueFoundry: TrueFoundry provides an enterprise-grade AI Gateway that encompasses an LLM Gateway, MCP Gateway, and Agent Gateway—enabling enterprises to connect, observe, and govern agentic AI applications across providers from a single control plane. For additional details about TrueFoundry go here.
Want to know more?
- Visit VMware.com/AIML for more information.
- Connect with us on Twitter at @VMwareVCF and on LinkedIn at VMware VCF.
Discover more from VMware Cloud Foundation (VCF) Blog
Subscribe to get the latest posts sent to your email.