[{"content":"Introduction \u0026ldquo;Run AI on our own infrastructure\u0026rdquo; sounds like a single project. With VMware Private AI Foundation with NVIDIA (PAIF) on VCF 9.1.x it isn\u0026rsquo;t, there are three different ways to use your GPUs, and the one you pick changes what you build before the first workload runs. Add a licensing model where two vendors each sell you a different piece, and the planning decisions matter more than the deployment steps. This post covers the fundamentals, what the domain is made of, the three ways to consume it, and what you need to license.\nWhat a Private AI Workload Domain Is PAIF builds its GPU capacity into a dedicated VI workload domain, made up of GPU-enabled ESX hosts, an NSX Edge cluster, and a Supervisor on the domain\u0026rsquo;s default cluster. The management domain stays outside that build, so keep it GPU-free. Broadcom\u0026rsquo;s documented starting point is at least three GPU-enabled ESX hosts in the initial cluster.\nSupervisor is the Kubernetes control plane built into vSphere, so there is no separate Kubernetes product to install. GPUs reach virtual machines in one of two ways, NVIDIA vGPU, which shares a physical GPU between workloads, or DirectPath I/O, which gives a VM or Kubernetes node exclusive access to a GPU.\nThree Ways to Use the GPUs PAIF doesn\u0026rsquo;t give you one way to run AI workloads, it gives you three, and the right one depends on who is consuming the GPUs:\n-\u0026gt; Deep Learning VMs (DLVMs): ready-made virtual machines validated by NVIDIA and VMware, with the NVIDIA drivers and common machine learning tooling already installed. Data scientists request one through a VCF Automation catalog item or kubectl, and VI administrators can deploy one from the vSphere Client. It is the quickest way to prove your GPU stack works end to end.\n-\u0026gt; Private AI Services: the path for building applications on large language models (LLMs), installed by a VI administrator as a Supervisor Service. It bundles a Model Gallery in Harbor for storing models, a Model Runtime that serves them, knowledge bases for retrieval-augmented generation (RAG), and an Agent Builder, all managed as one integrated service.\n-\u0026gt; GPU-accelerated VKS clusters: Kubernetes clusters whose worker nodes have GPUs, for teams that want to run their own AI containers. The AI Kubernetes Cluster catalog item in VCF Automation deploys them with the NVIDIA GPU Operator, which sets up the NVIDIA driver on the workers, and it asks for an NVIDIA NGC API key when you request one.\nLicensing Spans Two Vendors PAIF involves up to three entitlements. A VCF subscription and the Private AI Foundation license both come from Broadcom, which offers Private AI Foundation with NVIDIA as a solution license for VCF. The NVIDIA AI Enterprise (NVAIE) license comes directly from NVIDIA, and Broadcom\u0026rsquo;s requirements state it is needed for vGPU workloads, for the host driver and the guest drivers. With VCF 9.1, DirectPath workloads don\u0026rsquo;t need it, so the way you attach GPUs affects what you buy from NVIDIA.\nOne placement detail is worth knowing up front. The Private AI Foundation license is allocated to the GPU-enabled workload domains, but the guided deployment UI in the vSphere Client only appears if the license is also assigned to the management domain.\nArchitectural Overview What 9.1.1 Changed VCF 9.1.1 keeps the same three paths. Its Private AI Foundation release notes list two additions, a new version of Private AI Services, 3.0, and a new Deep Learning VM image. Check the release notes for your version before you upgrade an existing deployment, because AI Kubernetes blueprints and NVIDIA GPU Operator versions have had compatibility notes in earlier releases.\nA Quick Checklist Choosing a path? Deep Learning VMs to validate the GPU stack, Private AI Services to build LLM applications, GPU-accelerated VKS clusters for teams that want Kubernetes-native AI. Planning the hosts? A dedicated GPU workload domain with at least three GPU hosts, and a GPU-free management domain. Budgeting licenses? The Private AI Foundation license from Broadcom, plus NVAIE from NVIDIA if you use vGPU. Check whether DirectPath covers your workload before buying NVAIE. Upgrading an existing deployment? Read the release notes for your target version first. What\u0026rsquo;s Next Next in this series: Advanced Services for VCF, VPC, Load Balancing, and Network Observability.\nFurther Reading (Official Broadcom Documentation) VMware Private AI Foundation with NVIDIA 9.1 Requirements for Deploying VMware Private AI Foundation with NVIDIA Assign a Private AI Foundation License to the Management Domain Deploying AI Workloads on VKS Clusters Delivering Generative AI Applications by Using Private AI Services Deploy VCF on Supermicro HGX Servers with NVIDIA GPUs for AI Workloads VCF 9.1: The Secure, Cost-Effective Private Cloud Platform for Production AI VMware Private AI Foundation with NVIDIA 9.0.x Release Notes VMware Cloud Foundation 9.1.1.0 Release Notes VMware Private AI Foundation with NVIDIA 9.1.x Release Notes ","permalink":"https://virtualizationgurus.pages.dev/posts/vcf9-private-ai-workload-domain/","summary":"\u003ch2 id=\"introduction\"\u003eIntroduction\u003c/h2\u003e\n\u003cp\u003e\u0026ldquo;Run AI on our own infrastructure\u0026rdquo; sounds like a single project. With VMware Private AI Foundation with NVIDIA (PAIF) on VCF 9.1.x it isn\u0026rsquo;t, there are three different ways to use your GPUs, and the one you pick changes what you build before the first workload runs. Add a licensing model where two vendors each sell you a different piece, and the planning decisions matter more than the deployment steps. This post covers the fundamentals, what the domain is made of, the three ways to consume it, and what you need to license.\u003c/p\u003e","title":"Private AI Workload Domain: GPU Nodes, AI Kubernetes, and Private AI Services"},{"content":"Introduction Disaster recovery and ransomware recovery get planned in the same conversation more often than not, same team, same budget line, sometimes the same runbook with a different label on it. Broadcom\u0026rsquo;s own architecture disagrees with that framing. Operational DR assumes your primary environment is healthy but unreachable or impaired, the priority is RTO, automation, and minimal data loss. Cyber recovery assumes your primary environment is compromised, every action requires forensic isolation, immutability, and an air-gap mentality. Same team can run both. Structurally, they\u0026rsquo;re not the same discipline, and VCF 9.1\u0026rsquo;s tooling treats them as genuinely separate.\nThe Isolated Recovery Environment Exists for One Purpose Only Central to any cyber recovery strategy is the Isolated Recovery Environment, IRE, a cyber recovery \u0026ldquo;clean room\u0026rdquo; disconnected from the production data center, used to safely power on, inspect, and recover ransomware-infected workloads. Broadcom\u0026rsquo;s validated design is explicit about scope: the IRE is not used for test/development, burst capacity, or anything else. One purpose, full stop.\nThe reference design specifies exactly what that isolation looks like:\n-\u0026gt; A dedicated vSAN storage cluster using Express Storage Architecture (ESA) stores the replicated virtual machines, plus a separate compute cluster with its own NSX edge clusters and a remote datastore mounted from that storage cluster.\n-\u0026gt; A Tier-1 gateway and dedicated network segment, configured in the isolated workload domain\u0026rsquo;s NSX, is what the recovered test virtual machines actually connect to, with DHCP serving IP addresses and DNS settings on that segment.\n-\u0026gt; A Python isolation script creates graduated levels of network isolation on the IRE, letting you control how contained a given recovery session is depending on what phase of investigation you\u0026rsquo;re in. This does require its own Distributed Firewall license on NSX, worth budgeting for separately rather than discovering mid-incident.\n-\u0026gt; Outbound access to your EDR platform is deliberately allowed from inside the otherwise-isolated environment, specifically so recovered VMs can have security sensors installed and get analyzed before anything is trusted.\n-\u0026gt; DNS and NTP in the IRE are both intentionally separate from what the protected production instance uses, on the reasoning that the IRE\u0026rsquo;s path to the internet shouldn\u0026rsquo;t retrace anything the compromised environment touches.\nThe Recovery Workflow Itself VCF 9.1\u0026rsquo;s Protection and Recovery capability, paired with VMware Advanced Cyber Compliance, extends the familiar site-recovery pattern into something built specifically for cyber incidents rather than just outages. In Broadcom\u0026rsquo;s recovery states, a snapshot sits In backup until you validate it, then runs In validation, powered on inside the IRE (Protection and Recovery calls it the clean room), where it gets inspected with EDR tooling. That includes built-in AI/ML-powered detection and integrated analysis through Carbon Black (managed by Broadcom) or CrowdStrike (customer-managed). Once it\u0026rsquo;s Validated it moves to Staged, gets Recovered at the secondary site, and is then reprotected.\nCyber recovery is supported with vSAN snapshots and replication, and only vSAN-protected VMs are supported, so scope your protection groups with that in mind. VCF 9.1 ships two purpose-built presets rather than making you hand-configure retention every time: a Ransomware Recovery preset (1-hour RPO, last snapshot kept, hourly retained for a day, daily for a week, weekly for a month, monthly for six months) and a lighter Short-Term Retention preset for less critical workloads (same 1-hour RPO, but retention tapers off after the weekly tier). Protection groups can now be assigned by vSphere tag as well as by static or wildcard name, which matters once you\u0026rsquo;re managing this at real fleet scale rather than a handful of VMs.\nWhere VMware Live Site Recovery (Formerly SRM) Fits Now The naming has moved twice, so it\u0026rsquo;s worth getting straight. Broadcom rebranded Site Recovery Manager (SRM) to VMware Live Site Recovery in 2024, and VMware Live Recovery has since been renamed and integrated into VCF as VCF Protection and Recovery. Broadcom\u0026rsquo;s 9.1 documentation frames Protection and Recovery as three capabilities: operational recovery on vSAN local snapshots, disaster recovery that orchestrates array-based and host-based replication, and cyber recovery in a clean room. The 9.1 Protection and Recovery documentation doesn\u0026rsquo;t use the SRM name at all, but Broadcom\u0026rsquo;s validated on-premises ransomware recovery design still refers to Site Recovery Manager (now Live Site Recovery) and vSphere Replication for the recovery plan, so expect to meet both names depending on which document you\u0026rsquo;re reading. What changed is that SRM-style orchestration is no longer the only recovery path, since vSAN snapshots now cover operational recovery and the cyber recovery workflow is built on vSAN-protected VMs. Confirm which orchestrator your design actually relies on against the validated design before you commit to it.\nArchitectural Overview What Can Actually Go Wrong Broadcom\u0026rsquo;s own 9.1 release notes for Protection and Recovery list a real failure mode worth knowing before you\u0026rsquo;re mid-incident: the Ransomware Recovery End workflow can fail with \u0026ldquo;Operation interrupted: running,\u0026rdquo; caused by the srm service crashing and auto-restarting immediately after, which leaves the Recovery Plan stuck in active Ransomware Recovery Mode with no End action visible in the UI. The documented workaround is calling the rwrEnd API directly against the Protection and Recovery server\u0026rsquo;s Managed Object Browser, RecoveryManager object, rather than waiting for the UI option to reappear on its own.\nVPC Isolation Is the Same Mechanism, Different Context The network isolation techniques underpinning the IRE, Tier-1 gateways, DFW-based segmentation, scriptable isolation levels, aren\u0026rsquo;t a special-purpose invention for cyber recovery. They\u0026rsquo;re the same VPC isolation model this series has already covered for centralized vs. distributed connectivity in normal workload domains. What changes in the IRE context isn\u0026rsquo;t the mechanism, it\u0026rsquo;s the intent: isolation here exists to contain a suspected-compromised workload during forensic inspection, not to segment tenants or manage north-south traffic flow.\nA Quick Checklist Planning DR and cyber recovery as the same runbook? Split them, the assumptions about what state your primary environment is in are opposite, and a runbook that\u0026rsquo;s right for one is actively dangerous for the other. Building an IRE? Budget for the separate DFW license the isolation script depends on, and plan DNS/NTP infrastructure that\u0026rsquo;s genuinely separate from production, not just logically separate. Configuring retention? Start from the Ransomware Recovery or Short-Term Retention presets rather than hand-building a schedule, they\u0026rsquo;re tuned defaults, not just examples. Stuck in Ransomware Recovery Mode with no End option showing? Check for a crashed and auto-restarted srm service before assuming the UI is simply broken, the rwrEnd API workaround is the documented fix. What\u0026rsquo;s Next Next in this series: Private AI Workload Domain, GPU Nodes, AI Kubernetes, and Private AI Services.\nFurther Reading (Official Broadcom Documentation) Isolated Recovery Environment Design for On-Premises Ransomware Recovery Cyber Recovery in Protection and Recovery 9.1 Ransomware Recovery States VMware Live Site Recovery 9.0 Release Notes (Site Recovery Manager rebrand) Protection and Recovery 9.1 Release Notes VMware vSAN Protection and Recovery Enhancements for VCF 9.1 Continuous Compliance, Integrated Cyber Recovery and Enhanced Platform Security for VCF 9.1 ","permalink":"https://virtualizationgurus.pages.dev/posts/vcf9-dr-ransomware-recovery/","summary":"\u003ch2 id=\"introduction\"\u003eIntroduction\u003c/h2\u003e\n\u003cp\u003eDisaster recovery and ransomware recovery get planned in the same conversation more often than not, same team, same budget line, sometimes the same runbook with a different label on it. Broadcom\u0026rsquo;s own architecture disagrees with that framing. Operational DR assumes your primary environment is healthy but unreachable or impaired, the priority is RTO, automation, and minimal data loss. Cyber recovery assumes your primary environment is compromised, every action requires forensic isolation, immutability, and an air-gap mentality. Same team can run both. Structurally, they\u0026rsquo;re not the same discipline, and VCF 9.1\u0026rsquo;s tooling treats them as genuinely separate.\u003c/p\u003e","title":"DR \u0026 Ransomware Recovery: Isolated Recovery, Protection and Recovery, and VPC Isolation"},{"content":"Introduction Day-2 in most infrastructure conversations gets reduced to one word: patching. VCF 9.1 treats it as two separate, ongoing jobs that happen to share a control plane. The first is applying fixes without breaking anything that\u0026rsquo;s running. The second is proving, continuously, that what\u0026rsquo;s running still matches what you intended to deploy. Broadcom built genuinely different tooling for each, and conflating them is a good way to think you\u0026rsquo;re covered on one when you\u0026rsquo;ve only handled the other.\nPatching Isn\u0026rsquo;t One Workflow, It\u0026rsquo;s Three VCF 9.1\u0026rsquo;s patching architecture starts from a fact that\u0026rsquo;s easy to forget mid-incident: not every component carries the same disruption risk if something goes wrong while patching it. Broadcom\u0026rsquo;s own engineering team splits the entire stack into three layers, each with a patching mechanism matched to what that layer can tolerate:\n-\u0026gt; Management layer (VCF Management Services, VCF Operations, VCF Automation, Cloud Proxy, VCF Operations for Networks): architecturally separate from anything running workloads, so patching here carries no workload risk at all. This layer runs on a declarative model, you define a target version, and the Fleet Lifecycle service orchestrates the rest across the fleet, replacing what used to be manual, error-prone version-by-version steps.\n-\u0026gt; Control plane layer (vCenter, NSX Manager, vSphere Supervisor, VKS): everything here has to stay available during the patch, since every other operation depends on it. The mechanism is matched to the patch type: vCenter Quick Patch handles security and minor fixes, Reduced Downtime Upgrade handles full version transitions, and vSphere Supervisor and VKS clusters get rolling updates. NSX keeps at least two manager nodes active throughout so management never actually drops.\n-\u0026gt; Data plane layer (ESX, vSAN, NSX Edge): this is where workloads actually run, and it carries the strictest constraint of the three, patch the host without disrupting what\u0026rsquo;s on it. ESX Live Patch applies fixes directly in memory: no maintenance window, no evacuation, no reboot, and as of 9.1 this now extends to TPM-enabled hosts. When a reboot genuinely can\u0026rsquo;t be avoided, Quick Boot, pre-staging, and live vMotion evacuation keep the actual impact as small as possible.\nAcross all three layers, the same two disciplines apply: prechecks confirm a patch is likely to succeed before it commits to anything, and recoverable, migration-based designs give you a fallback when a precheck was wrong.\nDeclarative Lifecycle Management, in Practice The mental shift underneath all three layers: you stop issuing a sequence of individual patch commands and start declaring a target version for the environment. VCF Operations, through the Fleet Lifecycle service, works out and executes the actual sequence needed to get there. This isn\u0026rsquo;t just a UX simplification, it\u0026rsquo;s what makes the scale numbers in 9.1 possible: a single instance now supports up to 5,000 ESX hosts, and parallel upgrade capacity has quadrupled to 256 clusters at once. Neither of those numbers is reachable if a human is still sequencing individual patch steps by hand.\nCompliance Is a Separate, Continuous Job Patched doesn\u0026rsquo;t mean compliant. A host can be running the latest build and still have drifted away from the configuration baseline you actually intended, a changed MOTD, a modified role permission, a setting an engineer tweaked during troubleshooting and never reverted. VCF Operations treats catching that drift as its own workflow, not a side effect of patching:\n-\u0026gt; Configuration Management (the current name for what was previously called Configuration Drifts) lets you build configuration templates for vCenter and cluster objects, then schedule drift detection against them on a recurring basis rather than checking manually. For vSphere Configuration Profile-enabled clusters specifically, you get aggregated drift status across the whole fleet in one view.\n-\u0026gt; Drift reports are downloadable as PDFs, scoped to whichever vCenter instances and retention period you choose, useful as an actual audit artifact rather than something you have to screenshot.\n-\u0026gt; Git integration makes source control the template\u0026rsquo;s owner. Once you connect a Git repository, VCF Operations treats Git as authoritative for template versioning rather than the UI itself, which matters if you\u0026rsquo;re trying to run configuration-as-code discipline across a fleet rather than one-off manual edits.\n-\u0026gt; VMware Advanced Cyber Compliance (ACC) goes a step further than detection. Built on VMware Salt, it continuously monitors your VCF 9.1 stack against a chosen security baseline, either a Security Configuration Guide (SCG) or PCI-DSS, and when an ESX host drifts out of compliance, ACC can automatically remediate it back to the desired state rather than just flagging it for someone to fix later.\nArchitectural Overview What\u0026rsquo;s Actually New Architecturally in 9.1 Two structural changes are worth knowing if you\u0026rsquo;re coming from 9.0: the standalone VCF Operations Fleet Management Appliance no longer exists as its own component, its functionality (service registry, imported certificates, certificate signing requests) is absorbed into VCF Operations 9.1 directly, with the Fleet Lifecycle component taking over orchestration duties. And licensing moved out of VCF Operations entirely into a dedicated license server, installed automatically alongside VCF Operations 9.1 rather than living inside it.\nThis is the same pattern this series has already flagged once with the Identity Broker\u0026rsquo;s Appliance-to-Instance consolidation, VCF 9.1 is consistently folding what used to be standalone appliances into shared, unified services rather than just adding features to the existing ones.\nPatch Sequencing Gets Stricter in 9.1.1 The layered model above still holds exactly as described, but VCF 9.1.1\u0026rsquo;s own release notes add a hard sequencing rule worth knowing before your next maintenance window: when moving from 9.1.0.x to 9.1.1, you patch the VCF Management Services Fleet Lifecycle component first, before touching anything else. VCF Operations itself cannot be patched in parallel with other components, and before patching ESX hosts you must first patch the VCF Operations instance and the license servers connected to it.\nNone of this changes the three-layer model, it reinforces it: Management layer goes first because everything downstream depends on Fleet Lifecycle being current. It does mean \u0026ldquo;patch in whatever order is convenient\u0026rdquo; is no longer a safe assumption at 9.1.1, the order is now enforced, not just recommended.\nChoosing Where to Look First: A Quick Checklist Just deployed or upgraded to 9.1? Confirm your Fleet Lifecycle component is current before patching anything else, dependencies flow from it outward. Patching 9.1.0.x to 9.1.1? Patch order is enforced, not optional: VCF Management Services Fleet Lifecycle first, then VCF Operations and its connected license servers, before ESX hosts. Planning a patch cycle? Know which layer you\u0026rsquo;re touching before you start, the acceptable disruption and the mechanism both change by layer. Haven\u0026rsquo;t set up drift detection yet? It\u0026rsquo;s a separate configuration step from patching, being current on patches tells you nothing about configuration drift. Running a regulated environment? Advanced Cyber Compliance\u0026rsquo;s auto-remediation is worth deliberately choosing, or deliberately declining, rather than leaving on the default. What\u0026rsquo;s Next Next in this series: DR \u0026amp; Ransomware Recovery, Isolated Recovery, SRM, and VPC Isolation.\nFurther Reading (Official Broadcom Documentation) Faster Security Patching with Fewer Disruptions in VCF 9.1 Scale, Simplify, and Secure Your Private Cloud Operations with VCF 9.1 Securing your VMware Cloud Foundation 9.1 Environment Fleet Management: Configuration Management Overview VCF Operations 9.1.0.0 What\u0026rsquo;s New VMware Cloud Foundation 9.1.1.0 Release Notes, Getting to 9.1.1 ","permalink":"https://virtualizationgurus.pages.dev/posts/vcf9-day2-operations/","summary":"\u003ch2 id=\"introduction\"\u003eIntroduction\u003c/h2\u003e\n\u003cp\u003eDay-2 in most infrastructure conversations gets reduced to one word: patching. VCF 9.1 treats it as two separate, ongoing jobs that happen to share a control plane. The first is applying fixes without breaking anything that\u0026rsquo;s running. The second is proving, continuously, that what\u0026rsquo;s running still matches what you intended to deploy. Broadcom built genuinely different tooling for each, and conflating them is a good way to think you\u0026rsquo;re covered on one when you\u0026rsquo;ve only handled the other.\u003c/p\u003e","title":"Day-2 Operations: Lifecycle, Patching, and Compliance via VCF Operations"},{"content":"Introduction \u0026ldquo;Air-gapped\u0026rdquo; gets treated as a single scenario in most VCF conversations, you either have internet access or you don\u0026rsquo;t. VCF 9.1\u0026rsquo;s software depot architecture disagrees: it defines three distinct connection modes, and the two that cover no-internet environments, Offline Depot and Disconnected, are genuinely different operational models, not two names for the same workaround. Picking the wrong one, or assuming they\u0026rsquo;re interchangeable, is a common source of confused prechecks during upgrade planning.\nSoftware Depot: The Component Behind All Three Modes Starting with VCF 9, lifecycle management for every VCF component runs through a single component called the software depot, accessed via the VCF Operations UI. Every VCF fleet has one fleet-level software depot, deployed on the first VCF Instance, that can serve binaries to every other Instance in the fleet. When you deploy a new VCF Instance on version 9.1 or later, the fleet-level depot is assigned to it automatically. If you\u0026rsquo;re upgrading an existing fleet to 9.1, whatever depot configuration was already set in VCF Installer or SDDC Manager carries forward into the new fleet-level depot rather than needing to be redone.\nThere\u0026rsquo;s also a scaling knob worth knowing: if network latency between your first Instance\u0026rsquo;s software depot and another Instance in the fleet exceeds 150 ms, Broadcom\u0026rsquo;s guidance is to deploy a secondary, local software depot within that other Instance rather than have it pull binaries across a slow link.\nThree Connection Modes, Not Two Every software depot instance is configured in exactly one of three modes, and Broadcom\u0026rsquo;s own documentation is explicit that the choice depends entirely on how your environment reaches the internet:\n-\u0026gt; Connected: your environment has internet access, direct or through a proxy. The software depot registers directly with Broadcom, downloading binaries, hardware compatibility data, software interoperability data, and lifecycle metadata on its own. Registration itself has a specific flow: you copy the depot\u0026rsquo;s software depot ID, log into the VCF Business Services console, register the component under the correct tenant, generate an activation code against that ID, then paste the code back into VCF Operations to validate. Binaries are cached and managed internally, and there\u0026rsquo;s no manual delete option in this mode.\n-\u0026gt; Offline Depot: your environment can\u0026rsquo;t reach the internet, but you own a private server that can. You run the VCF Download Tool against that server to pull binaries down, and the software depot doesn\u0026rsquo;t download anything locally itself, it proxies every request from VCF components straight to your privately-owned server. If that server serves content over HTTPS, there\u0026rsquo;s a real prerequisite step: SSH into SDDC Manager, pull the depot server\u0026rsquo;s TLS certificate with openssl s_client, import it into the local Java certificate store with keytool, and restart the lcm service, all before the mode will actually connect.\n-\u0026gt; Disconnected: your environment can\u0026rsquo;t reach the internet and you don\u0026rsquo;t have an offline depot server at all. You run the VCF Download Tool on any separate computer that does have internet access, then manually upload the resulting binaries straight into the software depot. This is the one mode where a cleanup command exists to delete uploaded binaries again, Connected and Offline Depot both cache internally with no delete option.\nWhy the Offline Depot vs Disconnected Distinction Actually Matters Read those last two modes again: the practical difference is whether a dedicated, always-on depot server exists in your environment. Offline Depot assumes one does, and the software depot proxies to it continuously. Disconnected assumes one doesn\u0026rsquo;t, and every set of binaries you need gets manually carried in and uploaded as a one-off action. Treating \u0026ldquo;we\u0026rsquo;re air-gapped\u0026rdquo; as a single answer skips over a real infrastructure decision, standing up and maintaining a depot server is a genuinely different operational commitment than running the download tool on a laptop each time you need new binaries.\nThis is also where mismatched assumptions cause real friction during upgrade planning: someone configures Offline Depot mode expecting the software depot to actively fetch on demand, when what\u0026rsquo;s actually been built is a Disconnected workflow, binaries staged manually, no ongoing proxy behavior. The two look similar from a distance, \u0026ldquo;no internet, binaries come from somewhere else\u0026rdquo;, but the software depot treats them as fundamentally different connection types, and troubleshooting one as if it were the other wastes real time.\nThe VCF Download Tool Is the Common Thread Both Offline Depot and Disconnected modes rely on the same underlying utility: the VCF Download Tool, a command-line tool purpose-built for downloading and managing VCF component and ESX binaries and metadata in environments without internet access. It provides commands for downloading, uploading, listing, and cleaning up binaries, the same tool, pointed at either a depot server (Offline Depot) or a direct upload target (Disconnected), depending on which mode you\u0026rsquo;re actually running.\nArchitectural Overview Choosing Between Them: A Quick Checklist Full internet access, direct or via proxy? Connected, register once with an activation code and you\u0026rsquo;re done. No internet, but you can stand up and maintain a dedicated server? Offline Depot, plan for the TLS certificate import step if that server runs HTTPS. No internet, and no appetite to run a standing depot server? Disconnected, budget the recurring manual effort of downloading and uploading binaries per upgrade cycle. Multi-Instance fleet with a slow link to one of them? Deploy a secondary software depot in that Instance regardless of which of the three modes your fleet-level depot uses. What\u0026rsquo;s Next Next in this series: Day-2 Operations, Lifecycle, Patching, and Compliance via VCF Operations.\nFurther Reading (Official Broadcom Documentation) Binary Management for VMware Cloud Foundation Configure a Software Depot Connection Mode Download Binaries to an Offline Depot by Using the VCF Download Tool Download Binaries to Software Depot in Disconnected Mode by Using the VCF Download Tool VCF Download Tool Command Reference Information ","permalink":"https://virtualizationgurus.pages.dev/posts/vcf9-bundle-management/","summary":"\u003ch2 id=\"introduction\"\u003eIntroduction\u003c/h2\u003e\n\u003cp\u003e\u0026ldquo;Air-gapped\u0026rdquo; gets treated as a single scenario in most VCF conversations, you either have internet access or you don\u0026rsquo;t. VCF 9.1\u0026rsquo;s software depot architecture disagrees: it defines three distinct connection modes, and the two that cover no-internet environments, Offline Depot and Disconnected, are genuinely different operational models, not two names for the same workaround. Picking the wrong one, or assuming they\u0026rsquo;re interchangeable, is a common source of confused prechecks during upgrade planning.\u003c/p\u003e","title":"Bundle Management: Online vs Offline/Air-Gapped"},{"content":"Introduction If Fleet collapsed three consoles into one control plane, the Identity Broker does the same job for authentication. VMware Identity Manager (vIDM) is superseded in VCF 9 by a fleet-native component called the Identity Broker. The new component takes over VCF single sign-on going forward, though Broadcom doesn\u0026rsquo;t force a rip-and-replace cutover: existing vIDM instances can keep running (still managed via VMware Aria Suite Lifecycle 8.x) as an authentication source for components like VCF Automation while you migrate the rest of the fleet on your own schedule. This post covers what the Identity Broker actually is, the two ways you can deploy it, and what a real migration from vIDM looks like, based on Broadcom\u0026rsquo;s official 9.1 documentation.\nFrom VMware Identity Manager to a Fleet-Native Identity Broker VCF single sign-on is the mechanism that lets components across the fleet (vCenter, VCF Operations, VCF Automation, NSX, log management, VCF Operations for Networks, and VCF Operations Orchestrator) authenticate against a shared set of identity providers instead of maintaining separate logins. Two notable exceptions: SDDC Manager and ESX are not covered by VCF single sign-on and keep their own authentication paths. In earlier VCF releases, vIDM often filled the fleet-authentication role. In VCF 9.1, that job belongs to the Identity Broker, a purpose-built component for VCF single sign-on across the fleet.\nTwo Deployment Modes: Embedded or Instance According to Broadcom\u0026rsquo;s documentation, the Identity Broker supports two deployment modes, and the choice affects where the component actually lives. Worth flagging up front: the change between releases isn\u0026rsquo;t just a rename. In VCF 9.0, the non-embedded mode ran as a standalone, externally-deployed multi-node appliance cluster, and Broadcom called it \u0026ldquo;Appliance\u0026rdquo; mode. In VCF 9.1, that same role is filled by \u0026ldquo;Instance\u0026rdquo; mode, but the underlying architecture changed with it: it\u0026rsquo;s now consolidated into VCF Management Services, the unified runtime VCF 9.1 introduced for centralized lifecycle and operations across components like Identity Broker, Log Management, and the License Server. Broadcom\u0026rsquo;s own upgrade documentation confirms this directly: the 9.0 external vIDB appliance cluster is \u0026ldquo;migrated directly into VCF Management Services\u0026rdquo; as part of the 9.1 upgrade, not simply relabeled. If you\u0026rsquo;re cross-referencing older 9.0-era material or blog posts, keep this in mind, the terminology changed because the architecture did.\n-\u0026gt; Embedded mode, the Identity Broker is configured directly inside the management domain vCenter of a VCF Instance. This is the simpler option, typically used within a single VCF Instance. It\u0026rsquo;s also a single point of failure: if the management domain vCenter goes down, the embedded Identity Broker goes down with it.\n-\u0026gt; Instance mode, the Identity Broker is deployed as its own dedicated VCF management services component within the management domain, separate from vCenter, running as a three-node cluster that tolerates a single node failure. The first Identity Broker instance is deployed in the primary VCF Instance, and you can optionally deploy additional Identity Broker instances in other VCF Instances across the fleet, though Broadcom\u0026rsquo;s guidance caps a single Instance-mode broker at up to five connected VCF Instances.\nBroadcom\u0026rsquo;s Identity Broker Detailed Design documentation covers the specific requirements and recommendations for choosing between the two.\nWhat Migration From vIDM Actually Involves For environments coming from VMware Identity Manager 3.3.7 GA (or its latest patch), Broadcom provides a dedicated migration path in VCF 9.1, targeting an Instance-mode Identity Broker running 9.1 or later, built around export/import scripts rather than a one-click in-place upgrade:\n-\u0026gt; Users and groups are migrated from vIDM to the Identity Broker directly.\n-\u0026gt; Sync settings for existing identity providers are compared and displayed side by side, but not automatically migrated, you review the comparison and adjust the Identity Broker\u0026rsquo;s sync settings yourself if needed.\n-\u0026gt; Component updates, if VCF Operations, VCF Automation, or NSX currently authenticate through vIDM, a separate update step repoints each of them to the new Identity Broker once it\u0026rsquo;s configured.\nThe migration tooling ships as OS-specific export/import binaries (Windows x86_64, macOS ARM64, Linux x86_64) that you download from Broadcom Support, run against your existing vIDM instance to export data, then import into the target Identity Broker with built-in data-integrity and compatibility validation.\nMigration Limitations Worth Knowing A few constraints matter when planning a migration, straight from Broadcom\u0026rsquo;s documented limitations:\n-\u0026gt; Only the Instance deployment mode is a supported migration target, you can\u0026rsquo;t migrate directly into an Embedded-mode Identity Broker.\n-\u0026gt; Local accounts, and local accounts using multifactor authentication, aren\u0026rsquo;t supported on the Identity Broker, nor is multifactor authentication paired with Active Directory.\n-\u0026gt; OAuth clients don\u0026rsquo;t migrate automatically, they need to be manually regenerated against the Identity Broker.\n-\u0026gt; If a single vIDM instance currently serves multiple components, all of them get repointed to the same new Identity Broker as part of component migration, there\u0026rsquo;s no partial cutover.\n-\u0026gt; Component migration is only supported for three components: VCF Operations, VCF Automation, and NSX. Anything else authenticating through vIDM needs a separate plan.\nArchitecture at a Glance The Mental Model Shift Legacy (vIDM era) VCF 9 VMware Identity Manager as a bolted-on identity source Identity Broker as a native VCF single sign-on component One identity config per tool Shared Identity Broker across VCF Operations, VCF Automation, and NSX Manual, ad hoc cutover between identity tools Documented export/import/component-update migration path Local accounts and MFA handled inconsistently Local + MFA combinations explicitly unsupported on the Broker, third-party IdP/AD integration is the expected pattern \u0026ldquo;Appliance mode\u0026rdquo;: standalone external appliance cluster (VCF 9.0) \u0026ldquo;Instance mode\u0026rdquo;: consolidated into VCF Management Services (VCF 9.1), architecture changed, not just the name Embedded to Instance Migration (VCF 9.1) A separate migration path exists for a different scenario than the vIDM migration above: moving an Identity Broker that\u0026rsquo;s already running in Embedded mode into Instance mode, without touching vIDM at all. This is new in VCF 9.1, Broadcom\u0026rsquo;s release notes for VCF Operations 9.1 list it explicitly: \u0026ldquo;Migration from embedded to instance deployment of the identity broker: Support for the migration of the identity broker from embedded mode to instance mode in VCF Operations.\u0026rdquo;\nWorth correcting a common assumption up front: this is not a zero-touch, click-and-done migration. It\u0026rsquo;s a supported, no-data-loss path, but it involves real manual steps on your end, including reconfiguring your external identity provider.\nPrerequisites: your Embedded-mode Identity Broker must already be at version 9.1 (it upgrades automatically as part of the vCenter instance upgrade, no separate step needed), and you need an Instance-mode Identity Broker already deployed somewhere in the fleet to serve as the migration target. Only VCF Instances with an existing Instance-mode broker show up as valid targets, so if none exists yet, you deploy one first.\nWhat the migration actually involves, from VCF Operations (Manage \u0026gt; Fleet Management \u0026gt; Identity \u0026amp; Access \u0026gt; VCF SSO Overview \u0026gt; select the broker \u0026gt; Actions \u0026gt; Migrate from Embedded to Instance):\n-\u0026gt; Data transfer is SFTP-based. You provide an SFTP host, port, username, password, and path; the migration exports data from the embedded broker and imports it into the target instance over that connection. Broadcom\u0026rsquo;s own guidance: use a dedicated, single-purpose SFTP account scoped to only that folder, and delete it once the migration completes.\n-\u0026gt; Identity provider config and user/group provisioning are suspended for the duration of the transfer. Plan the migration window accordingly rather than treating it as a background operation.\n-\u0026gt; You manually update your external identity provider afterward. The wizard shows you the new Identity Broker instance\u0026rsquo;s service provider details, and you take those into your IdP\u0026rsquo;s own admin console to update its configuration, there\u0026rsquo;s no automatic push to a third-party IdP. A Test Login step lets you validate the new connection before committing further.\n-\u0026gt; Component reconnection is only partly automatic. The wizard\u0026rsquo;s \u0026ldquo;Update Components\u0026rdquo; step reconnects and sanity-tests the components it knows how to handle, but VCF Operations, HCX, log management, and VCF Operations for Networks all need manual follow-up, along with any automation scripts using API clients against the old configuration.\n-\u0026gt; Canceling mid-migration deletes the target\u0026rsquo;s configuration, not just aborts cleanly, so treat the confirmation step as a real go/no-go point rather than a formality.\nBroadcom documents the exact procedure in \u0026ldquo;Migration of Identity Broker Embedded to Identity Broker Instance,\u0026rdquo; linked below, worth a full read before scheduling this given the manual IdP and component-update work involved.\nWhat\u0026rsquo;s Next Next in this series: Bundle Management, Online vs Offline/Air-Gapped Depots.\nFurther Reading (Official Broadcom Documentation) Managing Identity and Access With VCF Single Sign-On Deployment Modes of the Identity Broker Migrating VMware Identity Manager to Identity Broker Upgrade to Identity Broker 9.1 Migration of Identity Broker Embedded to Identity Broker Instance VCF Management Services Models ","permalink":"https://virtualizationgurus.pages.dev/posts/vcf9-identity-broker/","summary":"\u003ch2 id=\"introduction\"\u003eIntroduction\u003c/h2\u003e\n\u003cp\u003eIf Fleet collapsed three consoles into one control plane, the Identity Broker does the same job for authentication. VMware Identity Manager (vIDM) is superseded in VCF 9 by a fleet-native component called the Identity Broker. The new component takes over VCF single sign-on going forward, though Broadcom doesn\u0026rsquo;t force a rip-and-replace cutover: existing vIDM instances can keep running (still managed via VMware Aria Suite Lifecycle 8.x) as an authentication source for components like VCF Automation while you migrate the rest of the fleet on your own schedule. This post covers what the Identity Broker actually is, the two ways you can deploy it, and what a real migration from vIDM looks like, based on Broadcom\u0026rsquo;s official 9.1 documentation.\u003c/p\u003e","title":"VCF 9 Identity Broker: Retiring VMware Identity Manager for Unified Fleet Authentication"},{"content":"Introduction Security is a foundational concern in any private cloud deployment. VMware Cloud Foundation 9 delivers a layered security architecture that spans from the physical network underlay through the workload networking and compute layers, with VCF Operations providing centralized visibility into the security posture of all domains.\nIn this post, we explore the VCF 9 security architecture, focusing on the Distributed Firewall design, VPC-level isolation, the Gateway Firewall default behavior change, and how VCF Operations handles security compliance monitoring.\nA quick terminology note: starting with VCF/NSX 9.0, Broadcom rebranded this firewall and threat-prevention stack as VMware vDefend. What was \u0026ldquo;NSX Distributed Firewall,\u0026rdquo; \u0026ldquo;NSX Gateway Firewall,\u0026rdquo; and \u0026ldquo;NSX Intelligence / NSX ATP\u0026rdquo; are now formally the vDefend Distributed Firewall, vDefend Gateway Firewall, and vDefend Advanced Threat Prevention (ATP). We\u0026rsquo;ll use the familiar DFW / Gateway Firewall shorthand throughout for readability, calling out the vDefend branding where it matters (worth knowing for interviews and documentation searches, since Broadcom TechDocs now files this content under vDefend, not NSX).\nArchitectural Overview The diagram below illustrates the VCF 9 security architecture, showing the defense-in-depth layers from the physical network through workload micro-segmentation:\nvDefend Distributed Firewall (DFW) in VCF 9 The vDefend Distributed Firewall (formerly \u0026ldquo;NSX Distributed Firewall\u0026rdquo;) is the primary security control for east-west (lateral) traffic between workloads in a VCF 9 environment.\nDFW Architecture in VCF 9 VCF 9.0 sets Enhanced Data Path (EDP) Standard as the default host switch mode for new VCF workload domains and new NSX installations, and the vDefend DFW is enforced within that same kernel datapath. In practice:\nDFW policy enforcement runs in the ESX kernel at the vNIC level, not in a user-space proxy EDP Standard delivers superior throughput, packet rate, and lower latency compared to the legacy Standard host switch stack, and supports NSX SPAN and Live Traffic Analysis in the fast path DFW intercepts and filters all VM-to-VM traffic at the source vNIC before it traverses the physical network fabric Workload domains that are upgraded or imported from earlier VCF releases keep running the legacy host switch stack until an administrator explicitly migrates them to EDP Standard Security Groups and Tags VCF 9\u0026rsquo;s DFW uses security groups and VM tags as the primary policy constructs:\nVM Tags: Labels applied to VMs in vCenter or NSX (e.g., \u0026ldquo;Tier=Web\u0026rdquo;, \u0026ldquo;Tier=App\u0026rdquo;, \u0026ldquo;Tier=DB\u0026rdquo;, \u0026ldquo;Env=Production\u0026rdquo;) Security Groups: Dynamic groups that include VMs based on tag membership criteria DFW Rules: Policies that define allowed or denied traffic between security groups, without requiring IP address management Example DFW policy structure for a 3-tier application:\nRule Name Source Destination Service Action Allow-Web-to-App SG-Web SG-App TCP/8080 Allow Allow-App-to-DB SG-App SG-DB TCP/3306 Allow Allow-Web-Inbound SG-External SG-Web TCP/443 Allow Default-Deny Any Any Any Drop DFW and VPC Integration In VCF 9\u0026rsquo;s VPC model, DFW policies apply within and between VPCs:\nWithin a VPC, DFW provides micro-segmentation between VMs on different subnets Between VPCs connected via a Transit Gateway (centralized CTGW or distributed DTGW), DFW policies at the attachment point control inter-VPC traffic VPC-level isolation provides the first layer; DFW provides the workload-level layer within each VPC Context-Aware DFW and Threat Prevention (vDefend ATP) As of VCF/NSX 9.0, Broadcom\u0026rsquo;s advanced security add-on is branded VMware vDefend Advanced Threat Prevention (ATP): this replaces the older \u0026ldquo;NSX Intelligence\u0026rdquo; / \u0026ldquo;NSX ATP\u0026rdquo; naming you may still see referenced in older material. vDefend ATP combines several detection technologies with aggregation, correlation, and context from Network Detection and Response (NDR):\nIDS/IPS: Signature-based intrusion detection and prevention, supported on both the Distributed Firewall and the Gateway Firewall; NSX Manager checks for new intrusion-detection signatures on the cloud every 4 hours by default Network Sandboxing (Malware Prevention): File-based threat analysis for traffic traversing the DFW or Gateway Firewall Network Traffic Analysis (NTA): Behavioral and anomaly detection to catch lateral movement and unusual traffic patterns within the VCF environment Licensing: the IDS/IPS capability requires a Threat Prevention license; Malware Prevention requires the separate Advanced Threat Prevention license: worth confirming which license tier a customer has before promising ATP capabilities in a design vDefend Gateway Firewall: Off by Default for New Gateways One of the significant security posture changes in VCF 9 / NSX 9.0 is that the vDefend Gateway Firewall is disabled by default for new gateways: but only on greenfield VCF 9.0 deployments. This is controlled by the Auto-Activate Gateway Firewall on New Gateways setting (Security → Gateway Firewall → Settings) and can be toggled globally; changing it never affects gateways that are already deployed.\nThe behavior is different for brownfield environments: for an installation upgraded from a previous VCF release to VCF 9.0, Gateway Firewall remains enabled by default for new gateways. If you want the new off-by-default posture on an upgraded environment, you have to explicitly switch the auto-activate setting to Off.\nWhy Was This Changed? In NSX 9.0\u0026rsquo;s VPC-centric design, the primary traffic model has shifted:\nVPC isolation provides tenant separation by default: cross-VPC communication requires explicit Transit Gateway attachment DFW provides micro-segmentation within VPCs The Gateway Firewall was historically used for perimeter-style controls, which are better implemented at the physical edge in the VPC model Disabling the Gateway Firewall by default (for greenfield deployments) improves performance and resource utilization on NSX Edge nodes, which no longer need to maintain stateful connection tracking for all workload traffic by default.\nWhen to Enable Gateway Firewall Enable the Gateway Firewall (on Tier-0 or Tier-1) only for specific use cases requiring stateful inspection at the gateway layer:\nPerimeter security: Stateful firewall enforcement for north-south traffic entering/exiting the VCF environment from external networks Compliance requirements: Specific regulatory frameworks requiring perimeter inspection (not covered by DFW alone) East-west isolation at the VPC boundary: Stateful inspection for inter-VPC traffic traversing the Transit Gateway (supplement to DFW) When enabling Gateway Firewall, define an explicit default-deny rule and allow only required traffic (similar to DFW policy design).\nVPC-Level Isolation: The Foundation of Multi-Tenancy Security VCF 9\u0026rsquo;s VPC model provides isolation-by-default between tenants:\nVMs in VPC-1 cannot communicate with VMs in VPC-2 without explicit Transit Gateway attachment and routing configuration No routing leakage between VPCs: each VPC has its own routing table Public subnets within a VPC allow external connectivity (via NAT or direct external IP), but this does not expose other VPCs VPC Security Best Practices One VPC per tenant or application boundary: Don\u0026rsquo;t share VPCs across security boundaries Minimize public subnet exposure: Place only VMs requiring external access in public subnets; all other VMs should be in private subnets Use Transit Gateway with explicit route filtering: Attach only required VPCs to a Transit Gateway and configure route filters to limit inter-VPC reachability to specific prefixes Apply DFW within VPCs: Never rely solely on VPC isolation; apply DFW policies for micro-segmentation within each VPC VCF Operations: Security Compliance and Monitoring VCF Operations provides centralized security compliance capabilities for all VCF domains:\nSecurity Configuration Compliance VCF Operations integrates with Broadcom\u0026rsquo;s security compliance frameworks to:\nScan ESX host, vCenter, NSX, and vSAN configurations against security baselines (VMware Security Configuration Guides) Report deviations from recommended security configurations Track remediation status for identified compliance findings Certificate Lifecycle Management VCF Operations centralizes certificate lifecycle management for both the VCF management components and the components inside every VCF instance/domain:\nTracks certificates and surfaces expiry alerts for ESX SSL, vCenter machine SSL, NSX Manager (LM and VIP), SDDC Manager SSL, VCF Identity Broker, VCF Operations, VCF Automation, and VCF Operations for Logs/Networks Flags certificates that have expired, are nearing expiration, or have been revoked by the issuing CA, with replace guidance built into the console Supports automatic (non-disruptive) renewal for eligible certificates, avoiding a manual replace-and-restart cycle Can generate certificate signing requests (CSRs) and configure an internal or external Certificate Authority directly from the console Note that SDDC Manager still appears in this certificate inventory in VCF 9.0: its management UI is deprecated and lifecycle workflows have moved to VCF Operations Fleet Management, but the SDDC Manager component itself, and its certificate, are still part of the deployed stack during this transition.\nAudit Logging All VCF Operations actions (domain creation, upgrades, license submissions, user access) are recorded in the VCF Operations audit log:\nProvides a tamper-evident record of all administrative actions Supports SIEM integration for centralized security monitoring Audit log retention is configurable based on compliance requirements VCF Health and Diagnostics (Skyline Parity) Rather than \u0026ldquo;integrating with\u0026rdquo; a separate Skyline Health product, VCF Operations 9.0 natively absorbs Skyline Advisor and Skyline Health Diagnostics functionality as built-in Health and Diagnostics findings:\nVCF Health continuously checks the environment against roughly 62 findings equivalent to former Skyline Advisor signatures, plus around 52 log-based findings equivalent to Skyline Health Diagnostics, re-evaluated automatically on a 4-hour cycle Findings cover known product issues, VMSA-based security exposures, expired/expiring certificates, NTP drift, DNS misconfiguration, and other best-practice deviations, each with root-cause detail and remediation steps Findings and affected objects can be exported to CSV for offline tracking or reporting Remediation for version-related findings typically routes through VCF Operations Fleet Management (applying the relevant update bundle) Security Hardening Checklist for VCF 9 Following the Broadcom VCF 9.0 Design documentation, here is a high-level security hardening checklist:\nNSX / vDefend:\nConfirm the Gateway Firewall auto-activate status for new gateways: off by default only applies to greenfield VCF 9.0 deployments; verify explicitly if this environment was upgraded from a prior VCF release Define DFW default-deny policy for all workload domains Use security groups with VM tags for DFW policy management (avoid IP-based rules where possible) Enable NSX audit logging and forward logs to SIEM vSphere / ESX:\nApply the ESX Security Configuration Guide settings via vSphere Lifecycle Manager desired state Enable vSphere Authentication Proxy for ESX host authentication Restrict management network access to ESX hosts (dedicated management VLAN with firewall controls) vSAN:\nEnable vSAN data-at-rest encryption (if required by compliance framework) Enable vSAN data-in-transit encryption for sensitive workload domains (adds performance overhead; evaluate for each domain) VCF Operations:\nEnable multi-factor authentication (MFA) for VCF Operations admin accounts Integrate VCF Operations with enterprise identity provider (LDAP/AD integration) Review VCF Operations audit logs regularly and forward to SIEM What\u0026rsquo;s Next In the next post, we cover the VCF 9 Identity Broker: how it replaces VMware Identity Manager for fleet-wide single sign-on, its two deployment modes (Embedded and Instance), and what a real migration from vIDM actually involves.\nFurther Reading (Official Broadcom Documentation) vDefend Distributed Firewall Gateway Firewall Settings Virtual Private Cloud in NSX Managing Certificates in VMware Cloud Foundation Overview of NSX IDS/IPS and NSX Malware Prevention (vDefend ATP) ","permalink":"https://virtualizationgurus.pages.dev/posts/vcf9-security-compliance/","summary":"\u003ch2 id=\"introduction\"\u003eIntroduction\u003c/h2\u003e\n\u003cp\u003eSecurity is a foundational concern in any private cloud deployment. VMware Cloud Foundation 9 delivers a layered security architecture that spans from the physical network underlay through the workload networking and compute layers, with VCF Operations providing centralized visibility into the security posture of all domains.\u003c/p\u003e\n\u003cp\u003eIn this post, we explore the VCF 9 security architecture, focusing on the Distributed Firewall design, VPC-level isolation, the Gateway Firewall default behavior change, and how VCF Operations handles security compliance monitoring.\u003c/p\u003e","title":"VCF 9 Security and Compliance: DFW, VPC Isolation, and Hardened Operations"},{"content":"Introduction vSAN\u0026rsquo;s storage architecture choice gets made before a single VM is ever placed \u0026ndash; it\u0026rsquo;s a hardware and cluster-topology decision baked in at workload domain creation. VCF 9.1 supports two architectures side by side, Express Storage Architecture (ESA) and the Original Storage Architecture (OSA), and they claim disks in fundamentally different ways. Picking the wrong one for your hardware, or your operational model, is expensive to unwind later.\nArchitectural Overview OSA: Cache and Capacity, Organized in Disk Groups Under the Original Storage Architecture, every host contributing storage needs at least one cache device and at least one capacity device, organized into one or more disk groups. Cache devices are SAS/SATA SSD or PCIe flash, and for hybrid configurations the cache tier needs to be sized at roughly 10% of anticipated capacity storage. Capacity devices differ by configuration: hybrid clusters use SAS/NL-SAS magnetic disks, all-flash clusters use SAS/SATA SSD or PCIe flash. OSA also requires a storage controller \u0026ndash; a SAS/SATA HBA or RAID controller running in passthrough or RAID-0 mode \u0026ndash; and host memory is sized based on how many disk groups and devices each host carries, typically worked out with the vSAN Sizer tool.\nESA: One Pool, Every Device Contributes With ESA, the cache/capacity split disappears. Every storage device claimed by vSAN contributes to both capacity and performance in a single, unified storage pool per host \u0026ndash; no disk groups to plan around. The hardware requirement is narrower but stricter: each storage pool needs at least one NVMe TLC device, and that device category is separately certified in the Broadcom Compatibility Guide from OSA\u0026rsquo;s cache/capacity certifications. Host memory requirements are also simpler to state \u0026ndash; ESA requires a flat minimum of 128GB per host, rather than the sizing-tool exercise OSA demands.\nESA also brings architectural extras that OSA doesn\u0026rsquo;t have: a Log-Structured File System (LSFS) with inline hardware-assisted compression and a Consistent Log-Aligned Data (CLAD) algorithm that keeps snapshot performance impact near zero. On the operational side, vSAN Proactive Hardware Management (PHM) rounds this out for ESA NVMe drives specifically \u0026ndash; once a supported Hardware Support Manager is registered to vCenter, PHM surfaces OEM predictive-failure signals for a dying device so you can remediate before it takes the storage pool down with it.\nCluster Types: HCI, Compute-Only, and Storage-Only The workload domain creation wizard exposes different cluster shapes depending on which architecture you pick. OSA supports a \u0026ldquo;vSAN HCI\u0026rdquo; cluster (compute and storage combined) or a \u0026ldquo;vSAN Compute Cluster\u0026rdquo; (compute only, which also supports a 2-node configuration without a witness host). ESA supports \u0026ldquo;vSAN HCI\u0026rdquo; as well, plus a disaggregated option \u0026ndash; a storage-only \u0026ldquo;vSAN Storage Cluster\u0026rdquo; that other vSAN ESA or compute clusters can mount as a datastore. This disaggregated model was previously branded vSAN Max in earlier VCF releases before being folded into vSAN ESA as the vSAN Storage Cluster. Data-in-transit encryption, with a configurable rekey interval defaulting to 1440 minutes, is available for both architectures.\nPolicy Mechanics That Apply to Both Whichever architecture you land on, vSAN\u0026rsquo;s fault-tolerance policies still drive host-count minimums: RAID-1 (FTT=1) needs at least 3 hosts, RAID-5 (FTT=1) needs at least 4 hosts and is more space-efficient than RAID-1, and RAID-6 (FTT=2) needs at least 6 hosts. These minimums shape whether ESA\u0026rsquo;s disaggregated storage-only cluster or a more traditional HCI shape is the better fit for a given domain\u0026rsquo;s host count.\nChoosing Between Them OSA remains the right call where you\u0026rsquo;re working with existing hybrid or mixed-media hardware, or where a full NVMe refresh isn\u0026rsquo;t in the current budget cycle. ESA is the better fit for new NVMe-based deployments where the simplified single-pool model, inline compression, and near-zero-impact snapshots outweigh the flat 128GB memory tax and stricter device certification requirements.\nWhat\u0026rsquo;s Next Next in this series: VCF 9 Security and Compliance: DFW, VPC Isolation, and Hardened Operations.\nFurther Reading (Official Broadcom Documentation) vSAN Deployment, Administration, and Monitoring vSAN Concepts Hardware Requirements for vSAN Managing Proactive Hardware Create a New Workload Domain ","permalink":"https://virtualizationgurus.pages.dev/posts/vcf9-vsan-esa-vs-osa/","summary":"\u003ch2 id=\"introduction\"\u003eIntroduction\u003c/h2\u003e\n\u003cp\u003evSAN\u0026rsquo;s storage architecture choice gets made before a single VM is ever placed \u0026ndash; it\u0026rsquo;s a hardware and cluster-topology decision baked in at workload domain creation. VCF 9.1 supports two architectures side by side, Express Storage Architecture (ESA) and the Original Storage Architecture (OSA), and they claim disks in fundamentally different ways. Picking the wrong one for your hardware, or your operational model, is expensive to unwind later.\u003c/p\u003e\n\u003ch2 id=\"architectural-overview\"\u003eArchitectural Overview\u003c/h2\u003e\n\u003cdiv class=\"diagram-embed\"\u003e\n  \u003cobject type=\"image/svg+xml\" data=\"/images/diagrams/vcf9-vsan-esa-vs-osa.svg\"\u003e\u003c/object\u003e\n\u003c/div\u003e\n\u003ch2 id=\"osa-cache-and-capacity-organized-in-disk-groups\"\u003eOSA: Cache and Capacity, Organized in Disk Groups\u003c/h2\u003e\n\u003cp\u003eUnder the Original Storage Architecture, every host contributing storage needs at least one cache device and at least one capacity device, organized into one or more disk groups. Cache devices are SAS/SATA SSD or PCIe flash, and for hybrid configurations the cache tier needs to be sized at roughly 10% of anticipated capacity storage. Capacity devices differ by configuration: hybrid clusters use SAS/NL-SAS magnetic disks, all-flash clusters use SAS/SATA SSD or PCIe flash. OSA also requires a storage controller \u0026ndash; a SAS/SATA HBA or RAID controller running in passthrough or RAID-0 mode \u0026ndash; and host memory is sized based on how many disk groups and devices each host carries, typically worked out with the vSAN Sizer tool.\u003c/p\u003e","title":"vSAN ESA vs OSA: Storage Architecture Decisions"},{"content":"Introduction Every workload domain eventually needs to talk to the outside world, and in NSX that conversation happens at the edge. The NSX Edge cluster is where policy meets physical: it hosts the Tier-0 gateway that peers with your physical network, terminates VPN tunnels, and enforces the firewall rules that decide what\u0026rsquo;s allowed to cross the north-south boundary. Get the Edge cluster\u0026rsquo;s HA design wrong and you inherit asymmetric routing, dropped stateful sessions, or a firewall that silently fails open on a node switchover. This post breaks down the Tier-0/Tier-1 split, the HA modes that govern them, and how VPN and firewall services layer on top.\nArchitectural Overview Tier-0 Gateway: The Fleet\u0026rsquo;s Front Door The Tier-0 gateway is the top-tier gateway in the NSX topology \u0026ndash; it\u0026rsquo;s the only place where NSX hands off to your physical network. Southbound, it connects to one or more Tier-1 gateways or directly to segments; northbound, it peers with physical routers. A given NSX Edge node supports a single Tier-0 gateway, which is why Edge cluster sizing is really a Tier-0 sizing exercise. The default T0-T1 transit subnet is 100.64.0.0/16, with an internal transit subnet of 169.254.0.0/24 used for intra-tier links \u0026ndash; worth knowing before you start allocating overlapping RFC 1918 ranges elsewhere in the fabric.\nTier-1 Gateway: Segment-Level Routing Tier-1 gateways sit below Tier-0 and own the downlinks to your segments. They have no direct physical uplink of their own \u0026ndash; everything they don\u0026rsquo;t know how to route gets pushed up to Tier-0. Tier-1 supports route advertisement back to Tier-0, static routes, and even recursive static routes for more deliberate control over what gets learned where. In practice, Tier-1 is where you scope tenancy or application boundaries, while Tier-0 stays focused on the fleet-wide north-south path.\nHigh Availability: Active-Active vs Active-Standby Both Tier-0 and Tier-1 gateways support two HA modes. Active-active is the default: both Edge nodes forward traffic simultaneously and load-balance across the pair, which is great for throughput but has a hard limitation \u0026ndash; stateful services (SNAT, DNAT, load balancing, stateful firewall, and VPN) are not supported in this mode, because there\u0026rsquo;s no guarantee a return packet lands on the node that saw the original flow.\nActive-standby elects a single active node and holds the second in reserve, with all stateful services allowed. Failover behavior itself is configurable: preemptive mode fails back to the original active node once it recovers, non-preemptive leaves the newly active node in place until the next failure. If your design needs NAT, a load balancer, or VPN anywhere on that gateway, active-standby isn\u0026rsquo;t optional \u0026ndash; it\u0026rsquo;s a prerequisite. An HA VIP is also available in active-standby mode, letting a single external interface failure be absorbed without a routing protocol re-convergence, though it\u0026rsquo;s intended to work with static routing rather than BGP.\nVPN Connectivity: IPSec and L2 VPN NSX Edge supports two VPN types for extending connectivity beyond the fleet. IPSec VPN builds site-to-site tunnels using ESP in tunnel mode (IP protocol 50) with IKE negotiation over UDP 500, or UDP 4500 when NAT-traversal is involved. You can run it policy-based or route-based, authenticating with a pre-shared key or certificates using SHA256/384/512 with RSA. L2 VPN, by contrast, extends a Layer 2 segment across data centers \u0026ndash; useful for migrations or stretched clusters rather than routed connectivity.\nBoth VPN types share the same HA constraint as other stateful services: IPSec VPN requires the Tier-0, Tier-0 VRF, or Tier-1 gateway to be running in active-standby mode, and on failover the VPN state is synchronized to the standby node so tunnels don\u0026rsquo;t need to renegotiate from scratch. One notable gap: an IPSec VPN session isn\u0026rsquo;t supported directly between a parent Tier-0 and an attached Tier-0 VRF.\nNorth-South Firewall with vDefend Firewalling in NSX comes in two form factors. The Distributed Firewall runs per-vSphere-workload and handles east-west traffic between VMs regardless of which segment they sit on. The Gateway Firewall is the north-south counterpart, deployed on a vSphere host as either a VM or a bare-metal ISO appliance, to inspect traffic crossing the Tier-0/Tier-1 boundary.\nBoth are part of the broader vDefend Firewall with Advanced Threat Prevention, a software-defined L2-7 stateful firewall stack that layers IDS/IPS, network sandboxing, and network traffic analysis with NDR-style aggregation and correlation on top of basic packet filtering. For the Edge cluster specifically, the Gateway Firewall is your enforcement point for anything leaving or entering the fleet \u0026ndash; and like NAT and VPN, its statefulness depends on the gateway running in active-standby HA.\nDeploying the Edge Cluster NSX Edge nodes can be installed through the NSX UI, the vSphere Client, or the CLI using the OVF tool, and are then grouped into an Edge cluster via the \u0026ldquo;Create an NSX Edge Cluster and Add Edge Nodes\u0026rdquo; workflow. This matters for workload domain connectivity models too: a domain using Centralized Connectivity only becomes VPC-ready after you\u0026rsquo;ve separately deployed an Edge cluster with a Tier-0 in active-standby mode, whereas Distributed Connectivity is VPC-ready immediately but still needs a Virtual Network Appliance cluster if you want stateful NAT or load balancing.\nWhat\u0026rsquo;s Next Next in this series: vSAN ESA vs OSA: Storage Architecture Decisions.\nFurther Reading (Official Broadcom Documentation) NSX Advanced Network Management Tier-0 Gateways Add an NSX Tier-0 Gateway Tier-1 Gateways Installing NSX Edge Understanding IPSec VPN vDefend Firewall with Advanced Threat Prevention ","permalink":"https://virtualizationgurus.pages.dev/posts/vcf9-nsx-edge-cluster-deep-dive/","summary":"\u003ch2 id=\"introduction\"\u003eIntroduction\u003c/h2\u003e\n\u003cp\u003eEvery workload domain eventually needs to talk to the outside world, and in NSX that conversation happens at the edge. The NSX Edge cluster is where policy meets physical: it hosts the Tier-0 gateway that peers with your physical network, terminates VPN tunnels, and enforces the firewall rules that decide what\u0026rsquo;s allowed to cross the north-south boundary. Get the Edge cluster\u0026rsquo;s HA design wrong and you inherit asymmetric routing, dropped stateful sessions, or a firewall that silently fails open on a node switchover. This post breaks down the Tier-0/Tier-1 split, the HA modes that govern them, and how VPN and firewall services layer on top.\u003c/p\u003e","title":"NSX Edge Cluster Deep Dive: Tier-0/Tier-1 Gateways, VPN, and North-South Firewall Design"},{"content":"Introduction Every VI workload domain you stand up in VCF asks the same networking question: does it join an existing NSX Manager, or get its own? The Networking page of the workload domain creation wizard boils this down to two buttons \u0026ndash; \u0026ldquo;Join Existing NSX Manager Instance\u0026rdquo; and \u0026ldquo;Create New NSX Manager Instance\u0026rdquo; \u0026ndash; but the operational consequences run much deeper than a single click. This is a per-domain decision, not a fleet-wide one, and a VCF instance scaling toward its 25-domain ceiling will likely end up with a mix of both.\nArchitectural Overview Shared NSX: Join an Existing Instance Choosing \u0026ldquo;Join Existing NSX Manager Instance\u0026rdquo; surfaces only version-compatible NSX Managers already in the fleet, and the new workload domain\u0026rsquo;s vCenter is automatically registered as a compute manager against it. VCF 9.1 adds a notable option here: a workload domain can now select the management domain\u0026rsquo;s own NSX Manager as its shared instance, rather than requiring a separate dedicated NSX deployment elsewhere. Shared NSX means a smaller overall footprint and one fewer NSX Manager cluster to patch and monitor, at the cost of a larger blast radius \u0026ndash; an NSX-side incident or maintenance window now touches every domain attached to that instance, and scaling is constrained by whatever headroom the shared instance has left. That headroom has a hard ceiling: an NSX Manager deployed at Large or Extra Large appliance size supports up to 16 attached vCenters (compute managers), while a Medium-sized NSX Manager supports only 2.\nDedicated NSX: Create a New Instance Choosing \u0026ldquo;Create New NSX Manager Instance\u0026rdquo; hands you a full deployment wizard: a Deployment Size of Simple (single-node) or High-Availability (three-node, recommended for anything beyond a lab), an Appliance Size of Medium, Large, or Extra Large, and FQDN requirements for each node plus a cluster-level FQDN. A dedicated NSX Manager gives the domain independent availability, independent scaling, and a lifecycle that moves on its own schedule \u0026ndash; nothing else. The tradeoff is a materially higher footprint: another HA cluster to size, deploy, and keep patched.\nVersion Compatibility Matters More With Shared NSX Because a shared NSX Manager serves multiple domains, its version has to stay compatible with every attached vCenter \u0026ndash; and that compatibility isn\u0026rsquo;t binary. Broadcom\u0026rsquo;s workload domain creation documentation lays it out as a state table:\nShared NSX Version Workload Domain vCenter Version Resulting State NSX 9.1 vCenter 9.1 Full features NSX 9.1 vCenter 9.0.x Pinned state (limited to vCenter 9.0 capabilities) NSX 9.1 vCenter 8.0.x Pinned state (limited to vCenter 8.0 capabilities) NSX 9.0.x vCenter 9.0.x Full features NSX 4.2.x vCenter 8.0.x Legacy mode Even after NSX Manager itself is upgraded, new NSX 9.1 features stay inactive until every ESXi host transport node across every attached workload domain is running the matching version. A dedicated NSX Manager sidesteps this entirely \u0026ndash; its version story only has to make sense for one domain.\nThe Tradeoff, Side by Side Dimension Dedicated NSX Shared NSX Footprint Higher \u0026ndash; own HA cluster Lower \u0026ndash; reuse existing cluster Availability Independent of other domains Shared blast radius Scalability Scales independently Constrained by shared instance headroom (16 vCenters max on Large/XL, 2 on Medium) Lifecycle Upgrades on its own schedule Version compatibility can pin features A Decision That\u0026rsquo;s Independent of VPC Connectivity Whichever NSX topology you pick, it\u0026rsquo;s a separate decision from how the domain reaches the outside world \u0026ndash; VCF 9\u0026rsquo;s networking consumption model is built on Virtual Private Clouds (VPCs) sitting behind Transit Gateways, and each workload domain chooses how that Transit Gateway connects out. Centralized Connectivity keeps north-south traffic funneled through a dedicated Edge cluster, and only becomes VPC-ready once you configure that Edge cluster with an active-standby Tier-0 gateway after domain creation. Distributed Connectivity is VPC-ready as soon as the domain is created, but needs a Virtual Network Appliance cluster deployed afterward if you want the Distributed Transit Gateway to provide stateful NAT or load balancing. Shared-vs-dedicated NSX and centralized-vs-distributed connectivity are two independent switches you flip for every domain \u0026ndash; not one combined choice.\nWhat\u0026rsquo;s Next Next in this series: NSX Edge Cluster Deep Dive: Tier-0/Tier-1 Gateways, VPN, and North-South Firewall Design.\nFurther Reading (Official Broadcom Documentation) Create a New Workload Domain VMware Cloud Foundation Architecture Models Managing Virtual Private Clouds in vCenter VMware Configuration Maximums for VCF 9 ","permalink":"https://virtualizationgurus.pages.dev/posts/vcf9-vi-workload-domains-shared-vs-dedicated-nsx/","summary":"\u003ch2 id=\"introduction\"\u003eIntroduction\u003c/h2\u003e\n\u003cp\u003eEvery VI workload domain you stand up in VCF asks the same networking question: does it join an existing NSX Manager, or get its own? The Networking page of the workload domain creation wizard boils this down to two buttons \u0026ndash; \u0026ldquo;Join Existing NSX Manager Instance\u0026rdquo; and \u0026ldquo;Create New NSX Manager Instance\u0026rdquo; \u0026ndash; but the operational consequences run much deeper than a single click. This is a per-domain decision, not a fleet-wide one, and a VCF instance scaling toward its 25-domain ceiling will likely end up with a mix of both.\u003c/p\u003e","title":"VI Workload Domains: Shared vs Dedicated NSX"},{"content":"Introduction Every VI workload domain in a VCF instance got there one of two ways: it was built from scratch through the workload domain wizard, or it was an existing vCenter environment that VCF Operations absorbed as-is. These aren\u0026rsquo;t just two UI paths to the same result \u0026ndash; they carry different prerequisites, different day-one side effects, and different long-term upgrade constraints. Picking the wrong one for your situation doesn\u0026rsquo;t just cost time in the wizard; it can lock a workload domain out of upgrade paths for its entire lifecycle.\nTwo Paths, One Fleet Greenfield creates a new workload domain from hosts that have never run a production vCenter \u0026ndash; commissioned ESX hosts, a chosen principal storage type, and a new or shared NSX Manager instance. VCF Operations builds the vCenter, configures the cluster, and enrolls the domain into fleet management from the moment it exists.\nImport takes a vCenter instance you already run \u0026ndash; with or without NSX \u0026ndash; and brings its entire inventory into the VCF instance as a workload domain. VCF Operations calls this \u0026ldquo;Import\u0026rdquo;; the same mechanism invoked through the VCF Installer with a JSON specification file is called \u0026ldquo;Converge.\u0026rdquo; Same outcome, different entry point depending on whether you\u0026rsquo;re driving it through the UI or automating it.\nGreenfield: Build It Fresh The wizard-driven path assumes a clean slate. Hosts must already be commissioned with the target principal storage type, and if the management domain hosts were imaged with an express patch, the new workload domain\u0026rsquo;s hosts need that same patch applied before they\u0026rsquo;ll pass validation. For NSX, you choose per domain: deploy a new NSX Manager instance, or join an existing shared instance from another workload domain \u0026ndash; though a shared instance deployed in an IPv4-only domain forces the new workload domain onto IPv4 as well, even if it was otherwise planned for dual-stack.\nvVols as a principal storage type still works in VCF 9, but it ships deprecated \u0026ndash; if you\u0026rsquo;re standing up new infrastructure today, that\u0026rsquo;s a call worth making deliberately rather than by default.\nImport: Bring What You Already Have Import has real teeth around version alignment and configuration shape, because it\u0026rsquo;s absorbing infrastructure VCF didn\u0026rsquo;t build:\nMinimum versions: vCenter 8.0 Update 1+, ESX 8.0 Update 1+, and if NSX Manager is present, a three-node cluster on 4.1.0.2 or later. No NSX Manager present? VCF deploys one during the import, at the latest version compatible with that vCenter. All-or-nothing at the vCenter level: every cluster managed by the imported vCenter comes in together. There\u0026rsquo;s no picking a subset \u0026ndash; if even one cluster in that vCenter doesn\u0026rsquo;t meet the requirements, the whole import blocks. Storage priority is fixed: when a cluster has multiple datastore types, VCF picks the primary automatically \u0026ndash; vSAN first, then NFS v3, then VMFS, then NFS 4.1, then iSCSI, with vVols last and not recommended. A JSON convergence spec can override this with existingDatastoreName, but the UI-driven import can\u0026rsquo;t. What\u0026rsquo;s explicitly unsupported: Dell VxRail-managed clusters, vSphere Configuration Profiles, vCenter instances still running Enhanced Linked Mode (ELM must be broken first), clusters on baseline-based lifecycle management, and clusters with DRS set to manual or partially automated rather than fully automated. Two prerequisites are easy to miss because they sit outside the wizard entirely: SSH has to be enabled on the existing vCenter appliance before you start, and every ESX host in scope needs to be using FQDNs in vSphere inventory rather than short names \u0026ndash; both are precheck failures if skipped.\nWhat Import Changes on Day One The import workflow doesn\u0026rsquo;t just register the vCenter \u0026ndash; it actively reconfigures the network layer underneath it. NSX gets activated on every Distributed Virtual Port Group in the imported cluster, which in turn activates the Distributed Firewall on those same DVPGs. The default DFW ruleset allows all Layer 2 and Layer 3 traffic, so nothing breaks immediately \u0026ndash; but existing workloads are now sitting behind a firewall layer that wasn\u0026rsquo;t there an hour before, and someone now owns writing real rules for it. You can deactivate NSX on DVPGs afterward through the NSX UI or the Transport Node Collection API if that\u0026rsquo;s not what you wanted on day one.\nThe Forward-Only Upgrade Trap This is the constraint that makes \u0026ldquo;just import it and sort out versions later\u0026rdquo; a bad plan. VCF enforces forward-only upgrade validation: you can only upgrade to a target version released after your current version\u0026rsquo;s release date. When you import components, your workload domain\u0026rsquo;s effective baseline becomes the earliest release date among everything you brought in. If the NSX instance you imported was released after a given VCF minor or maintenance version, that workload domain is permanently blocked from upgrading to that specific version \u0026ndash; not delayed, blocked. Checking the VMware Interoperability Matrix before importing, not after, is the difference between a clean path forward and a workload domain stuck a step behind the rest of the fleet.\nChoosing Between Them: A Checklist No existing vSphere estate to bring in? Greenfield \u0026ndash; there\u0026rsquo;s nothing to import, and the wizard is simpler than staging a migration. Existing production vCenter with workloads you can\u0026rsquo;t re-platform? Import \u0026ndash; greenfield would mean rebuilding and migrating, and import exists specifically to skip that. Existing vCenter using ELM, VxRail, baselines, or manual DRS? Remediate those first; none of them pass import prechecks as-is. Mixed NSX versions across what you\u0026rsquo;re importing? Check the Interoperability Matrix against your target VCF version before you start \u0026ndash; not after the import completes. Either way: confirm SSH is enabled and every host is on FQDN before you open the wizard. Both are the most common precheck failures, and both take five minutes to fix in advance. What\u0026rsquo;s Next Next in this series: VI Workload Domains \u0026ndash; Shared vs Dedicated NSX.\nFurther Reading (Official Broadcom Documentation) Import an Existing vCenter to Create a Workload Domain Create a New Workload Domain Create a Workload Domain by Using the VCF Operations API VMware Interoperability Matrix ","permalink":"https://virtualizationgurus.pages.dev/posts/vcf9-workload-domain-creation-greenfield-vs-import/","summary":"\u003ch2 id=\"introduction\"\u003eIntroduction\u003c/h2\u003e\n\u003cp\u003eEvery VI workload domain in a VCF instance got there one of two ways: it was built from scratch through the workload domain wizard, or it was an existing vCenter environment that VCF Operations absorbed as-is. These aren\u0026rsquo;t just two UI paths to the same result \u0026ndash; they carry different prerequisites, different day-one side effects, and different long-term upgrade constraints. Picking the wrong one for your situation doesn\u0026rsquo;t just cost time in the wizard; it can lock a workload domain out of upgrade paths for its entire lifecycle.\u003c/p\u003e","title":"Workload Domain Creation: Greenfield vs Import Existing vCenter"},{"content":"Introduction Every workload domain you\u0026rsquo;ll ever stand up in VCF inherits whatever the physical network underneath it can actually deliver. By the time you\u0026rsquo;re in the workload domain creation wizard picking a VDS profile, the rack layout, the ToR trunk configuration, and the MTU on every switch port in between have already decided how much of that wizard\u0026rsquo;s promise you can keep. This post works bottom-up: what the physical fabric needs to provide, how vSphere Distributed Switches carve that fabric into traffic types, and how NSX Edge nodes hand off to it over BGP.\nArchitectural Overview Rack-Level Fabric: Where Bottlenecks Actually Come From Broadcom\u0026rsquo;s data center network design guidance frames physical fault tolerance in layers \u0026ndash; sites, availability zones, witness sites, and racks \u0026ndash; and racks are where most day-to-day network design decisions actually live. Two hosts on the same ToR switch typically see no fabric-induced bottleneck; hosts on different switches, whether in the same rack or across racks, can be constrained by over-subscription on the links between switches or through the spine. That\u0026rsquo;s not a theoretical concern \u0026ndash; it\u0026rsquo;s the reason \u0026ldquo;which rack is this host in\u0026rdquo; is a real design input, not just a facilities detail.\nToR Switch Requirements: Trunks, Uplinks, and MTU Each ToR switch pair is configured to carry every VLAN a host needs over an 802.1Q trunk, and each ESX host connects to that ToR pair with two or more redundant ports, recommended at 25GbE or higher. MTU is the detail that\u0026rsquo;s easy to under-provision: NSX Geneve encapsulation sets the don\u0026rsquo;t-fragment bit on overlay traffic between host and Edge TEPs, so the whole path \u0026ndash; host physical NICs, ToR ports, spine, and every hop in between \u0026ndash; has to carry it without fragmenting. The minimum required MTU is 1600, but 1700 is the recommended figure, sized to leave headroom for future Geneve header growth rather than the bare minimum.\nVDS Separation: Five Profiles for Splitting Traffic The workload domain creation wizard doesn\u0026rsquo;t make you hand-design a VDS from nothing \u0026ndash; it offers preconfigured profiles, each mapping traffic types onto one or more distributed switches with dedicated physical NICs:\nProfile What it separates Default Single VDS, unified fabric for every traffic type Storage Traffic Separation Two VDS: one dedicated to storage, one for everything else NSX Traffic Separation Two VDS: one dedicated to NSX (overlay) traffic, one for everything else Storage and NSX Traffic Separation Three VDS: storage, NSX, and everything else, each isolated Custom Switch Configuration Build your own VDS layout, copying from a profile or starting fresh Whichever profile you pick, some traffic types \u0026ndash; Management, vMotion, vSAN, and NSX \u0026ndash; can only be configured once per cluster, while NSX and Public traffic types can span multiple VDS if your custom configuration calls for it. Uplinks per switch are either individual VDS Uplinks (software load-balanced via LBT or virtual-port-ID hashing, no special physical switch configuration required) or a VDS LAG bundling multiple physical NICs under LACP \u0026ndash; which does require matching LACP configuration on the ToR side, active or passive.\nBGP Uplinks: Getting NSX Edge Talking to the Physical Fabric Once traffic reaches an NSX Edge node, the Tier-0 gateway\u0026rsquo;s uplinks are where the overlay world hands off to the physical one, and that handoff is BGP in the overwhelming majority of VCF designs. eBGP neighbors must sit directly connected in the same subnet as the Tier-0 uplink \u0026ndash; multi-hop eBGP is supported but requires explicit configuration \u0026ndash; and both a local AS number and the physical router\u0026rsquo;s remote AS number are required before the session comes up. Active-active Tier-0 gateways default to ASN 65000 and enable BGP automatically; active-standby gateways have no default ASN and start with BGP disabled.\nFor load distribution across multiple uplinks, ECMP hashes on the standard 5-tuple (protocol, source/destination address, source/destination port). If your physical topology advertises routes with the same AS-path length but different AS-path content \u0026ndash; common with certain multi-uplink or multi-router designs \u0026ndash; ECMP alone won\u0026rsquo;t use all of them; Multipath Relax has to be enabled on top of ECMP to allow load-sharing across paths that only differ in AS-path values. On the resiliency side, BFD intervals as low as 500ms are supported for Edge nodes running as VMs, and MD5 authentication is available on the BGP session itself if your network team requires it.\nPutting It Together: A Design Checklist None of these layers are independent decisions. A VDS profile that isolates NSX traffic onto its own switch is only as good as the ToR trunk and MTU underneath it; a beautifully tuned BGP uplink is only as resilient as the rack fault-tolerance model it\u0026rsquo;s peering out of. Before a workload domain goes live, it\u0026rsquo;s worth walking the stack top to bottom: confirm the VDS profile matches the isolation the workload actually needs, confirm every switch in the Geneve path carries at least 1700 MTU, confirm ToR uplinks are redundant at 25GbE or better, and confirm the Tier-0 BGP configuration \u0026ndash; AS numbers, ECMP, Multipath Relax if needed \u0026ndash; matches what the physical network team actually configured on their side.\nWhat\u0026rsquo;s Next Next in this series: Workload Domain Creation: Greenfield vs Import Existing vCenter.\nFurther Reading (Official Broadcom Documentation) Data Center Network Requirements Create a New Workload Domain Configure BGP Tier-0 Gateways ","permalink":"https://virtualizationgurus.pages.dev/posts/vcf9-physical-network-design/","summary":"\u003ch2 id=\"introduction\"\u003eIntroduction\u003c/h2\u003e\n\u003cp\u003eEvery workload domain you\u0026rsquo;ll ever stand up in VCF inherits whatever the physical network underneath it can actually deliver. By the time you\u0026rsquo;re in the workload domain creation wizard picking a VDS profile, the rack layout, the ToR trunk configuration, and the MTU on every switch port in between have already decided how much of that wizard\u0026rsquo;s promise you can keep. This post works bottom-up: what the physical fabric needs to provide, how vSphere Distributed Switches carve that fabric into traffic types, and how NSX Edge nodes hand off to it over BGP.\u003c/p\u003e","title":"Physical Network Design: VDS Separation, ToR Switches, and BGP Uplinks"},{"content":"Introduction Every VCF Instance starts with a management domain \u0026ndash; it\u0026rsquo;s the first workload domain deployed, and it never stops being special. This post breaks down exactly what components live inside a management domain per Broadcom\u0026rsquo;s official documentation, why the very first one in a fleet carries extra weight, and the shared-vs-dedicated NSX tradeoff that shows up in almost every design conversation.\nWhat Every VCF Domain Contains Per the VCF Taxonomy documentation, any VCF domain \u0026ndash; management or workload \u0026ndash; is built from the same base components: one vCenter instance, one or more vSphere clusters with vSphere HA and DRS enabled, at least one vSphere Distributed Switch per cluster plus NSX segments for workload traffic, a dedicated or shared NSX Manager instance, optional NSX Edge or Virtual Network Appliance clusters added after domain creation, and one or more shared storage allocations.\nThe Management Domain\u0026rsquo;s Extra Responsibilities What makes the management domain different is what else it carries on top of that baseline, according to the VCF Domain Models documentation:\n-\u0026gt; Deployed first, always \u0026ndash; the management domain is the first workload domain deployed in a VCF Instance, created during initial deployment or convergence by VCF Installer, before any workload domain exists.\n-\u0026gt; Houses the fleet\u0026rsquo;s control plane \u0026ndash; in the first VCF Instance of a fleet, the management domain contains the VCF fleet management components in addition to its own instance-level components. Additional management domains (for additional Instances in the same fleet) carry instance-level components only, not the fleet-level ones.\n-\u0026gt; Runs SDDC Manager and a separate fleet management appliance \u0026ndash; the management domain includes the SDDC Manager appliance alongside a distinct fleet management appliance. It\u0026rsquo;s worth keeping these two straight: SDDC Manager is still installed and still listed as a management-domain component, but its UI and lifecycle-management workflows are deprecated as of VCF 9.0 and have moved into VCF Operations Fleet Management. The fleet management appliance is the component that actually extends the VCF Operations instance with the infrastructure-automation workflows for the fleet \u0026ndash; not SDDC Manager itself.\n-\u0026gt; Hosts a dedicated License Server \u0026ndash; as of VCF 9.1, licensing runs in its own VCF License Server appliance (auto-deployed alongside VCF Operations during install or upgrade) rather than inside VCF Operations, as it did previously.\n-\u0026gt; Can still run business workloads \u0026ndash; the management domain isn\u0026rsquo;t purely infrastructure; per Broadcom\u0026rsquo;s Domain Models documentation, it\u0026rsquo;s explicitly permitted to host general workloads and NSX Edge nodes alongside its management role, though most designs keep it dedicated to management functions for lifecycle and resource isolation reasons.\nShared vs. Dedicated NSX: A Real Tradeoff Workload domains \u0026ndash; and by extension, the NSX relationship they have with the management domain \u0026ndash; can either get a dedicated NSX Manager cluster or share one. Broadcom\u0026rsquo;s documentation lays out the tradeoff plainly:\nAttribute Dedicated NSX per Domain Shared NSX Across Domains Footprint Higher \u0026ndash; one NSX Manager cluster per domain Lower \u0026ndash; one cluster serves multiple domains Availability Independent control plane per domain Larger blast radius if the shared cluster fails Scalability Each domain scales independently Constrained by the shared instance\u0026rsquo;s limits Lifecycle Upgrades/patches applied independently Upgrades/patches affect every domain sharing it A VCF Instance can scale up to 25 total domains \u0026ndash; one management domain plus as many as 24 VI workload domains \u0026ndash; and each workload domain individually chooses dedicated or shared NSX; it isn\u0026rsquo;t a fleet-wide, all-or-nothing setting. A single shared NSX Manager has its own ceiling too: an Extra Large or Large form factor supports up to 16 attached vCenters/Compute Managers, while a Medium form factor supports only 2.\nArchitecture at a Glance A Practical Tip Size the management domain with room to grow \u0026ndash; Broadcom\u0026rsquo;s own guidance notes that you may need to expand its resources over time to accommodate additional workload domains and management components, so treat initial sizing as a floor, not a ceiling.\nWhat\u0026rsquo;s Next Next in this series: Physical Network Design: VDS Separation, ToR, and BGP Uplinks.\nFurther Reading (Official Broadcom Documentation) VCF Taxonomy VCF Domain Models Deploy VCF Management Services and License Server as Part of VCF Upgrade to 9.1 Licensing Overview VMware Configuration Maximums for VCF 9 ","permalink":"https://virtualizationgurus.pages.dev/posts/vcf9-management-domain-anatomy/","summary":"\u003ch2 id=\"introduction\"\u003eIntroduction\u003c/h2\u003e\n\u003cp\u003eEvery VCF Instance starts with a management domain \u0026ndash; it\u0026rsquo;s the first workload domain deployed, and it never stops being special. This post breaks down exactly what components live inside a management domain per Broadcom\u0026rsquo;s official documentation, why the very first one in a fleet carries extra weight, and the shared-vs-dedicated NSX tradeoff that shows up in almost every design conversation.\u003c/p\u003e\n\u003ch2 id=\"what-every-vcf-domain-contains\"\u003eWhat Every VCF Domain Contains\u003c/h2\u003e\n\u003cp\u003ePer the VCF Taxonomy documentation, any VCF domain \u0026ndash; management or workload \u0026ndash; is built from the same base components: one vCenter instance, one or more vSphere clusters with vSphere HA and DRS enabled, at least one vSphere Distributed Switch per cluster plus NSX segments for workload traffic, a dedicated or shared NSX Manager instance, optional NSX Edge or Virtual Network Appliance clusters added after domain creation, and one or more shared storage allocations.\u003c/p\u003e","title":"VCF 9 Management Domain Anatomy: What Actually Runs Inside"},{"content":"Introduction Once you\u0026rsquo;ve internalized \u0026ldquo;one Fleet across every site,\u0026rdquo; the next question is practical: how do you actually design that fleet across a headquarters site, a DR site, and edge or sovereign-cloud locations? This post breaks down the core VCF Instance/Fleet constructs and the four fleet deployment designs Broadcom documents for VCF 9.1, and maps them to the multi-site patterns operators actually build.\nThe Three Constructs, Revisited Broadcom\u0026rsquo;s VCF Taxonomy defines three nested constructs worth keeping straight:\n-\u0026gt; VCF Instance \u0026ndash; the compute, storage, and networking infrastructure that runs actual workloads: a management domain plus, optionally, workload domains.\n-\u0026gt; VCF fleet \u0026ndash; the environment managed by a single set of fleet-level components (VCF Operations and VCF Automation), which can span one or more VCF Instances and even standalone vCenter deployments not otherwise part of an Instance.\n-\u0026gt; VCF private cloud \u0026ndash; the highest-level construct, containing one or more VCF fleets.\nA single Instance can be headquarters-sized on its own. Multiple Instances under one fleet are what let a HQ, a DR site, and edge locations share centralized operations and automation instead of running as isolated silos.\nOne placement detail matters when you\u0026rsquo;re actually drawing this topology: per Broadcom\u0026rsquo;s taxonomy, the fleet-level components (VCF Operations and VCF Automation) live in the management domain of whichever VCF Instance is first in the fleet. So when you\u0026rsquo;re designing an HQ + DR + Edge fleet, the choice of which Instance goes first effectively decides where your single control plane physically sits.\nFour Fleet Deployment Designs Broadcom\u0026rsquo;s VCF Fleet Deployment Models documentation lays out four designs, each building on the previous one, that map naturally onto common multi-site topologies:\n-\u0026gt; Basic Design \u0026ndash; a single VCF fleet in one availability zone or region. This is the starting point: one management domain instance centrally controlling one or more workload domain instances, without cross-site protection. A natural fit for a standalone HQ deployment.\n-\u0026gt; Site High Availability (Across Zones) Design \u0026ndash; adds fault domains that spread resources across two availability zones with vSphere HA, vSAN stretched clustering, and NSX stretched segments, giving active-active protection against a single zone outage.\n-\u0026gt; Disaster Recovery (Across Regions) Design \u0026ndash; builds on the Basic Design and adds a second Instance in a separate region, using VMware Live Recovery for failover and failback \u0026ndash; the closest match to a dedicated HQ + DR pairing.\n-\u0026gt; Fault Domains and Disaster Recovery Design \u0026ndash; combines both: fault-domain high availability within a site plus cross-region disaster recovery, for organizations that need protection from both localized hardware failures and full site-level disasters.\nEdge locations, unlike the general HQ/DR patterns above, do get their own named model: VCF Edge (formerly called Remote Clusters), a purpose-built configuration for edge deployments with its own sizing rules \u0026ndash; a minimum of 10 sites, at least 8 CPU cores per host, a cap of 256 CPU cores per site, and a requirement that Edge hosts sit in a physically distinct location (a separate rack or switch) from your data center workloads unless you have strict segregation controls in place. VCF Operations is a mandatory licensing requirement for any VCF Edge deployment. Sovereign-cloud requirements, by contrast, aren\u0026rsquo;t tied to a single named topology \u0026ndash; they\u0026rsquo;re typically met by combining one of the four fleet designs above with in-country data residency and recovery controls (for example, vSAN-based ransomware and data recovery) rather than a distinct named architecture.\nArchitecture at a Glance Architecture of HQ Instance Basic / Site-HA\nArchitecture of DR Instance (cross-region recovery)\nArchitecture of Edge / Sovereign Instance (VCF Edge model)\nComparing the Four Designs Design Protects Against Typical Role Basic Host/storage failure only Single-site HQ Site HA (Across Zones) Zone-level outage HQ with local resilience Disaster Recovery (Across Regions) Site-wide disaster HQ + DR pairing Fault Domains + DR Both zone and region failures HQ with local HA, plus DR Each design also maps to a named infrastructure blueprint in Broadcom\u0026rsquo;s Design Library \u0026ndash; what you actually select from when building the design in practice:\nDesign Infrastructure Blueprint Basic VCF Fleet in a Single Site (or \u0026hellip;with Minimal Footprint) Site HA (Across Zones) VCF Fleet with Multiple Sites in a Single Region Disaster Recovery (Across Regions) VCF Fleet with Multiple Sites Across Multiple Regions Fault Domains + DR VCF Fleet with Multiple Sites in a Single Region plus Additional Region(s) A Practical Tip Before committing to a topology, check the network latency and bandwidth requirements between sites \u0026ndash; stretched-cluster and cross-region designs both depend on meeting specific latency thresholds, and Broadcom\u0026rsquo;s Ports and Protocols documentation is the place to verify them before finalizing a design.\nWhat\u0026rsquo;s Next Next in this series: VCF 9 Management Domain Anatomy \u0026ndash; what actually lives inside a management domain, and how the first one differs from the rest.\nFurther Reading (Official Broadcom Documentation) VCF Taxonomy VCF Fleet Deployment Models VCF Edge Models Architectural Options in VMware Cloud Foundation ","permalink":"https://virtualizationgurus.pages.dev/posts/vcf9-instance-model/","summary":"\u003ch2 id=\"introduction\"\u003eIntroduction\u003c/h2\u003e\n\u003cp\u003eOnce you\u0026rsquo;ve internalized \u0026ldquo;one Fleet across every site,\u0026rdquo; the next question is practical: how do you actually design that fleet across a headquarters site, a DR site, and edge or sovereign-cloud locations? This post breaks down the core VCF Instance/Fleet constructs and the four fleet deployment designs Broadcom documents for VCF 9.1, and maps them to the multi-site patterns operators actually build.\u003c/p\u003e\n\u003ch2 id=\"the-three-constructs-revisited\"\u003eThe Three Constructs, Revisited\u003c/h2\u003e\n\u003cp\u003eBroadcom\u0026rsquo;s VCF Taxonomy defines three nested constructs worth keeping straight:\u003c/p\u003e","title":"VCF 9 Instance Model: Designing HQ, DR, and Edge/Sovereign Topologies"},{"content":"Introduction Managing the lifecycle of a VMware Cloud Foundation environment has historically required coordinated upgrades across multiple products, ESX, vCenter, NSX, and vSAN, with complex compatibility matrices and manual orchestration. VCF 9 changes this experience by moving lifecycle management into VCF Operations: the SDDC Manager UI is deprecated, and its LCM workflows now live in VCF Operations Fleet Management, giving you a single place to plan and orchestrate upgrades across every domain.\nIn this post, we explore the VCF 9 lifecycle management architecture, the unified upgrade workflow, the new licensing model, and fleet-level health monitoring capabilities.\nArchitectural Overview The diagram below shows the VCF Operations lifecycle management architecture, including the relationships between the management plane, fleet components, and the upgrade orchestration workflow:\nVCF Operations: The Unified Management Plane Starting with VCF 9.0, the SDDC Manager UI is deprecated and its workflows move into VCF Operations and the vSphere Client. SDDC Manager itself is still installed as a component of every VCF 9 instance during this transition, but its lifecycle management capabilities, and the LCM API that used to live on SDDC Manager, now belong to VCF Operations Fleet Management. In practice, VCF Operations is where you go to manage one or more VCF instances. Key responsibilities include:\nFleet Management: Centralized inventory and health visibility across all management and workload domains, including multi-site VCF deployments, with SDDC Manager\u0026rsquo;s former LCM workflows absorbed into this capability Lifecycle Management (LCM): Orchestrated upgrades for all VCF components (ESX, vCenter, NSX, vSAN) with pre-upgrade validation, staged rollout, and post-upgrade health checks License Management: Tracking and submission of VCF usage data under the per-core primary license model Cost and Capacity Management: Resource utilization visibility and capacity forecasting across all domains Security Compliance: Integration with Broadcom\u0026rsquo;s security compliance frameworks for VCF component configuration validation VCF 9 Lifecycle Management: Upgrade Workflow The VCF 9 upgrade workflow is orchestrated entirely through VCF Operations, providing a single workflow for upgrading all VCF components across one or multiple domains.\nStep 1: Bundle Management VCF 9 uses upgrade bundles that contain all required component images:\nOnline Depot: Connect VCF Operations to Broadcom Customer Connect (online) to download bundles automatically. VCF Operations checks for available updates and presents them in the Lifecycle Management dashboard. Offline/Air-gapped: Download bundles from Broadcom Customer Connect and import to a local depot. VCF Operations consumes bundles from the local depot without internet connectivity. A VCF update bundle includes:\nESX 9.0 update image (compatible with vSphere Lifecycle Manager) vCenter Server update or patch image NSX update image: NSX virtual networking kernel modules (VIBs) ship bundled with ESX by default in VCF 9, and NSX VIB lifecycle is now tied to ESX lifecycle rather than managed separately vSAN update (typically part of the ESX image for ESA deployments) Step 2: Pre-Upgrade Validation Before applying any upgrade, VCF Operations performs automated pre-upgrade checks:\nHardware compatibility: Validates all ESX hosts against the Broadcom Compatibility Guide (BCG) for the target upgrade version Component compatibility: Verifies the upgrade bundle is compatible with the current VCF component versions (interoperability matrix check) Health checks: Validates vSAN cluster health, NSX control plane health, and vCenter inventory state License validation: Confirms VCF licenses are valid and usage has been submitted within the required 180-day window DNS/NTP validation: Verifies DNS resolution and NTP synchronization for all management domain components Address all pre-upgrade failures before proceeding. VCF Operations clearly identifies which checks failed and provides remediation guidance.\nStep 3: Upgrade Orchestration Once validation passes, VCF Operations orchestrates the upgrade in the correct sequence:\nESX hosts: ESX hosts are upgraded using vSphere Lifecycle Manager (vLCM) integration. Hosts are placed in maintenance mode sequentially (or using custom scheduling), updated, and returned to service. DRS ensures workload migration during maintenance. vCenter Server: vCenter is upgraded after all ESX hosts in the domain are updated. The upgrade uses the vCenter appliance upgrade mechanism and is orchestrated by VCF Operations. NSX: NSX Manager and transport nodes are upgraded. Because NSX VIBs ship with ESX, NSX VIB updates land as part of the ESX host upgrade rather than a separate maintenance cycle, and NSX VIBs on ESX hosts support ESX Live Patch: allowing certain updates to apply without maintenance mode or impact to vMotion and DRS. vSAN: vSAN upgrades are typically part of the ESX upgrade in ESA deployments. VCF Operations monitors vSAN health throughout the upgrade. Step 4: Post-Upgrade Validation After upgrade completion, VCF Operations automatically runs post-upgrade health checks:\nAll components show expected version numbers NSX control plane is healthy and all transport nodes show as Up vSAN cluster health is green with no rebuild operations in progress vCenter inventory is consistent and all hosts show as Connected VCF 9 Licensing Model VCF 9 replaces the many separate per-product license keys of earlier VCF versions with subscription-based license files assigned at the vCenter level. A few mechanics are worth understanding before you plan a deployment:\nPrimary licenses are per-core. VCF and vSphere Foundation licenses are consumed by ESX hosts, calculated from total physical CPU cores across your environment. Most products enforce a 16-core-per-CPU minimum: a CPU with fewer physical cores (say, 8) still consumes 16 cores of license capacity. You license vCenter instances, not hosts directly. You assign a primary license to a vCenter instance, and the other components connected to that vCenter (ESX, and by extension NSX) are licensed automatically. vSAN capacity is a separate, pooled license measured in TiB. A VCF subscription delivers both a default VMware Cloud Foundation (per-core) license and a default vSAN Enterprise (per-TiB) license. If your storage needs exceed the vSAN capacity included with your core purchase, you buy additional vSAN Capacity Add-on licenses in per-TiB increments. Capacity pools automatically. Multiple subscriptions for the same product, same unit of measure, and same Site ID combine into one default license rather than staying as separate license entries: useful when you add capacity incrementally over time. New installs get a 90-day evaluation period before licensing is required (stateless ESX hosts provisioned via Auto Deploy have no evaluation period in 9.0). License Submission VCF 9 uses a usage-based license submission model:\nLicense usage data must be submitted to Broadcom at least once every 180 days: miss the window and licenses are treated as expired, disconnecting hosts from vCenter and blocking workload operations In connected mode, VCF Operations generates and submits usage files automatically (daily) In disconnected mode, you export a signed usage file from VCF Operations and upload it to Broadcom Customer Connect manually on a regular interval Licensing in Disconnected/Air-gapped Environments VCF Operations supports fully disconnected (air-gapped) operation:\nNo internet connectivity required for LCM operations when using a local depot License usage data can be exported from VCF Operations as a signed file and manually uploaded to Broadcom Customer Connect VCF Operations continues to operate normally in disconnected mode; license submission is the only operation requiring external connectivity (and can still be done offline via manual file upload) Fleet-Level Health Monitoring VCF Operations provides fleet-level health monitoring across all VCF domains:\nDashboard and Alerts Fleet Overview: Single pane showing all management and workload domains with health status (green/yellow/red) Component Health: Per-component health (ESX, vCenter, NSX, vSAN) with drill-down from domain to individual host Proactive Alerts: Integration with Broadcom Skyline Health proactively identifies issues, including predictive hardware alerts from Proactive Hardware Management (PHM), NSX transport node issues, and certificate expiry warnings Certificate Lifecycle Management VCF Operations manages certificate lifecycle for all VCF components:\nTracks certificate expiry dates across all domains and components Automates certificate rotation for management domain components (vCenter, NSX Manager, VCF Operations itself) Sends proactive alerts when certificates approach expiry VCF SDK and Automation VCF 9 provides comprehensive automation capabilities through a unified VCF SDK that consolidates what used to be separate per-component SDKs:\nREST API: Roughly 90% of VCF APIs now follow the OpenAPI 3.0 specification, improving consistency across lifecycle management, domain management, and license operations Python and Java SDK: A single unified VCF SDK for both languages, installable online via PyPI (Python) and Maven (Java), or offline through the Broadcom Developer Portal for regulated/air-gapped environments VCF PowerCLI: Auto-generated PowerShell modules providing 1:1 bindings with the VCF REST APIs, part of the broader PowerCLI suite Example: Triggering an Upgrade via API # Example: Using VCF Python SDK to trigger an LCM upgrade from vcf_sdk import VcfClient client = VcfClient(host=\u0026#34;vcf-operations.example.com\u0026#34;, username=\u0026#34;admin\u0026#34;, password=\u0026#34;...\u0026#34;) # Get available upgrade bundles bundles = client.lifecycle.get_bundles(domain_id=\u0026#34;management-domain-id\u0026#34;) print(f\u0026#34;Available bundles: {[b.version for b in bundles]}\u0026#34;) # Run pre-upgrade checks validation = client.lifecycle.validate_upgrade( domain_id=\u0026#34;management-domain-id\u0026#34;, bundle_id=bundles[0].id ) print(f\u0026#34;Validation status: {validation.status}\u0026#34;) # Initiate upgrade if validation passes if validation.status == \u0026#34;SUCCEEDED\u0026#34;: upgrade = client.lifecycle.start_upgrade( domain_id=\u0026#34;management-domain-id\u0026#34;, bundle_id=bundles[0].id ) print(f\u0026#34;Upgrade started: {upgrade.id}\u0026#34;) What\u0026rsquo;s Next In the next post, we will explore VCF 9 security and compliance, covering the NSX Distributed Firewall design, VPC-level isolation, and how VCF Operations integrates with Broadcom\u0026rsquo;s security compliance frameworks to maintain a hardened VCF environment.\nFurther Reading (Official Broadcom Documentation) VCF Operations: What\u0026rsquo;s New Licensing Model Updating Licenses and Viewing the License Usage File VCF SDKs, APIs, and CLIs: What\u0026rsquo;s New ","permalink":"https://virtualizationgurus.pages.dev/posts/vcf9-lifecycle-management/","summary":"\u003ch2 id=\"introduction\"\u003eIntroduction\u003c/h2\u003e\n\u003cp\u003eManaging the lifecycle of a VMware Cloud Foundation environment has historically required coordinated upgrades across multiple products, ESX, vCenter, NSX, and vSAN, with complex compatibility matrices and manual orchestration. VCF 9 changes this experience by moving lifecycle management into VCF Operations: the SDDC Manager UI is deprecated, and its LCM workflows now live in VCF Operations Fleet Management, giving you a single place to plan and orchestrate upgrades across every domain.\u003c/p\u003e","title":"VCF 9 Lifecycle Management with VCF Operations: Unified Upgrades and Fleet Management"},{"content":"Introduction vSAN Express Storage Architecture (ESA) is the recommended storage architecture for all new VMware Cloud Foundation 9 deployments. ESA represents a ground-up redesign of the vSAN storage stack, purpose-built to exploit the performance characteristics of modern NVMe flash media. In this post, we explore the vSAN ESA architecture, design decisions, VCF 9.1\u0026rsquo;s new Auto-RAID capability, and best practices for storage policy design.\nArchitectural Overview The diagram below illustrates the vSAN ESA architecture in a VCF 9 environment, from the NVMe hardware layer through the vSAN data path and management plane:\nESA vs OSA: Architecture Comparison vSAN ESA fundamentally differs from the original vSAN Original Storage Architecture (OSA):\nFeature OSA ESA Drive Tiers Cache + Capacity (2-tier) Single NVMe pool (1-tier) NVMe Support Limited (cache tier only) Full NVMe-native throughout Compression Software only (post-process) Always-on cluster service (VCF 9.1), hardware-assisted Snapshot Performance Degrades under load Metadata-based, near-zero impact RAID-5/6 Efficiency Requires 6 hosts (RAID-6) Available with 4 hosts (RAID-5), auto-selected via Auto-RAID in VCF 9.1 Drive Health Monitoring Basic health alerts Proactive Hardware Management (PHM) for predictive NVMe failure detection File System VMFS-based (COW) Log-Structured File System (LFS) Minimum Hosts 3 3 Minimum Drives/Host 1 cache + 1 capacity 3 NVMe (all-flash single tier) ESA Key Design Decisions Single-Tier NVMe Pool ESA eliminates the two-tier cache/capacity split used in OSA. All NVMe drives contribute to a single unified storage pool, delivering:\nConsistent low-latency I/O across all stored objects regardless of working set size Simplified capacity planning with no cache-to-capacity ratio calculations required Minimum 3 NVMe drives per host for ESA (the Broadcom Compatibility Guide lists ESA-certified NVMe drives) All drives must be certified for ESA: not all NVMe drives qualify; refer to the Broadcom Compatibility Guide (BCG) for the ESA Compatibility category Log-Structured File System (LFS) ESA uses a purpose-built Log-Structured File System optimized for NVMe characteristics:\nSequential write patterns to all NVMe devices, maximizing write throughput In VCF 9.1, compression is now an always-on cluster service rather than a per-policy toggle, reducing write amplification and extending drive life with no configuration required Efficient space reclamation through background compaction without impacting foreground I/O Near-Zero Impact Snapshots ESA\u0026rsquo;s log-structured architecture enables a metadata-based snapshot engine that is fundamentally different from OSA\u0026rsquo;s copy-on-write approach:\nSnapshots are tracked through metadata pointers rather than copying data blocks, so creation is crash-consistent without stunning the VM Snapshot deletion is largely a metadata operation, acknowledged immediately, with underlying data reclaimed asynchronously, and is dramatically faster than OSA\u0026rsquo;s redo-log-based mechanism Snapshot trees do not degrade read/write performance, enabling efficient use of snapshots for backup integration (VMware Live Recovery) Supports VMware Live Recovery with RPO as low as 1 minute for vSAN-to-vSAN replication in supported configurations New in VCF 9.1: Auto-RAID VCF 9.1 introduces Auto-RAID, a fully system-managed approach to data resilience that replaces manual RAID/FTT policy selection as the recommended default for vSAN ESA clusters:\nA single \u0026ldquo;vSAN ESA Auto RAID Policy\u0026rdquo; governs all vSAN 9.1 clusters cluster-wide: no explicit resilience settings are stored in the policy itself; vSAN senses cluster characteristics (host count, topology) and applies the optimal RAID level automatically Standard clusters with 6+ hosts: FTT=2 using RAID-6 (1.5x capacity overhead) Standard clusters with 3–5 hosts: FTT=1 using RAID-5, always using the 2+1 erasure code (1.5x capacity overhead) Fewer than 3 hosts: FTT=0 (1.0x capacity overhead) until the cluster scales up Stretched clusters and 2-node clusters have their own Auto-RAID resilience tables, layering a mirror-based site/host disaster tolerance on top of the same erasure-coding logic Cluster changes (adding/removing hosts) are re-evaluated dynamically, so resilience adjusts automatically as the cluster grows All new VCF 9.1 clusters use Auto-RAID by default. Existing clusters upgraded to 9.1 keep their prior storage policy or Auto-Policy Management configuration until migrated, though vSAN surfaces a health alert recommending the switch This is a meaningful shift from earlier vSAN ESA guidance (including VCF 9.0), where administrators manually selected RAID-1/5/6 per policy based on host count and workload criticality.\nProactive Hardware Management (PHM) for NVMe Drives vSAN\u0026rsquo;s Proactive Hardware Management (PHM) capability applies to ESA NVMe drives in VCF 9:\nPredictive detection: PHM surfaces disk predictive-failure events generated by the OEM vendor through a registered Hardware Support Manager (HSM), rather than waiting for a hard failure Continuous polling: PHM checks for HSM-generated hardware failure events every 10 minutes Administrator-driven remediation: Based on the predictive failure signal, PHM lets you take the appropriate remediation action (such as proactively evacuating data from the affected drive) before an unplanned failure occurs Alert integration: PHM events surface through the vSAN management service on vCenter and integrate with vSAN Health and Skyline Health for fleet-wide visibility in VCF Operations Requires a supported Hardware Support Manager registered to vCenter: without an HSM, vSAN cannot receive OEM predictive-failure signals for PHM.\nvSAN ESA Storage Policies in VCF 9 Storage policies in vSAN ESA are defined through VM Storage Policies applied at the VM or VMDK level. As of VCF 9.1, Auto-RAID is the recommended default for resilience settings (see above); the guidance below reflects the underlying RAID mechanics and remains relevant for pre-9.1 clusters or environments not yet migrated to Auto-RAID.\nKey Policy Parameters Failures to Tolerate (FTT): the number of host failures the cluster can sustain:\nFTT=1 with RAID-1 (mirroring): minimum 3 hosts required FTT=1 with RAID-5 (erasure coding): minimum 4 hosts, more space-efficient than RAID-1 FTT=2 with RAID-6 (erasure coding): minimum 6 hosts, recommended for production workloads Storage Policy Best Practices for VCF 9:\nOn VCF 9.1, start with the vSAN ESA Auto RAID Policy as the datastore default: it removes the guesswork of matching RAID level to host count and adjusts automatically as the cluster scales If manual policies are still required (pre-9.1 clusters, or specific IOPS limit / Object Space Reservation / stretched-cluster site-locality needs), use RAID-5/FTT=1 for general workloads with 4+ hosts and RAID-6/FTT=2 for business-critical VMs requiring tolerance of 2 simultaneous host failures Define separate storage policies for management VMs and workload VMs only where Auto-RAID\u0026rsquo;s cluster-wide policy doesn\u0026rsquo;t fit the requirement (for example, custom IOPS limits) Enable storage-based policy management (SPBM) integration with VCF Automation for self-service policy assignment vSAN File Services VCF 9 supports vSAN File Services on ESA clusters:\nSupports up to 500 file shares per cluster, of which up to 100 can be SMB Provides NFS v3/v4.1 and SMB protocol support for workloads requiring shared file storage VCF 9.1 delivers up to 2x faster metadata operations for SMB workloads File services VMs are deployed automatically within the vSAN cluster Useful for Kubernetes persistent volumes (NFS-backed) and legacy application shared storage requirements vSAN ESA and VMware Live Recovery Integration VCF 9\u0026rsquo;s vSAN ESA integrates with VMware Live Recovery for disaster recovery:\nHost-based vSAN-to-vSAN replication with RPO as low as 1 minute for protected VMs Metadata-based snapshots enable efficient replication without degrading production I/O VCF 9.1 adds seeding for vSAN replication, using existing replicas to avoid full re-synchronization when creating new replication pairs Failover orchestration is managed through VCF Operations or the dedicated Live Recovery management interface Supports both on-premises-to-on-premises and on-premises-to-VMware Cloud replication topologies Operational Best Practices Host maintenance and replacement:\nBefore removing a host from the cluster, use vSphere maintenance mode with \u0026ldquo;Ensure Accessibility\u0026rdquo; or \u0026ldquo;Full Data Migration\u0026rdquo; to safely migrate vSAN objects For ESA, a minimum of 3 fully operational hosts (not including the host in maintenance) is required to maintain FTT=1 protection Capacity management:\nMonitor vSAN capacity from VCF Operations \u0026gt; Cost and Capacity Management VCF 9.1\u0026rsquo;s Auto-RAID standardizes capacity overhead per cluster type (1.5x for standard clusters, 3x for stretched, 2x for 2-node), which simplifies capacity forecasting compared to mixed manual policies Plan for vSAN resync capacity: maintain 10-15% slack capacity for background resync after a host failure or replacement Monitoring and health:\nEnable Skyline Health integration in VCF Operations for proactive hardware health alerts Review PHM alerts regularly: a predictive-failure signal indicates a drive is near end of life and should be scheduled for replacement vSAN performance diagnostics are available from vCenter \u0026gt; Monitor \u0026gt; vSAN \u0026gt; Performance What\u0026rsquo;s Next In the next post, we cover VCF 9 Lifecycle Management with VCF Operations: Unified Upgrades and Fleet Management, walking through the unified upgrade workflow for ESX, vCenter, NSX, and vSAN, the new simplified licensing model, and fleet-level health monitoring across every VCF domain.\nFurther Reading (Official Broadcom Documentation) vSAN 9.1 New Features: Release Notes Selecting the Best RAID Configuration for a vSAN Storage Cluster Managing Proactive Hardware vSAN File Service ","permalink":"https://virtualizationgurus.pages.dev/posts/vcf9-vsan-esa-deep-dive/","summary":"\u003ch2 id=\"introduction\"\u003eIntroduction\u003c/h2\u003e\n\u003cp\u003evSAN Express Storage Architecture (ESA) is the recommended storage architecture for all new VMware Cloud Foundation 9 deployments. ESA represents a ground-up redesign of the vSAN storage stack, purpose-built to exploit the performance characteristics of modern NVMe flash media. In this post, we explore the vSAN ESA architecture, design decisions, VCF 9.1\u0026rsquo;s new Auto-RAID capability, and best practices for storage policy design.\u003c/p\u003e\n\u003ch2 id=\"architectural-overview\"\u003eArchitectural Overview\u003c/h2\u003e\n\u003cp\u003eThe diagram below illustrates the vSAN ESA architecture in a VCF 9 environment, from the NVMe hardware layer through the vSAN data path and management plane:\u003c/p\u003e","title":"VCF 9 vSAN ESA Deep Dive: Architecture, Performance, and New Features"},{"content":"Introduction One of the most impactful networking changes in VCF 9 is the introduction of Virtual Private Cloud (VPC) networking as the primary multi-tenancy model in NSX 9.0. VPCs replace the manual overlay segment and Tier-1 gateway model of previous VCF releases with a cloud-native, self-service networking abstraction that aligns with industry standards.\nIn this post, we dive deep into the VCF 9 VPC architecture, Transit Gateway design patterns, and how application teams can consume VPC-based networking from vCenter, VCF Automation, and vSphere Supervisor.\nArchitectural Overview The diagram below illustrates the complete NSX 9.0 VPC networking architecture in VCF 9, from the physical network underlay through the VPC overlay consumed by workloads:\nKey VPC Concepts in NSX 9.0 Virtual Private Cloud (VPC) A VPC in NSX 9.0 is a self-contained networking environment with:\nComplete isolation from other VPCs by default (no inter-VPC communication without explicit Transit Gateway attachment) Private – VPC subnets for internal workload-to-workload communication within the same VPC (NAT required for any outside communication) Private – Transit Gateway subnets for inter-VPC connectivity below the Transit Gateway without NAT (IP translation is still required if those workloads must be reachable from outside the environment) Public subnets for workloads requiring external IP exposure: reachable from other VPCs and from outside the environment, above the Transit Gateway DHCP enhancements supporting advanced configurations beyond basic use cases Multiple namespaces support: multiple Kubernetes namespaces can be assigned to a single VPC Transit Gateway (TGW) The NSX Transit Gateway is a central routing hub for inter-VPC and VPC-to-external-network communication:\nTGW Type Use Case Key Benefit Centralized TGW (CTGW) Shared external routing across multiple VPCs Simplified routing, single point of control Distributed TGW (DTGW) Direct host-to-fabric connectivity Lower latency, reduced Edge node dependency, non-blocking performance VCF 9 also supports Transit Gateways with Distributed VLAN Connectivity (DTGW), which provides a direct, high-performance datapath from ESX hosts to the network fabric without requiring additional Edge node infrastructure.\nNSX automatically creates a default Transit Gateway for every NSX Project (tenant), including the Default project. Additional CTGWs and DTGWs can be added as requirements grow, and a VPC can be re-pointed to a different TGW simply by changing its connectivity profile, NSX runs IPAM and VPC span validation checks whenever a VPC is moved. VCF 9.1 also adds the option of a Distributed VXLAN Connection alongside Distributed VLAN Connectivity for DTGWs, and independent HA modes (active-active or active-standby) per CTGW instead of inheriting HA from the Tier-0/VRF gateway, giving architects more flexibility in how the transit layer is built.\nConnectivity and Service Profiles VPC 9 introduces Connectivity and Service Profiles that allow cloud administrators to define:\nExternal connectivity patterns (which TGW a VPC connects to) Services available to VPCs (NAT, DHCP, Load Balancing) Centralized policy that multiple VPCs can reference This simplification means application teams can create VPCs by selecting a profile, without needing to understand the underlying NSX infrastructure topology.\nEnhanced Data Path (EDP) Standard NSX 9.0 introduces EDP Standard as the default host switch mode for all new VCF installations and workload domains. EDP Standard:\nDelivers superior throughput, packet rate, and reduced latency compared to the legacy Standard stack Supports NSX Switch Port Analyzer (SPAN) and Live Traffic Analysis in the fast path Is available for new deployments: upgraded/imported workload domains continue using the legacy stack until explicitly migrated Network Span (New in VCF 9.1) VCF 9.1 introduces Network Span, a logical construct that scopes which vSphere clusters can see the VPC subnets attached to a given Transit Gateway (Centralized or Distributed), instead of every subnet being available across every cluster in the environment by default. Two span types matter in practice:\nDefault span: every vCenter cluster belongs to it unless explicitly removed. Exclusive span: a dedicated set of clusters carved out for a specific workload; assigning clusters to an exclusive span also removes them from the default span, so capacity isn\u0026rsquo;t double-counted and spans don\u0026rsquo;t silently overlap. For architects, Network Span is the lever for keeping a VPC\u0026rsquo;s subnets confined to the clusters that should actually host that tenant\u0026rsquo;s workloads, rather than exposing them fleet-wide.\nVPC Consumption Patterns From vCenter (Infrastructure Team) Infrastructure administrators can create and manage VPCs directly from the vCenter networking pane:\nNavigate to vCenter → Networking → VPCs Create a VPC with name, CIDR range, and Connectivity Profile Create subnets (private or public) within the VPC Attach VMs to subnets via standard vCenter port group assignment From VCF Automation (Application Teams - Self-Service) VCF Automation provides a cloud-like self-service experience for VPC consumption:\nApplication teams request VPCs and subnets from a catalog VMs and Kubernetes clusters are automatically placed in the appropriate VPC Network isolation and security policies are automatically applied From vSphere Supervisor (Platform Engineering) vSphere Supervisor integrates VPCs as the fundamental networking building block for Kubernetes workloads:\nVKS clusters and vSphere Pods are placed inside VPCs K8s NetworkPolicy and NSX SecurityPolicy work together for workload-level micro-segmentation StaticRoutes within VPCs enable specific routing requirements for multi-tier applications Security Considerations Gateway Firewall - Disabled by Default In NSX 9.0, Gateway Firewall is disabled by default for all new Tier-0 and Tier-1 gateway deployments. This improves performance and resource utilization. Enable Gateway Firewall only for perimeter security requirements where stateful inspection at the gateway level is needed.\nDistributed Firewall (DFW) NSX Distributed Firewall remains active and is the primary security mechanism for east-west traffic within and between VPCs. DFW in VCF 9:\nRuns in EDP Standard mode for improved performance Supports context-aware policies using VM tags and security groups Operates at the hypervisor level for zero-trust micro-segmentation VPC-Level Isolation VPCs provide built-in isolation: by default, VMs in different VPCs cannot communicate. Inter-VPC communication requires explicit Transit Gateway attachment and routing configuration, providing an isolation-by-default security posture.\nWhat\u0026rsquo;s Next In the next post, we cover the VCF 9 vSAN ESA Deep Dive: Architecture, Performance, and New Features, taking a close look at how the vSAN Express Storage Architecture reshapes storage design, performance characteristics, and what\u0026rsquo;s new for VCF 9 deployments.\nFurther Reading (Official Broadcom Documentation) Virtual Private Clouds Overview Transit Gateways NSX: What\u0026rsquo;s New in VCF 9.0 Add a Network Span ","permalink":"https://virtualizationgurus.pages.dev/posts/vcf9-nsx-vpc-deep-dive/","summary":"\u003ch2 id=\"introduction\"\u003eIntroduction\u003c/h2\u003e\n\u003cp\u003eOne of the most impactful networking changes in VCF 9 is the introduction of \u003cstrong\u003eVirtual Private Cloud (VPC)\u003c/strong\u003e networking as the primary multi-tenancy model in NSX 9.0. VPCs replace the manual overlay segment and Tier-1 gateway model of previous VCF releases with a cloud-native, self-service networking abstraction that aligns with industry standards.\u003c/p\u003e\n\u003cp\u003eIn this post, we dive deep into the VCF 9 VPC architecture, Transit Gateway design patterns, and how application teams can consume VPC-based networking from vCenter, VCF Automation, and vSphere Supervisor.\u003c/p\u003e","title":"VCF 9 NSX VPC Deep Dive: Cloud-Native Networking for Your Private Cloud"},{"content":"Introduction With the VCF 9 management domain up and running, the next step in your private cloud journey is creating workload domains. A workload domain groups ESX hosts under a dedicated vCenter instance, along with NSX networking and storage, so it can host tenant or departmental workloads. Depending on how you configure NSX connectivity during creation, a workload domain can become VPC-ready either immediately or after a short follow-up step: more on that below.\nThree Ways to Create a Workload Domain Per VCF Operations\u0026rsquo; Workload Domain wizard, there are three supported paths:\nFull deployment with cluster: deploys a fully provisioned workload domain with an initial vSphere cluster. Requires ESX hosts already commissioned with the target principal storage type. Domain infrastructure only: deploys and configures a new vCenter instance and a new-or-shared NSX Manager instance, without requiring any unassigned ESX hosts. You add a vSphere cluster to it later. Import an existing vCenter: brings an already-running vCenter and its managed ESX hosts under VCF as a workload domain, so it\u0026rsquo;s included in centralized identity, certificate, and lifecycle management. Architectural Overview The diagram below illustrates the relationship between the VCF 9 management domain, VCF Operations, and multiple workload domains, including VPC-based networking topology:\nWorkload Domain Planning Considerations Before creating a workload domain in VCF 9, use the VCF Planning and Preparation Workbook to document all design decisions. Key areas to plan:\nCompute Sizing Minimum of 3 ESX hosts for a vSAN ESA workload domain Size CPU and RAM based on the expected VM density and workload type (compute-intensive, memory-intensive, or mixed) Account for ESX host management overhead and vSAN overhead when calculating usable capacity For high availability, plan for N+1 host capacity so one host can fail without a performance hit Storage Design A workload domain can be built on one of four principal storage types, selected during the wizard\u0026rsquo;s Storage step:\nvSAN: either ESA (the current recommended architecture) or OSA; requires SSD or NVMe disks free of pre-existing partitions NFS: an external NFS datastore, identified by server IP and export path VMFS on FC: Fibre Channel-backed VMFS datastore vVols: supported for compatibility, but Broadcom has marked vVols as deprecated as of VCF/vSphere Foundation 9.0, with removal planned in a future release; new designs should avoid it Networking Design Network Segment Purpose Configuration Notes Management VLAN ESX management, vCenter Must be reachable from VCF Operations vMotion VLAN Live VM migration Dedicated VLAN for performance isolation vSAN VLAN Storage traffic Dedicated VLAN, jumbo frames recommended NSX Overlay (Host TEP) Tenant Geneve-encapsulated traffic Needs a static IP pool or DHCP scope NSX Edge Uplink North-south routing to physical network Dedicated VLANs for T0 connectivity NSX and VPC Planning Every workload domain needs an NSX Manager, which can either be a new dedicated instance or a shared instance already used by another domain (subject to version-compatibility rules between the shared NSX Manager and each domain\u0026rsquo;s vCenter version). When configuring VPC Gateway Connectivity during creation, you choose between:\nCentralized Connectivity: simpler to set up, but the domain becomes VPC-ready only after you separately deploy an NSX Edge cluster with a Tier-0 gateway afterward Distributed Connectivity: requires a dedicated VLAN, gateway CIDR, and external/private IP blocks up front, but the workload domain is VPC-ready immediately after creation Deploying a Workload Domain via VCF Operations Step 1: Commission ESX Hosts Before creating a full-deployment workload domain, commission the target ESX hosts (skip this if you\u0026rsquo;re using the \u0026ldquo;domain infrastructure only\u0026rdquo; or \u0026ldquo;import existing vCenter\u0026rdquo; paths):\nIn VCF Operations, commission hosts with the correct target principal storage type VCF Operations validates host connectivity, DNS, NTP, and hardware compliance Hosts move to an unassigned state once commissioned, ready to be selected during workload domain creation Step 2: Create the Workload Domain In VCF Operations, click Operate in the top navigation bar, then in the left pane select Overview → Inventory Expand VCF Instances and select the instance where you\u0026rsquo;re creating the domain From the Add workload domain drop-down, choose Create new, review prerequisites, and proceed Work through the wizard: General Information → vCenter → Cluster → Image → Networking (NSX) → Storage → Hosts → vSphere Distributed Switch → vSphere Supervisor (optional) → Finish VCF Operations performs pre-deployment validation, then orchestrates the full domain bringup Track progress under Fleet Management → Tasks Step 3: Post-Creation VPC Setup If you chose Centralized Connectivity, complete VPC-readiness by deploying an NSX Edge cluster with an active-standby Tier-0 gateway. If you chose Distributed Connectivity, the domain is already VPC-ready: proceed straight to creating VPCs, subnets, and Transit Gateway connectivity for each tenant or application team.\nDay-2 Workload Domain Operations VCF Operations provides built-in automation for common Day-2 operations, and VCF 9.1 adds guardrails so that certain out-of-band changes (made outside the VCF Operations UI, such as expanding a cluster or changing primary storage directly in vCenter) no longer create unresolvable configuration drift:\nCluster expansion/contraction: Add or remove ESX hosts from an existing cluster via VCF Operations Host remediation: Apply ESX updates, patches, or desired-state images across the cluster using vSphere Lifecycle Manager images Certificate rotation: Automate certificate lifecycle for all workload domain components from VCF Operations License management: Track and submit VCF and vSAN license usage periodically directly from VCF Operations Sources Validated against Broadcom\u0026rsquo;s official VCF 9.1 documentation: Building Cloud Infrastructure, Managing VCF Domains, and Create a New Workload Domain.\nWhat\u0026rsquo;s Next In the next post, we will do a deep dive into VCF 9 NSX VPC networking, covering Transit Gateway design, VPC consumption from vCenter, and practical examples of deploying application workloads using VPC-based isolation.\n","permalink":"https://virtualizationgurus.pages.dev/posts/vcf9-workload-domain-planning/","summary":"\u003ch2 id=\"introduction\"\u003eIntroduction\u003c/h2\u003e\n\u003cp\u003eWith the VCF 9 management domain up and running, the next step in your private cloud journey is creating workload domains. A workload domain groups ESX hosts under a dedicated vCenter instance, along with NSX networking and storage, so it can host tenant or departmental workloads. Depending on how you configure NSX connectivity during creation, a workload domain can become VPC-ready either immediately or after a short follow-up step: more on that below.\u003c/p\u003e","title":"VCF 9 Workload Domain Planning and Deployment"},{"content":"Introduction VMware Cloud Foundation 9 collapses what used to be three separate consoles into one control plane. If you\u0026rsquo;re coming from a 5.x background, this is arguably the single biggest mental model shift in the release: you stop thinking \u0026ldquo;one SDDC Manager per site\u0026rdquo; and start thinking \u0026ldquo;one Fleet across every site.\u0026rdquo; This post breaks down what Fleet actually is, which components feed into it, and how the pieces map to Broadcom\u0026rsquo;s official architecture documentation.\nFrom Per-Site Consoles to a Single Fleet In VCF 5.x, operators juggled Cloud Builder for initial bring-up, SDDC Manager for ongoing lifecycle work, and a separate automation tool layered on top for self-service provisioning. Each of those tools generally spoke to a single site. Running HQ, a DR site, and an edge location meant stitching together multiple, mostly disconnected management planes.\nVCF 9 restructures this around two core constructs, defined in Broadcom\u0026rsquo;s official VCF Taxonomy documentation: the VCF Instance (the compute, storage, and networking infrastructure running actual workloads: a management domain plus optional workload domains) and the VCF fleet (the environment managed by a single set of fleet-level components, which can span one or more Instances and even standalone vCenter deployments).\nThe Components That Make Up Fleet According to the official documentation, three components are involved, though they play distinct roles rather than being three equal peers sitting \u0026ldquo;inside\u0026rdquo; Fleet:\nVCF Installer is a dedicated virtual appliance used for Day-0 work: planning, validating, and deploying a new VCF or vSphere Foundation platform, converging existing infrastructure into VCF, or extending an existing fleet with an additional Instance. It ships as part of the SDDC Manager appliance OVA and can operate in two modes: a standalone \u0026ldquo;installer mode\u0026rdquo; for deploying multiple platforms, or it transitions into the SDDC Manager role for the Instance it just deployed. This is the direct functional replacement for the old Cloud Builder appliance.\nVCF Operations (formerly Aria Operations) is the persistent, fleet-level operations plane. Per Broadcom\u0026rsquo;s overview, it helps build, manage, operate, and secure the private cloud across four functional areas, build, manage, operate, and protect, covering monitoring, alerting, capacity, and compliance across the whole fleet, not just a single site. This is what absorbs SDDC Manager\u0026rsquo;s Day-2 operational role in the new model.\nVCF Automation (formerly Aria Automation) provides the self-service layer: a multi-tenant Infrastructure-as-a-Service catalog with policy-based governance, spanning cloud services, provider management, organization management, and vSphere Supervisor for Kubernetes workloads.\nStrictly speaking, the persistent \u0026ldquo;VCF fleet\u0026rdquo; is composed of VCF Operations and VCF Automation running together in the management domain of the first Instance: that pairing is what the taxonomy defines as the fleet-level component set. VCF Installer is the tool you use to create and extend that fleet, rather than a service that runs continuously inside it. It\u0026rsquo;s a subtle distinction, but one worth knowing if you\u0026rsquo;re mapping this to the official architecture rather than a simplified mental model.\nArchitecture at a Glance The Mental Model Shift VCF 5.x VCF 9 Cloud Builder per deployment VCF Installer bootstraps and extends the fleet SDDC Manager per site VCF Operations manages Day-2 across every Instance Bolted-on automation tool VCF Automation as a native, policy-driven catalog Separate logins per site One login, one inventory, one Fleet A Note on Instance Topologies Multi-site patterns like a primary/HQ Instance, a DR Instance, and edge or sovereign-cloud Instances are common real-world deployment topologies under the VCF Instance/Fleet model, rather than a fixed set of named categories in the core taxonomy documentation itself. Broadcom\u0026rsquo;s Design guidance covers specific reference architectures for these patterns in more depth: worth a read if you\u0026rsquo;re designing a multi-site fleet.\nSources This post was validated against Broadcom\u0026rsquo;s official VCF 9.1 documentation, specifically the VCF Taxonomy, VCF Installer Overview, VCF Operations Overview, and VCF Automation Overview pages under VMware Cloud Foundation 9.1 documentation.\nUp Next Next in this series: VCF 9 Workload Domain Planning and Deployment: putting the fleet management layer to work by planning, provisioning, and standing up your first workload domain.\n","permalink":"https://virtualizationgurus.pages.dev/posts/vcf9-fleet-management-layer/","summary":"\u003ch2 id=\"introduction\"\u003eIntroduction\u003c/h2\u003e\n\u003cp\u003eVMware Cloud Foundation 9 collapses what used to be three separate consoles into one control plane. If you\u0026rsquo;re coming from a 5.x background, this is arguably the single biggest mental model shift in the release: you stop thinking \u0026ldquo;one SDDC Manager per site\u0026rdquo; and start thinking \u0026ldquo;one Fleet across every site.\u0026rdquo; This post breaks down what Fleet actually is, which components feed into it, and how the pieces map to Broadcom\u0026rsquo;s official architecture documentation.\u003c/p\u003e","title":"VCF 9 Fleet Management Layer: One Control Plane for Every Instance"},{"content":"Introduction Deploying VMware Cloud Foundation 9 is a structured process that, when followed correctly, results in a fully configured SDDC in a matter of hours. This guide walks through the complete deployment flow from planning through a functional management domain.\nNote: VCF 9 replaces the Cloud Builder appliance used in VCF 5.x with the new VCF Installer: a new virtual appliance downloaded from the Broadcom Support portal that provides automated deployment and configuration workflows for the VCF environment. Unlike Cloud Builder which was a single-purpose OVA, the VCF Installer also supports VCF Converge (bringing existing vSphere infrastructure into VCF) and downloading/staging all required VCF component binaries.\nDeployment Flow Overview The following diagram illustrates the end-to-end VCF 9 deployment flow, from initial planning through a running management domain:\nPrerequisites Before starting the deployment, ensure you have the following in place:\nHardware Requirements Minimum hosts: 3 physical ESX hosts for management domain (production deployments should plan according to workload requirements; see the VCF Planning and Preparation Workbook for sizing guidance) CPU: Intel or AMD processors with VT-x and VT-d support; VCF 9 requires a minimum of 16 cores per CPU for ESX 9.0 licensing RAM: Size according to management workload requirements; management VMs (vCenter, NSX Manager, VCF Operations) require dedicated resources Storage: NVMe SSDs required for vSAN ESA (check the Broadcom Compatibility Guide for ESA-certified NVMe drives) Networking: 10 GbE minimum, 25 GbE recommended; 2 uplinks per host Important: Always cross-reference hardware against the Broadcom Compatibility Guide for ESX 9.0 HCL compliance before purchasing.\nNetwork Prerequisites Network Purpose Recommended VLAN Management ESX management, vCenter, VCF Ops Dedicated VLAN vMotion Live migration traffic Dedicated VLAN vSAN Storage traffic Dedicated VLAN NSX Overlay Geneve-encapsulated workload traffic Trunk NSX Edge Uplink North-south routing Dedicated VLANs DNS and NTP All FQDNs must be resolvable (forward and reverse DNS) before starting deployment:\nVCF Installer FQDN vCenter Server FQDN NSX Manager FQDN (and optional 3x NSX Manager cluster FQDNs for HA) ESX host FQDNs for all management domain hosts NTP server reachable from all hosts and management VMs Step 1: Download and Deploy the VCF Installer VCF 9 introduces the VCF Installer as the successor to Cloud Builder. The VCF Installer is:\nA virtual appliance (OVA) downloaded from the Broadcom Support portal Deployed to an existing ESX host or vCenter instance that is separate from your VCF management domain target hosts Part of the VCF SDK (with Python and Java bindings), integrated with PowerCLI, and provides a comprehensive OpenAPI 3.0 specification Capable of deploying new VCF environments, converging existing vSphere infrastructure to VCF, and downloading all necessary install binaries # VCF Installer deployment workflow: # 1. Download VCF Installer OVA from Broadcom Support portal # 2. Deploy OVA to an existing ESX/vCenter host (NOT the target management domain hosts) # 3. Power on the VCF Installer appliance # 4. Access the VCF Installer UI at: # https://\u0026lt;vcf-installer-appliance-fqdn\u0026gt;/ # Key differences from VCF 5.x Cloud Builder: # - Both are OVA virtual appliances, but VCF Installer supports Converge workflows # - VCF Installer integrates with VCF SDK for automation # - Deployment JSON replaces the Excel-based deployment parameter workbook Once deployed, access the VCF Installer UI and log in with your admin credentials.\nStep 2: Prepare the Deployment JSON In VCF 9, the Excel-based deployment parameter workbook used in VCF 5.x has been replaced by a JSON-based deployment specification. Use the VCF Planning and Preparation Workbook (available from the VCF documentation portal) to gather all required parameters, then use the VCF Installer UI or CLI to generate and validate the deployment JSON.\nFill in all required fields:\nManagement network parameters: IP addresses, subnet masks, gateway, DNS servers Host credentials: Root password for all ESX hosts License keys: VCF and vSAN licenses (VCF 9 uses 2 license types instead of 11 in older releases) DNS entries: Pre-populate all required DNS records (forward and reverse) before running the validation Step 3: Import and Validate the Deployment JSON In the VCF Installer:\nNavigate to Workflow → Deploy VMware Cloud Foundation Upload your completed deployment JSON specification Click Validate: the VCF Installer will perform over 200 pre-deployment checks including: DNS resolution (forward and reverse) for all FQDNs Network connectivity between hosts NTP synchronization Host hardware compatibility (HCL check) License key validity Tip: Address all validation failures before proceeding. Common issues include missing reverse DNS entries and NTP drift between hosts. The VCF Installer clearly indicates which checks failed and why.\nStep 4: Initiate Management Domain Bringup Once validation passes (all green), click Deploy to begin the management domain bringup. The process includes:\nPhase 1: ESX Configuration (15-20 min) Configures vSwitches and port groups on all hosts Sets NTP and DNS configuration Prepares hosts for vSAN cluster formation Phase 2: vSAN Cluster Formation (20-30 min) Creates vSAN ESA cluster Formats and initializes NVMe disks for ESA Configures vSAN storage policies ESA initialization includes hardware-level format steps specific to NVMe-native operation Phase 3: vCenter Deployment (30-45 min) Deploys vCenter Server 9.0 appliance to vSAN datastore Configures vCenter inventory (datacenter, cluster, hosts) Applies DRS and HA settings Phase 4: NSX 9.0 Deployment (45-60 min) Deploys NSX Manager (3-node cluster for production HA, or single node for resource-constrained/lab environments) Configures transport zones and host switch profiles NSX VIBs are already bundled with ESX 9.0: no separate VIB installation is required Configures VTEP pool and host transport nodes Enhanced Data Path (EDP) Standard is configured as the default host switch mode for new VCF installations Phase 5: VCF Operations and Management Domain Registration (20-30 min) Deploys VCF Operations (formerly SDDC Manager) appliance Imports and registers management domain inventory Configures initial license assignments (VCF 9 uses 2 license types: \u0026ldquo;VMware Cloud Foundation (cores)\u0026rdquo; and \u0026ldquo;VMware vSAN (TiBs)\u0026rdquo;) Enables fleet management, lifecycle management, and cost/capacity visibility Step 5: Post-Deployment Validation After the VCF Installer reports successful completion:\nLog in to VCF Operations at https://\u0026lt;vcf-operations-fqdn\u0026gt;/ Verify inventory: All ESX hosts should show as ACTIVE under the management domain Check NSX 9.0 status: Navigate to NSX Manager and confirm all transport nodes show as Up; verify EDP Standard mode is active Validate vSAN health: In vCenter, check vSAN health under the cluster: all checks should be green; verify ESA disk groups are formed correctly Run VCF Operations health check: From VCF Operations, review system health across all management domain components Verify licensing: Confirm VCF and vSAN licenses are applied and usage is being tracked; note that license usage must be submitted from VCF Operations every 180 days Common Deployment Issues and Fixes vSAN Disk Claim Failures If vSAN ESA fails to claim disks during Phase 2, verify:\nDisks are not presenting existing partition tables (wipe with esxcli storage core device partition delete) NVMe disks are certified for vSAN ESA (check the Broadcom Compatibility Guide: ESA requires specific NVMe certification) All ESA-required NVMe disks are visible and healthy in ESX NSX Manager Deployment Timeout If NSX Manager deployment times out:\nCheck vSAN health: insufficient capacity is the most common cause Verify the NSX Manager FQDN resolves correctly (including reverse DNS) from the VCF Installer appliance Ensure the management network has connectivity to all target hosts License Validation Failures VCF 9 uses a new licensing model. If license validation fails:\nConfirm you are using VCF 9-compatible licenses (2 license types: VCF cores and vSAN TiB) Pre-version 9 licenses are supported during upgrades but new deployments require VCF 9 licenses VCF Operations can operate in disconnected mode if the environment has no internet access Up Next Now that your management domain is running, the next post steps back before we get to workload domains: VCF Operations, VCF Automation, and the new fleet management appliance work together as a single control plane, and it\u0026rsquo;s worth understanding how that layer actually operates before building anything on top of it.\n","permalink":"https://virtualizationgurus.pages.dev/posts/vcf9-deployment-flow/","summary":"\u003ch2 id=\"introduction\"\u003eIntroduction\u003c/h2\u003e\n\u003cp\u003eDeploying VMware Cloud Foundation 9 is a structured process that, when followed correctly, results in a fully configured SDDC in a matter of hours. This guide walks through the complete deployment flow from planning through a functional management domain.\u003c/p\u003e\n\u003cblockquote\u003e\n\u003cp\u003e\u003cstrong\u003eNote:\u003c/strong\u003e VCF 9 replaces the Cloud Builder appliance used in VCF 5.x with the new \u003cstrong\u003eVCF Installer\u003c/strong\u003e: a new virtual appliance downloaded from the Broadcom Support portal that provides automated deployment and configuration workflows for the VCF environment. Unlike Cloud Builder which was a single-purpose OVA, the VCF Installer also supports VCF Converge (bringing existing vSphere infrastructure into VCF) and downloading/staging all required VCF component binaries.\u003c/p\u003e","title":"VCF 9 Deployment Flow: From Planning to a Running Management Domain"},{"content":"Overview VMware Cloud Foundation 9 (VCF 9) represents a significant evolution in how Broadcom delivers its software-defined data center stack. Released on June 17, 2025, VCF 9 unifies ESX 9.0, vCenter 9.0, NSX 9.0, and vSAN under a single platform version. In this post, we explore the core architectural changes that differentiate VCF 9 from its predecessors and understand why these changes matter for enterprise deployments.\nVCF 9 Architecture Overview The diagram below shows the high-level architecture of a VCF 9 environment, spanning the management domain and workload domains with all core components:\nManagement Domain Redesign One of the most impactful changes in VCF 9 is the redesigned management domain. Unlike previous versions, VCF 9 introduces a more flexible and unified approach:\nReduced minimum footprint: VCF 9 supports a 3-host management domain, lowering the barrier to entry for smaller deployments and edge/ROBO scenarios VCF Operations as the single management plane: VCF Operations replaces the SDDC Manager-centric model, providing a unified interface for fleet management, lifecycle operations, licensing, and cost management across all VCF instances Single NSX Manager support: VCF 9 introduces the option to deploy a single NSX Manager (instead of the traditional 3-node cluster) for resource-constrained environments. A 3-node NSX Manager cluster remains the recommended configuration for production high availability. Unified versioning: VCF 9.0 ships ESX 9.0, vCenter 9.0, and NSX 9.0 as a single, co-versioned platform, simplifying compatibility and lifecycle management NSX 9.0 Integration VCF 9 ships with NSX 9.0 as the default networking layer. Key improvements include:\nEnhanced Data Path (EDP) Standard as Default NSX 9.0 introduces Enhanced Data Path (EDP) Standard as the default host switch mode of operation for all new VCF Workload Domains. EDP Standard delivers superior performance in terms of throughput, packet rate, latency, and CPU utilization compared to the legacy Standard stack. Workload Domains upgraded or imported from VCF 5.x will continue using the legacy stack until the mode is explicitly changed.\nVirtual Private Cloud (VPC) Networking VCF 9 introduces Virtual Private Cloud (VPC) as the multi-tenancy networking model, replacing the need for manual NSX segment configuration for tenant isolation. VPCs are consumable from vCenter, VCF Automation, and vSphere Supervisor (VKS). Key VPC capabilities include:\nTransit Gateways (TGW): Centralized or distributed gateway hubs for inter-VPC and VPC-to-external routing, eliminating the need for tenants to configure infrastructure components directly VPC-Ready Workload Domains: All prerequisites for VPC consumption are fulfilled automatically when a workload domain is created Terraform support: The NSX Terraform Provider supports Transit Gateway, VPCs, and related constructs NSX VIBs Bundled with ESX A significant operational improvement in VCF 9: NSX kernel modules (VIBs) are now bundled with ESX 9.0 by default, removing the separate NSX VIB installation/upgrade step. This enables Live Patch for NSX transport node upgrades without requiring ESX maintenance mode.\nGateway Firewall Disabled by Default Starting with NSX 9.0, Gateway Firewall is automatically disabled by default for all new Tier-0 and Tier-1 Gateway deployments, improving performance and resource utilization for modern VPC-based network designs.\nvSAN ESA (Express Storage Architecture) VCF 9 continues with vSAN ESA (Express Storage Architecture) as the recommended storage architecture, with new capabilities:\nFeature OSA ESA NVMe Support Limited Full NVMe-native Compression Software only Hardware-assisted Snapshot Performance Degraded under load Near-zero impact Cache Drive Monitoring Basic Dying Disk Handling (DDH) for cache drives Disaster Recovery Basic VMware Live Recovery integration (RPO ≥ 1 min) New in VCF 9: vSAN ESA now supports Dying Disk Handling (DDH) for cache drives, enabling proactive detection and automated corrective actions when disk latency exceeds a predefined threshold.\nWorkload Domain Orchestration The workload domain creation and management process in VCF 9 has been significantly modernized:\nVCF Installer for deployment: The new VCF Installer virtual appliance (replacing Cloud Builder) orchestrates management domain bringup. It is downloaded from the Broadcom Support portal and deployed as a virtual appliance. VCF Operations for Day-2: VCF Operations provides a unified interface for fleet and domain management, lifecycle operations, license management, cost and capacity visibility, and security compliance, all from a single pane of glass VPC-Ready domains: Every new workload domain is provisioned VPC-ready, meaning networking isolation via NSX VPCs is available immediately upon domain creation API-first design: The VCF SDK (with Python and Java bindings) and PowerCLI provide comprehensive automation capabilities for all VCF operations What This Means for Your Deployment If you are planning a new VCF deployment or evaluating an upgrade from VCF 4.x or 5.x, here are the key takeaways:\nSmaller entry point: 3-host management domains and single NSX Manager option make VCF 9 accessible for smaller environments and branch offices Better storage performance with ESA: vSAN ESA delivers dramatically better performance for mixed I/O workloads, with new DDH protection for cache drives Modern networking with VPCs: The VPC model provides cloud-native tenant isolation without complex manual NSX segment management Operational efficiency: VCF Operations consolidates all management, licensing, cost visibility, and lifecycle tasks into a single interface EDP performance gains: Enhanced Data Path Standard is now the default, delivering superior network performance for all new workload domains Simplified NSX lifecycle: NSX VIBs bundled with ESX and Live Patch support eliminate the operational complexity of separate NSX upgrade cycles Next Steps In our next post, we will walk through the complete VCF 9 deployment flow, from initial planning through a working management domain using the new VCF Installer. Stay tuned!\n","permalink":"https://virtualizationgurus.pages.dev/posts/vcf9-architecture-deep-dive/","summary":"\u003ch2 id=\"overview\"\u003eOverview\u003c/h2\u003e\n\u003cp\u003eVMware Cloud Foundation 9 (VCF 9) represents a significant evolution in how Broadcom delivers its software-defined data center stack. Released on June 17, 2025, VCF 9 unifies ESX 9.0, vCenter 9.0, NSX 9.0, and vSAN under a single platform version. In this post, we explore the core architectural changes that differentiate VCF 9 from its predecessors and understand why these changes matter for enterprise deployments.\u003c/p\u003e\n\u003ch2 id=\"vcf-9-architecture-overview\"\u003eVCF 9 Architecture Overview\u003c/h2\u003e\n\u003cp\u003eThe diagram below shows the high-level architecture of a VCF 9 environment, spanning the management domain and workload domains with all core components:\u003c/p\u003e","title":"VCF 9 Architecture Deep Dive: What Changed and Why It Matters"},{"content":"About This Blog Virtualization Gurus is a technical blog focused on enterprise virtualization, VMware Cloud Foundation (VCF), and modern data center technologies. Our content is written by practitioners with hands-on experience designing, deploying, and operating VCF environments at scale.\nWhat We Cover Our primary focus areas include:\nVMware Cloud Foundation (VCF): Architecture, deployment, operations, and lifecycle management vSphere \u0026amp; vSAN: Deep dives into the core compute and storage stack NSX: Network virtualization, microsegmentation, and NSX Advanced Load Balancer Aria Suite: Operations, Automation, and Log Insight Multi-cloud: VCF on-premises connecting to public cloud environments Content Philosophy Every post on this blog is based on real-world experience. We avoid generic overviews in favor of actionable technical content: the kind of detail you need when you are actually building or troubleshooting a VCF environment.\nAbout the Author Mohamed Rabiee Senior Cloud Infrastructure Consultant, evoila I\u0026rsquo;m a Senior Cloud Infrastructure Consultant specializing in VMware Cloud Foundation (VCF) and multi-cloud architecture. I\u0026rsquo;ve delivered 50+ VCF deployments and upgrade engagements for enterprise customers, moving from VMware technical support through a Site Reliability Engineer role at VMware, Pod Lead on 50+ VCF upgrade engagements, to multi-cloud consulting (OCI, GCP) and now partner-level SDDC advisory work at evoila. This blog is where I write up the architecture details, upgrade gotchas, and design decisions from that work, VCF 9 in particular.\nViews and content on this blog are my own, based on personal experience, and do not represent the views of any employer, past or present.\nCommunity Recognition Named a VMware vExpert 2026, the Broadcom community program that recognizes practitioners who share their VMware knowledge publicly.\nCertifications VCP - VCF Architect VCP - VCF Administrator VCAP - VCF VKS VCAP - VCF Networking VCF Certified Specialist vSAN Certified Specialist VCIX - DCV Double VCP - DCV \u0026amp; NV Broadcom PP - VKS Architecture Broadcom PP - VKS Implementation Broadcom PP - VKS Pre-Sales Broadcom PP - VKS Support Broadcom PP - VCF Networking Architecture Broadcom PP - VCF Networking Implementation Broadcom PP - VCF Networking Pre-Sales Broadcom PP - VCF Networking Support Broadcom CE - VCF Implementation Broadcom CE - VKS Implementation Broadcom CE - VCF Networking Implementation Broadcom CE - VCF Sales Broadcom PP - VCF Sales Broadcom PP - VCF Implementation vSphere 6.7 Foundations VCP - Network Virtualization VCAP - DCV Design VCAP - DCV Deploy VCP - Data Center Virtualization All badges verified on Credly. Click any badge to view its verification page.\nContact Have a topic you want covered? Encountered an issue with VCF that you cannot find documentation for? Reach out via GitHub Issues on the blog repository.\nReach out directly! ✉ Email mohamed.amr.rabiee@gmail.com 👥 LinkedIn Mohamed Rabiee\n","permalink":"https://virtualizationgurus.pages.dev/about/","summary":"\u003ch2 id=\"about-this-blog\"\u003eAbout This Blog\u003c/h2\u003e\n\u003cp\u003eVirtualization Gurus is a technical blog focused on enterprise virtualization, VMware Cloud Foundation (VCF), and modern data center technologies. Our content is written by practitioners with hands-on experience designing, deploying, and operating VCF environments at scale.\u003c/p\u003e\n\u003ch2 id=\"what-we-cover\"\u003eWhat We Cover\u003c/h2\u003e\n\u003cp\u003eOur primary focus areas include:\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e\u003cstrong\u003eVMware Cloud Foundation (VCF)\u003c/strong\u003e: Architecture, deployment, operations, and lifecycle management\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003evSphere \u0026amp; vSAN\u003c/strong\u003e: Deep dives into the core compute and storage stack\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eNSX\u003c/strong\u003e: Network virtualization, microsegmentation, and NSX Advanced Load Balancer\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eAria Suite\u003c/strong\u003e: Operations, Automation, and Log Insight\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eMulti-cloud\u003c/strong\u003e: VCF on-premises connecting to public cloud environments\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch2 id=\"content-philosophy\"\u003eContent Philosophy\u003c/h2\u003e\n\u003cp\u003eEvery post on this blog is based on real-world experience. We avoid generic overviews in favor of actionable technical content: the kind of detail you need when you are actually building or troubleshooting a VCF environment.\u003c/p\u003e","title":"About Virtualization Gurus"},{"content":"A practitioner-level walkthrough of VMware Cloud Foundation 9, one focused topic per post, each one validated against the official Broadcom TechDocs. This is a 21-part series, and 20 posts are live. Published VCF 9 Architecture Deep Dive: What Changed and Why It Matters, June 15, 2026 VCF 9 Deployment Flow: From Planning to a Running Management Domain, June 22, 2026 VCF 9 Fleet Management Layer: One Control Plane for Every Instance, July 3, 2026 VCF 9 Workload Domain Planning and Deployment, July 6, 2026 VCF 9 NSX VPC Deep Dive: Cloud-Native Networking for Your Private Cloud, July 8, 2026 VCF 9 vSAN ESA Deep Dive: Architecture, Performance, and New Features, July 11, 2026 VCF 9 Lifecycle Management with VCF Operations: Unified Upgrades and Fleet Management, July 14, 2026 VCF 9 Instance Model: Designing HQ, DR, and Edge/Sovereign Topologies, July 17, 2026 VCF 9 Management Domain Anatomy: What Actually Runs Inside, July 20, 2026 Physical Network Design: VDS Separation, ToR Switches, and BGP Uplinks, July 24, 2026 Workload Domain Creation: Greenfield vs Import Existing vCenter, July 25, 2026 VI Workload Domains: Shared vs Dedicated NSX, August 2, 2026 NSX Edge Cluster Deep Dive: Tier-0/Tier-1 Gateways, VPN, and North-South Firewall Design, August 9, 2026 vSAN ESA vs OSA: Storage Architecture Decisions, August 16, 2026 VCF 9 Security and Compliance: DFW, VPC Isolation, and Hardened Operations, August 23, 2026 VCF 9 Identity Broker: Retiring VMware Identity Manager for Unified Fleet Authentication, August 30, 2026 Bundle Management: Online vs Offline/Air-Gapped, September 6, 2026 Day-2 Operations: Lifecycle, Patching, and Compliance via VCF Operations, September 26, 2026 DR \u0026 Ransomware Recovery: Isolated Recovery, Protection and Recovery, and VPC Isolation, October 3, 2026 Private AI Workload Domain: GPU Nodes, AI Kubernetes, and Private AI Services, October 11, 2026 Coming up Advanced Services for VCF: VPC, Load Balancing, and Network Observability, October 19, 2026 Check back here for the running index, or subscribe via RSS to catch new posts as they publish.\n","permalink":"https://virtualizationgurus.pages.dev/vcf9-series/","summary":"\u003cp\u003eA practitioner-level walkthrough of VMware Cloud Foundation 9, one focused topic per post, each one validated against the official Broadcom TechDocs. This is a 21-part series, and 20 posts are live.\n\u003c/p\u003e\n\u003ch2 id=\"published\"\u003ePublished\u003c/h2\u003e\n\u003col\u003e\n  \u003cli\u003e\u003ca href=\"/posts/vcf9-architecture-deep-dive/\"\u003eVCF 9 Architecture Deep Dive: What Changed and Why It Matters\u003c/a\u003e, June 15, 2026\u003c/li\u003e\n  \u003cli\u003e\u003ca href=\"/posts/vcf9-deployment-flow/\"\u003eVCF 9 Deployment Flow: From Planning to a Running Management Domain\u003c/a\u003e, June 22, 2026\u003c/li\u003e\n  \u003cli\u003e\u003ca href=\"/posts/vcf9-fleet-management-layer/\"\u003eVCF 9 Fleet Management Layer: One Control Plane for Every Instance\u003c/a\u003e, July 3, 2026\u003c/li\u003e\n  \u003cli\u003e\u003ca href=\"/posts/vcf9-workload-domain-planning/\"\u003eVCF 9 Workload Domain Planning and Deployment\u003c/a\u003e, July 6, 2026\u003c/li\u003e\n  \u003cli\u003e\u003ca href=\"/posts/vcf9-nsx-vpc-deep-dive/\"\u003eVCF 9 NSX VPC Deep Dive: Cloud-Native Networking for Your Private Cloud\u003c/a\u003e, July 8, 2026\u003c/li\u003e\n  \u003cli\u003e\u003ca href=\"/posts/vcf9-vsan-esa-deep-dive/\"\u003eVCF 9 vSAN ESA Deep Dive: Architecture, Performance, and New Features\u003c/a\u003e, July 11, 2026\u003c/li\u003e\n  \u003cli\u003e\u003ca href=\"/posts/vcf9-lifecycle-management/\"\u003eVCF 9 Lifecycle Management with VCF Operations: Unified Upgrades and Fleet Management\u003c/a\u003e, July 14, 2026\u003c/li\u003e\n  \u003cli\u003e\u003ca href=\"/posts/vcf9-instance-model/\"\u003eVCF 9 Instance Model: Designing HQ, DR, and Edge/Sovereign Topologies\u003c/a\u003e, July 17, 2026\u003c/li\u003e\n  \u003cli\u003e\u003ca href=\"/posts/vcf9-management-domain-anatomy/\"\u003eVCF 9 Management Domain Anatomy: What Actually Runs Inside\u003c/a\u003e, July 20, 2026\u003c/li\u003e\n  \u003cli\u003e\u003ca href=\"/posts/vcf9-physical-network-design/\"\u003ePhysical Network Design: VDS Separation, ToR Switches, and BGP Uplinks\u003c/a\u003e, July 24, 2026\u003c/li\u003e\n  \u003cli\u003e\u003ca href=\"/posts/vcf9-workload-domain-creation-greenfield-vs-import/\"\u003eWorkload Domain Creation: Greenfield vs Import Existing vCenter\u003c/a\u003e, July 25, 2026\u003c/li\u003e\n  \u003cli\u003e\u003ca href=\"/posts/vcf9-vi-workload-domains-shared-vs-dedicated-nsx/\"\u003eVI Workload Domains: Shared vs Dedicated NSX\u003c/a\u003e, August 2, 2026\u003c/li\u003e\n  \u003cli\u003e\u003ca href=\"/posts/vcf9-nsx-edge-cluster-deep-dive/\"\u003eNSX Edge Cluster Deep Dive: Tier-0/Tier-1 Gateways, VPN, and North-South Firewall Design\u003c/a\u003e, August 9, 2026\u003c/li\u003e\n  \u003cli\u003e\u003ca href=\"/posts/vcf9-vsan-esa-vs-osa/\"\u003evSAN ESA vs OSA: Storage Architecture Decisions\u003c/a\u003e, August 16, 2026\u003c/li\u003e\n  \u003cli\u003e\u003ca href=\"/posts/vcf9-security-compliance/\"\u003eVCF 9 Security and Compliance: DFW, VPC Isolation, and Hardened Operations\u003c/a\u003e, August 23, 2026\u003c/li\u003e\n  \u003cli\u003e\u003ca href=\"/posts/vcf9-identity-broker/\"\u003eVCF 9 Identity Broker: Retiring VMware Identity Manager for Unified Fleet Authentication\u003c/a\u003e, August 30, 2026\u003c/li\u003e\n  \u003cli\u003e\u003ca href=\"/posts/vcf9-bundle-management/\"\u003eBundle Management: Online vs Offline/Air-Gapped\u003c/a\u003e, September 6, 2026\u003c/li\u003e\n  \u003cli\u003e\u003ca href=\"/posts/vcf9-day2-operations/\"\u003eDay-2 Operations: Lifecycle, Patching, and Compliance via VCF Operations\u003c/a\u003e, September 26, 2026\u003c/li\u003e\n  \u003cli\u003e\u003ca href=\"/posts/vcf9-dr-ransomware-recovery/\"\u003eDR \u0026 Ransomware Recovery: Isolated Recovery, Protection and Recovery, and VPC Isolation\u003c/a\u003e, October 3, 2026\u003c/li\u003e\n  \u003cli\u003e\u003ca href=\"/posts/vcf9-private-ai-workload-domain/\"\u003ePrivate AI Workload Domain: GPU Nodes, AI Kubernetes, and Private AI Services\u003c/a\u003e, October 11, 2026\u003c/li\u003e\n\u003c/ol\u003e\n\n\u003ch2 id=\"coming-up\"\u003eComing up\u003c/h2\u003e\n\n\u003col start=\"21\"\u003e\u003cli\u003eAdvanced Services for VCF: VPC, Load Balancing, and Network Observability, October 19, 2026\u003c/li\u003e\u003c/ol\u003e\n\n\u003cp\u003eCheck back here for the running index, or \u003ca href=\"/index.xml\"\u003esubscribe via RSS\u003c/a\u003e to catch new posts as they publish.\u003c/p\u003e","title":"VCF 9 Series"}]