Skip to content
WindowsDeep Dive Published Updated 7 min readViews unavailable

Windows Shielded VM and Guarded Fabric Operations: Attestation, Key Custody, and Recovery

Operate a Windows guarded fabric by validating HGS attestation, host authorization, shielded VM protection, key custody, and recovery dependencies.

A guarded fabric is a Hyper-V environment in which the Host Guardian Service (HGS) decides whether a host satisfies attestation requirements and may obtain the key material needed to start protected virtual machines. Shielded VMs add protections for tenant workloads against inspection or tampering by fabric administrators, depending on the VM’s protection mode, template, key protector, and the guarded-fabric design. The assurance is not produced by a single checkbox. It depends on host integrity, attestation configuration, HGS availability, certificate/private-key custody, VM configuration, migration compatibility, and rehearsed recovery.

This article is an operations runbook for an existing environment. It is not a deployment recipe: initial HGS and guarded-host deployment must follow the Microsoft guide for the exact Windows Server release and attestation mode. Before production deployment, apply the current cumulative update and verify compatibility matrices. Do not assume a standard Hyper-V VM is shielded because it has a virtual TPM, encryption enabled, or a secure-boot setting; verify its protection state and key protector explicitly.

Distinguish VM protection layers

Hyper-V virtual machine encryption, a virtual TPM, secure boot, and shielding address different properties. A vTPM can protect guest secrets and support BitLocker within a VM; encryption can protect selected VM state; shielding adds guarded-fabric controls intended to restrict fabric-admin access to protected guest state and keys. The exact guarantees depend on generation, configuration, template, policy, guest OS, and host support. Model the adversary and document which operations a fabric administrator should or should not be able to perform before selecting a protection mode.

Shielded VM provisioning normally involves a trusted template disk and shielding data that specify the approved owner and fabric protections. The VM’s key protector binds startup to the configured HGS trust model. A guardian or fabric signing/encryption key can be critical to VM portability and recovery. Protect all private key copies with strict role separation, offline or otherwise appropriately protected backup, access logging, and periodic restore tests. If a required key is lost or compromised, treat it as a recovery/security incident, not a routine certificate renewal.

Inventory HGS mode, hosts, and workload state

Capture the HGS server version, configured attestation mode, HGS cluster health, encryption and signing certificate thumbprints and validity, guarded host inventory, host software/firmware compliance policy, VM protection type, key protector, and migration targets. Store the inventory in the restricted operations system. Do not export private keys or shielding data to general-purpose tickets. An HGS object that reports healthy does not by itself prove every host can attest or that every VM has recoverable key material.

Run the following read-only commands on an HGS management context and a Hyper-V host with the appropriate modules. Confirm cmdlet availability and permissions for the installed release before operational use.

Import-Module HostGuardian -ErrorAction Stop
Import-Module Hyper-V -ErrorAction Stop

Get-HgsServer
Get-HgsClientConfiguration
Get-VM | Select-Object Name, State, Generation
Get-VMSecurity -VMName 'APP01'

The exact property set available can vary by module/version, so treat the output as an inventory starting point and check the specific VM protection cmdlets documented for your server release. Compare host attestation policy with the approved baseline and confirm the HGS URLs resolved by each host. A host can fail attestation because of a legitimate firmware or code-integrity change; do not permanently relax policy simply to get a workload running.

Diagnose attestation failures from evidence

When a host cannot attest, identify the precise HGS error and whether the fabric uses TPM-trusted, host-key, or admin-trusted attestation. Validate DNS and network reachability to the HGS endpoints, certificate expiration and trust, HGS cluster state, host identity, required roles/features, secure boot and TPM prerequisites for the selected mode, code integrity policy, and recent patch/firmware changes. Use the supported guarded-fabric diagnostics on both HGS and the Hyper-V host, then correlate HGS, Hyper-V, and system logs by timestamp.

Avoid mode switching during an incident without a reviewed change plan. HGS attestation modes provide different assurance properties and compatibility. Moving from TPM-based validation to an administrative trust model can alter the threat boundary and create new dependencies. If a planned mode change is necessary, follow the Microsoft procedure, inventory trusted host groups and trusts, test in a nonproduction fabric, update every host consistently, and document how to revert without weakening the security baseline.

Plan key rotation and certificate lifecycle

HGS uses signing and encryption certificates as part of its guarded-fabric trust. Identify all certificates, private key locations, expiration, renewal owner, dependent VMs, and recovery copies before rotation. Keep old and new keys available during a supported transition if Microsoft’s procedure requires rekeying or dual-key overlap. Do not delete old keys just because a replacement certificate has been installed: an offline VM or a disaster-restored VM may still depend on the prior key protector.

For each private key, define who can access it, how an emergency access is approved, how use is audited, and how the key is restored after HGS disaster recovery. Ensure backups are encrypted, access is separated from routine host administration, and restoration is practiced in an isolated test environment. A loss of the only signing/encryption private-key copies can make protected workloads unrecoverable; a compromise can undermine the fabric’s protection claims. Document both risks and the response path.

Validate VM mobility and maintenance before a change

Before patching, rebooting, or replacing a guarded host, prove that affected shielded VMs can run on the remaining authorized hosts and that each host can reach a healthy HGS instance. Test the actual live migration or planned failover workflow for the VM’s protection type and cluster configuration. Verify that destination hosts meet the attestation policy and that the VM’s key protector allows the intended movement. A generic Hyper-V migration test with an unshielded test VM does not establish that protected workloads can move.

Check cluster health, storage paths, network placement, HGS reachability, certificate state, and spare host capacity. Change one dependency at a time and watch both HGS attestation and guest service health. Keep maintenance windows large enough to restore the host or move workloads back without rushing a change to the trust configuration. For host replacement, enroll and attest the new host before moving production VMs; do not trust a machine solely because it is in the same rack or cluster.

Troubleshoot a VM that cannot start

Separate a guest boot error from a key-protector or HGS denial. Record the VM’s protection state and exact error, host identity and attestation state, HGS endpoint selected, certificate status, VM configuration path, cluster ownership, and recent changes. If the host itself is not attested, debug host trust first. If the host is attested but one VM fails, validate that the VM’s shielding data and key protector are consistent with the current guardian set and that the host is authorized to obtain the required key.

Do not bypass shielded protections by copying VM state to an untrusted host, exporting secrets, changing a key protector blindly, or restoring an old host snapshot. Use the documented repair procedure for the actual error and retain evidence. If a required certificate/private key is missing, stop before attempting destructive reconfiguration; engage the key-custody and recovery owners and use the tested backup. Any temporary exception must be formally approved, scoped to a specific VM/host, time-bounded, and monitored.

Recovery and acceptance criteria

Recovery planning must include HGS system-state or supported backup data, configuration, certificate/private-key material, guardian metadata, host attestation policy, VM shielding data, and the protected VM storage. Test recovery in isolation without exposing tenant data. Document the ordering of HGS restoration and host configuration for the relevant product version; do not assume a generic cluster backup is sufficient. Validate that restored HGS endpoints, certificate thumbprints, trust mode, and guarded hosts agree before starting production workloads.

Acceptance for a production change includes: HGS cluster health; the intended attestation mode; successful attestation by each relevant host; valid certificate chains with expiry monitoring; protected VM state confirmed through supported tools; migration or failover tested between authorized hosts; backup and key recovery evidence; and no unauthorized change to key or shielding policy. Retain change IDs and test evidence, but store sensitive shielding data and key material only in the approved protected repository.

Keep the trust boundary current

Review Windows Server support and compatibility documentation before upgrades, maintain current cumulative updates, and reassess host firmware, code integrity, and certificate lifecycles. Remove retired hosts and stale trust entries using supported procedures. Alert on new HGS trusted hosts, policy changes, certificate/key access, failed attestation spikes, and unplanned changes to shielding data. A guarded fabric only remains trustworthy if its trust and recovery dependencies are maintained as actively as its compute cluster.

Related:

Sources:

Comments