# Prerequisites

> Before you deploy Self-Hosted Deployment

## Access to the container registry

Ensure that you have credentials for the Unity private container registry `uccmpprivatecloud.azurecr.io`. Unity provides these credentials. You need this information to download shdctl, pull release packages, and sync container images and ORAS artifacts to your private registry.

## Tooling

Ensure that you have this tooling:

* Access to a terminal, and basic knowledge of the command line

* One of these tools if you use Microsoft Windows:

  * Windows Subsystem for Linux (WSL)
  * Git for Windows

* [ORAS CLI](https://oras.land/docs/installation), for the shdctl install script, which downloads shdctl from the registry

* [Helm](https://helm.sh/docs/intro/install/) (version 3.9 or later), to install the Helm charts

* [kubectl](https://kubernetes.io/docs/tasks/tools/), to interact with your Kubernetes cluster

* [Docker](https://docs.docker.com/get-docker/), to sync the container images to your private registry

* [ArgoCD](https://argo-cd.readthedocs.io/en/stable/getting_started/) (recommended), for continuous delivery via GitOps. If you choose the ArgoCD deployment method, install it in the target cluster before you deploy: shdctl applies a single bootstrap `Application` and ArgoCD takes over from there. Every application is created in the deployment namespace, so install ArgoCD in that namespace, or configure it to watch it: ArgoCD's "Applications in any namespace" needs the namespace in both its `application.namespaces` setting and the `default` project's `sourceNamespaces`. An ArgoCD that watches only its own namespace leaves the applications unsynced, and reports no error. If you deploy with Helm, you don't need ArgoCD at all.

* **shdctl** is the Unity CLI tool that manages the entire deployment lifecycle. Its install steps, requirements, and registry credential setup are in [Install shdctl](../../shdctl/install). If you used vpctl before, install shdctl 1.0.0 or later: vpctl can't pull Self-Hosted Deployment 2.0.0 or later.

> **Note:**
>
> Unity recommends ArgoCD as the deployment method. ArgoCD handles CRD installation ordering and detects configuration drift automatically, which makes ongoing operations more reliable. Helm is also fully supported if you prefer direct deployments without a GitOps workflow.

## Kubernetes

The deployment requires a Kubernetes cluster that you manage.

Use Kubernetes version 1.34 or later.

Your cluster must include at least **3 schedulable nodes** in the general workloads pool. MongoDB (Percona Server for MongoDB) and RabbitMQ run as 3-replica clusters by default and place each replica on a different node by using `podAntiAffinity` on `kubernetes.io/hostname`. PostgreSQL and Garage object storage also run 3 replicas by default, but don't require separate nodes. If your cluster has only two nodes, some pods remain in the `Pending` state because the scheduler can't satisfy the anti-affinity requirement.

### Storage classes

The cluster must provide two storage classes:

* A **default storage class** for general-purpose persistent volumes (block storage), marked as the cluster default with the `storageclass.kubernetes.io/is-default-class: "true"` annotation. For example, a local-path provisioner, a SAN-backed CSI driver, or any block storage provisioner.
* A **ReadWriteMany (RWX) storage class** for shared volumes that multiple pods can mount simultaneously. For example, an NFS provisioner, a distributed filesystem, or any CSI driver that supports the `ReadWriteMany` access mode.

In the manifest file, `configuration.kubernetes.storage.readWriteManyStorageClass` names the RWX class. `defaultStorageClass` names the class for Garage and the Unity Licensing Server only: every other ReadWriteOnce volume uses the cluster's default storage class, so the cluster must have one even if the manifest names a class.

### Metrics API

The cluster must serve the `metrics.k8s.io` API, which the platform's horizontal pod autoscalers require. Most distributions provide it through `metrics-server`. Without it, autoscaled workloads can't read their resource usage and don't scale.

### Namespaces

Services come preconfigured for use within a single Kubernetes namespace. You configure the target namespace in the manifest file.

Create the namespace before you deploy. The deployment doesn't create it for you.

### Node pools

The platform separates general application workloads from asset transformation jobs. Configure your cluster with three node pools (or equivalent node groups):

| Node pool               | Node label                    | Taint                                    | Recommended node size                                                                     | Purpose                                                                                                               |
| ----------------------- | ----------------------------- | ---------------------------------------- | ----------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------- |
| General workloads       | None required                 | None                                     | 8–32 vCPUs; compute, general-purpose, or memory-optimized instances                       | Application services, databases, messaging                                                                            |
| Transformations         | `aks-node-pool=argocpu`       | None                                     | 8–32 vCPUs and at least 64 GiB of memory; general-purpose or memory-optimized instances   | Transformation workflow pods                                                                                          |
| Transformations (large) | `aks-node-pool=argocpu-large` | `aks-node-pool=argocpu-large:NoSchedule` | 32–64 vCPUs and at least 256 GiB of memory; general-purpose or memory-optimized instances | Escalation pool for memory-intensive transformations; only pods that explicitly tolerate the taint are scheduled here |

Transformation steps request up to 32 GiB of memory on the `argocpu` pool, and their retries up to 220 GiB on the `argocpu-large` pool. A step that no node can fit stays `Pending` until its workflow times out.

The exact mechanism to create node pools depends on your Kubernetes distribution (for example, static node labels, a node autoscaler, or a cluster API provider). If an autoscaler other than Karpenter scales the transformation pool from zero, `shdctl cluster check` reports the pool as missing while it is empty; that failure is expected.

GPU nodes are optional. By default, transformation actions that can use a GPU run on the `argocpu` pool with a software renderer. If your cluster has GPU nodes that advertise the `nvidia.com/gpu` resource, for example through the NVIDIA device plugin, set `configuration.overrides.automation-manager.values.env.vpcGpuAvailable: true` in the manifest so that those actions request a GPU.

### Network policies

During deployment to an existing cluster, you may need to control the flow of network traffic by using network policies. You can deploy most Kubernetes resources in a single namespace, which you can use to scope network isolation from other services that run in your cluster.

### Clusters behind an HTTP proxy

If your AKS cluster sends its outbound traffic through an HTTP proxy ([AKS HTTP proxy](https://learn.microsoft.com/azure/aks/http-proxy)), AKS injects `HTTP_PROXY`, `HTTPS_PROXY` and `NO_PROXY` into every pod. Platform services call each other by their short Service names, which no `NO_PROXY` entry matches, so those calls go to the proxy and fail.

To keep the platform pods off the proxy, add the AKS opt-out annotation to every pod in the manifest file. Nodes still pull images through the proxy, but the pods get no internet access:

```yaml
configuration:
  kubernetes:
    podAnnotations:
      kubernetes.azure.com/no-http-proxy-vars: "true"
```

The automation-manager job pulls the release manifest from your registry itself, so it still needs the proxy. Unless your cluster reaches the registry without the proxy, pass it through the chart's overrides:

```yaml
configuration:
  overrides:
    automation-manager:
      values:
        env:
          httpsProxy: http://<proxy>:<port>
          noProxy: .svc,.cluster.local,localhost,127.0.0.1
```

`configuration.kubernetes.podAnnotations` applies to every pod the release creates. To change or remove an annotation for a single chart, refer to [Per-chart Helm value overrides](../../shdctl/manifest#per-chart-helm-value-overrides).

The `azure-baseline-example` in the release package can create such a cluster: `aks_http_proxy` sets the proxy, `aks_private_cluster_enabled` makes the API server private, and `aks_egress_via_proxy_only` denies all other outbound traffic from the cluster's subnet. Its README describes each variable.

#### What the proxy must allow

* **AKS's own outbound traffic.** Nodes bootstrap and keep running through the proxy, so it must allow the destinations in [AKS outbound network and FQDN rules](https://learn.microsoft.com/azure/aks/outbound-rules-control-egress), for example `mcr.microsoft.com` and `packages.aks.azure.com`.
* **Your container registry.** Nodes pull every image through the proxy, and the automation-manager job reaches the registry through it too: the registry in `configuration.kubernetes.docker.repository`, or `uccmpprivatecloud.azurecr.io` if you pull from it directly. Allow the endpoints the registry serves image layers from as well; for Azure Container Registry, refer to [Configure rules to access an Azure container registry behind a firewall](https://learn.microsoft.com/azure/container-registry/container-registry-firewall-access-rules).
* **Requests from the nodes.** Pod traffic reaches the proxy from its node's address, so allow the node subnet as a source.
* **HTTP and HTTPS only.** HTTPS goes through the proxy as `CONNECT` on port 443, and plain HTTP on port 80. Outbound traffic that isn't HTTP, such as SMTP or SSH, has no path through the proxy.
* **TLS inspection.** If the proxy decrypts HTTPS, give the cluster its CA certificate (`trustedCa` in the AKS HTTP proxy configuration). The certificates must use subject alternative names.

## Container registry

By default, the platform pulls container images directly from the Unity source registry (`uccmpprivatecloud.azurecr.io`). This is the simplest setup and doesn't require a separate registry.

If your environment is air-gapped or if you need full control over artifact distribution, you can mirror artifacts to your own private container registry. Any OCI-compliant registry is supported, for example Harbor, JFrog Artifactory, or a cloud-managed registry such as Amazon ECR or Azure Container Registry. Use these commands to mirror artifacts from the Unity source registry to your private registry:

* `shdctl artifact sync images`: this command syncs the Docker container images.
* `shdctl artifact sync oras`: this command syncs the ORAS artifacts, for example, the Pixyz workflow templates.
* `shdctl artifact sync verify`: this command checks that your registry holds every image your manifest needs.

## System requirements

Ensure that you have these elements:

* A valid hostname that can be updated to point to the IP of the load balancer
* A valid Unity Version Control (UVCS) license to run the UVCS server
* A valid Unity Asset Transformer SDK license to run transformations on assets

  A valid Unity Asset Transformer license is required for all users who perform asset transformations. The platform uses static licenses, so a floating license server is not required.

## Validate your cluster

Once your `manifest.yaml` exists, check the prerequisites on this page automatically. The command reads the namespace, the storage classes, and the ArgoCD settings from the manifest, and step 3 of the [deployment](./deployment.md) runs it:

```sh
shdctl cluster check
```

The command verifies:

* The Kubernetes version.
* The number of schedulable nodes.
* The transformation node pools.
* The storage classes.
* The target namespace.
* The metrics API.
* The in-cluster ArgoCD installation if your manifest configures ArgoCD. On the Helm deployment path, these checks are skipped.

The command reads from the cluster and doesn't change it. It prints a remediation hint for anything that doesn't pass.

Add `--probe-storage` to verify that each storage class can provision a volume, including `ReadWriteMany` support.

For the full list of checks and options, refer to [shdctl cluster command](../../shdctl/commands/cluster).

## Next steps

[Deploy Self-Hosted Deployment](./deployment.md)
