# shdctl cluster command

> Validate that your Kubernetes cluster meets the deployment prerequisites before you deploy

The `cluster check` command validates your manifest, and then verifies that the cluster your current kubeconfig context points at meets the deployment prerequisites.

```sh
shdctl cluster check
```

The command requires a manifest, because it reads the target namespace, the storage class names, and the deployment method from it. It doesn't require a pulled release, so you can run it before you download the release package. It looks for `manifest.yaml` in the current directory, then in each parent directory. To use a manifest that this auto-discovery won't find, pass `--manifest`: refer to [Global flags](./_index.md#global-flags).

The command is read-only by default: against the cluster, it runs only `kubectl get` and `kubectl version` commands.

## Cluster check reference

The following table lists each check the command runs, in order, and the condition it verifies.

| Check                            | What it verifies                                                                                                                                                                                                                                                           |
| -------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `kubectl available`              | `kubectl` is on your `PATH`                                                                                                                                                                                                                                                |
| `cluster reachable`              | the current kubeconfig context's API server answers                                                                                                                                                                                                                        |
| `kubernetes version`             | the server runs Kubernetes 1.34 or later                                                                                                                                                                                                                                   |
| `general worker nodes`           | at least three nodes are Ready, not cordoned, and carry no `NoSchedule` or `NoExecute` taint — one node with `configuration.infrastructure.singleNode: true`                                                                                                               |
| `transformations pool`           | a schedulable, untainted node is labeled `aks-node-pool=argocpu`, or a Karpenter NodePool provisions one. Transformation workflows select that label and stay Pending without it                                                                                           |
| `large transformations pool`     | nodes labeled `aks-node-pool=argocpu-large` carry the matching `NoSchedule` taint and one is Ready — or, with none at rest, a Karpenter NodePool provisions them with that taint declared. Warns rather than fails: only memory-intensive transformation retries run there |
| `default storage class`          | the cluster has a default storage class, and the class `configuration.kubernetes.storage.defaultStorageClass` names, if any, exists (see below)                                                                                                                            |
| `rwx storage class`              | `configuration.kubernetes.storage.readWriteManyStorageClass` is set and exists; only `--probe-storage` proves it provisions ReadWriteMany volumes                                                                                                                          |
| `namespace`                      | `configuration.kubernetes.namespace` exists. Nothing in the release creates it, so create it before the first deploy                                                                                                                                                       |
| `metrics api`                    | the `v1beta1.metrics.k8s.io` APIService is available; warns otherwise. Horizontal pod autoscaling requires it                                                                                                                                                              |
| `argocd application crd`         | the ArgoCD `Application` CRD is installed — checked for the `argocd` format only                                                                                                                                                                                           |
| `argocd controller`              | an `argocd-application-controller` pod is Ready — checked for the `argocd` format only, and skipped when your kubeconfig can't list pods in every namespace                                                                                                                |
| `node ipv6 addresses`            | a node reports an IPv6 InternalIP — checked, as a warning, only when `configuration.networking.ipFamily` is `ipv6`                                                                                                                                                         |
| `storage probe (default class)`  | with `--probe-storage` only: a 1-GiB test PVC binds on the cluster's default class                                                                                                                                                                                         |
| `storage probe (manifest class)` | with `--probe-storage` only: a 1-GiB test PVC binds on `defaultStorageClass`, when that is a different class                                                                                                                                                               |
| `storage probe (rwx class)`      | with `--probe-storage` only: a 1-GiB `ReadWriteMany` test PVC binds on `readWriteManyStorageClass`                                                                                                                                                                         |

If a node pool has no running node, the command looks for a matching Karpenter `NodePool` instead. Clusters that scale transformation capacity from zero with Karpenter therefore pass the node pool checks even when no node is running. If Karpenter isn't installed, or your credentials can't read its resources, the command ignores the lookup, because Karpenter is optional: a pool that another autoscaler scales from zero reports as missing while it is empty.

Most checks read cluster-scoped objects — nodes, storage classes, the namespace, the metrics APIService, the Application CRD. A kubeconfig denied the node, storage-class, namespace or CRD read fails those checks, so a namespace-scoped one fails the command: run it with a kubeconfig that can read cluster-scoped objects.

The storage-class check fails when the cluster has no default storage class, even if the manifest names `defaultStorageClass`: that field only applies to garage and the license server, and every other ReadWriteOnce volume uses the cluster default (ReadWriteMany volumes use `readWriteManyStorageClass`). It warns when the named class exists but isn't the cluster default.

## Read the results

Each check reports one of these statuses:

| Symbol | Status  | Effect on the exit code                  |
| ------ | ------- | ---------------------------------------- |
| `✔`    | Pass    | None.                                    |
| `⚠`    | Warning | None.                                    |
| `✘`    | Failure | The command exits with a nonzero status. |
| `-`    | Skipped | None.                                    |

Failed and warning checks also print a remediation hint. Because the command exits with a nonzero status when any check fails, you can use it to gate a CI pipeline before the deployment steps. When `kubectl` is missing or the cluster is unreachable, every check after that one is skipped.

The output looks like the following example:

```text
Context: my-cluster (https://kubernetes.example.com)
Manifest: namespace asset-solutions

✔ kubectl available              client v1.34.1
✔ cluster reachable              server v1.34.1
✔ kubernetes version             1.34 ≥ 1.34 required
✔ general worker nodes           3 schedulable node(s) (need ≥ 3): node-1, node-2, node-3
✔ transformations pool           2 node(s) labeled aks-node-pool=argocpu: node-4, node-5
✔ large transformations pool     no node at rest; provisioned on demand by Karpenter NodePool transformations-large (taint declared)
✔ default storage class          gp3 (cluster default, named in manifest)
✔ rwx storage class              efs-sc exists (RWX capability verified only with --probe-storage)
✔ namespace                      asset-solutions exists
✔ metrics api                    v1beta1.metrics.k8s.io available
✔ argocd application crd         applications.argoproj.io installed
✔ argocd controller              pod/argocd-application-controller-0 is Ready
- node ipv6 addresses            manifest ipFamily is ipv4 (default)

Summary: 12 passed, 0 warning, 0 failed, 1 skipped
```

## Parameters for cluster check

The `cluster check` command accepts the following parameters:

| Parameter         | Description                                                                                                                                    | Default                                                                         |
| ----------------- | ---------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------- |
| `--probe-storage` | Provisions and deletes a temporary 1-GiB test volume for each configured storage class to verify that provisioning works. Mutates the cluster. | `false`                                                                         |
| `--probe-timeout` | Time to wait for a probe volume to bind.                                                                                                       | `60s`                                                                           |
| `--dry-run`       | Prints the command plan without executing anything.                                                                                            | `false`                                                                         |
| `--format`        | Deployment format to validate for: `argocd` or `helm`. Selects whether the ArgoCD checks run.                                                  | `argocd` if the manifest contains a `deployment.argocd` block, otherwise `helm` |

## Verify that storage provisioning works

> **Important:**
>
> `--probe-storage` creates and deletes resources in your cluster. Every other `cluster check` option is read-only.

Add `--probe-storage` to verify that each configured storage class can provision a volume, including `ReadWriteMany` support:

```sh
shdctl cluster check --probe-storage
```

This is the only option that changes anything in your cluster. For each configured storage class (the three `storage probe` rows above), the command creates a temporary 1-GiB `PersistentVolumeClaim` named `shdctl-cluster-check-*` and labeled `shdctl.unity.com/cluster-check=true` in the target namespace, waits for it to bind, and then deletes it. If a probe fails to bind, the command reads the events in that namespace to report why.

The kubeconfig needs `get`, `list`, `watch`, `create` and `delete` on persistentvolumeclaims in the target namespace, and `list` on events there. A probe that binds provisions a real volume, which a class with `reclaimPolicy: Retain` keeps after the claim is gone.

The command skips, rather than probes, storage classes that use `volumeBindingMode: WaitForFirstConsumer`, because binding such a class requires a scheduled pod. Skipped checks don't affect the exit code.

## Preview the commands without contacting the cluster

Add `--dry-run` to print every `kubectl` command that the checks will run. The command executes nothing and doesn't contact the cluster, so you can use `--dry-run` for a security review:

```sh
shdctl cluster check --dry-run
```

A real run also reads your kubeconfig to name the context it checks; the plan doesn't list that read.

## Select the deployment format

By default, the command runs the ArgoCD checks if your manifest contains a `deployment.argocd` block, and skips them otherwise. To override that inference, pass `--format`:

```sh
shdctl cluster check --format argocd
```

## Additional resources

* [Prerequisites](../../shd/installation/prerequisites)
* [Global flags](./_index.md#global-flags)
* [shdctl manifest commands](./manifest.md)
* [Manifest reference](../manifest)
