Case study 01 · Platform engineering

Enterprise-style Kubernetes platform.

A six-node home-lab cluster designed to practice production-minded orchestration, GitOps, persistent storage, security policy, observability, backup, and application operations.

Control-plane expansion and failover

Expanded the original single-control-plane cluster to three control planes and three workers. Updated kubeadm with a stable API endpoint, added the endpoint to API certificate SANs, joined the additional control planes, and deployed kube-vip for virtual-IP leadership.

  • All six nodes were observed Ready after the expansion.
  • The automation runner kubeconfig was switched from an individual control plane to the shared API endpoint.
  • The current runbook records a hard power-off test of the original control plane with API availability maintained through the surviving nodes.
  • A worker power-off test confirmed ingress VIP movement and continued ingress availability, with three ingress controllers distributed across the workers.

These tests cover the API and ingress paths. Stateful application recovery is assessed separately.

Objective

Build a realistic platform that could host useful services while providing hands-on experience with the operational concerns that exist beyond a basic Kubernetes installation: networking, storage, certificates, identity, policy, upgrades, monitoring, logging, backups, and declarative delivery.

Architecture

  • Three Ubuntu control-plane nodes and three Ubuntu worker nodes running Kubernetes 1.36.2 with containerd 2.2.1. Each Hyper-V host carries one control plane and one standalone worker VM.
  • Calico for pod networking and enforceable NetworkPolicies.
  • MetalLB for LoadBalancer services and three ingress-nginx controllers, one per worker.
  • cert-manager for automated certificate lifecycle management.
  • Longhorn distributed block storage on dedicated 1 TB worker disks, with NFS backup targets and an offsite Azure Blob backup CronJob.
  • Argo CD app-of-apps pattern for GitOps reconciliation.
  • Sealed Secrets for encrypted secret material stored in Git.

Delivery and operations

Applications are defined in Git and reconciled by Argo CD. Kustomize and Helm are used where appropriate, and automated image workflows can update digest-pinned deployments. Current workloads include documentation, password management, monitoring, uptime checks, logging, DNS automation, tooling, and portfolio services.

Operational practices

  • Namespace-based separation for applications and platform components.
  • Resource requests and limits, health probes, disruption budgets, and topology spreading.
  • Persistent-volume backups and recovery planning.
  • Prometheus metrics, Grafana dashboards, Loki logs, and Alertmanager email delivery.
  • Git history as the change record and Argo CD as the drift detector.

Security controls demonstrated by this portfolio

  • Restricted Pod Security enforcement pinned to the cluster version.
  • Non-root execution, read-only root filesystem, RuntimeDefault seccomp, no privilege escalation, and all Linux capabilities dropped.
  • Dedicated ServiceAccount with token automount disabled.
  • Default-deny ingress and egress with only ingress-controller traffic allowed.
  • Restricted Argo CD AppProject limited to one repository and one namespace.
  • ResourceQuota and LimitRange to constrain blast radius.

Engineering lessons

The project reinforced that a useful Kubernetes platform is an integration of multiple control planes. Storage, certificates, DNS, ingress, monitoring, and GitOps must all agree on naming, identity, networking, and lifecycle behavior. Troubleshooting therefore requires moving methodically across layers instead of treating every symptom as an application problem.

See another case study.

Continue with Microsoft infrastructure engineering or observability.