Case study 04 · Reliability engineering

Observability and alerting.

A layered monitoring design that combines infrastructure metrics, logs, synthetic availability checks, traditional-server visibility, dashboards, and actionable email notification.

Infrastructure Homepage and host monitoring

Migrated Zabbix to Rocky Linux and assembled a Homepage dashboard for infrastructure and application access. Added metrics-server and scoped ServiceAccount permissions, verified node/pod and metrics access, and resolved Kubernetes widget errors.

Integrated internal certificate trust and reorganized services around the current Hyper-V platform. The custom Normandy/EDI theme required separate checks at the pod, service, and external route; external theme delivery remained a troubleshooting item.

Monitoring layers

  • Prometheus: cluster, node, controller, and application metrics.
  • Grafana: dashboards for operational visibility and trend analysis.
  • Loki and Promtail: centralized Kubernetes log collection and search.
  • Alertmanager: routing, grouping, repeat intervals, and email delivery through an internal relay.
  • Uptime Kuma: synthetic checks for service and endpoint availability.
  • Zabbix: infrastructure monitoring on a dedicated Rocky Linux host outside the Kubernetes-native stack.

Alert delivery

Alertmanager was integrated with an internal SMTP relay and verified end to end with a dedicated PrometheusRule. The configuration separates actionable alerts from noise, routes Watchdog-style signals away from the user receiver, and applies a repeat interval to reduce unnecessary messages.

Operational outcomes

  • One place to inspect cluster health, application state, and resource consumption.
  • Centralized logs for troubleshooting restarts, ingress failures, and application errors.
  • Email notification for conditions that need attention away from the dashboard.
  • Independent uptime checks that validate the user-facing path rather than only pod health.
  • Traditional-server monitoring alongside cloud-native telemetry; VMware and OpenStack have been retired from the current lab.

Reliability mindset

Monitoring is treated as part of the service design, not an afterthought. New workloads should expose health checks, declare resource requirements, produce useful logs, and have a clear alerting path. The objective is early detection with enough context to shorten diagnosis and recovery.

Return to the complete portfolio.

Review the experience, skills, credentials, and live platform architecture.

Back to portfolio