We're hiring a Senior DevOps Engineer to operate production Kubernetes clusters on bare-metal/private infrastructure (self-managed via kubeadm — not EKS/GKE/AKS).
Responsibilities
Build and maintain kubeadm Kubernetes clusters (control-plane + worker nodes)
Manage containerd, CNI, CoreDNS, Ingress, Helm, RBAC, PV/PVC
Handle cluster upgrades, node maintenance, backup/recovery
Troubleshoot Linux and Kubernetes production issues at the root cause
Set up monitoring/alerting (Prometheus, Grafana, Alertmanager)
Build CI/CD and automation
Operate infra services: Nginx, Redis, PostgreSQL, RabbitMQ, Elasticsearch, MinIO
Required Skills
Linux: systemd/journalctl, process & resource management, CPU/memory/disk troubleshooting, LVM/filesystems, permissions/SSH, kernel-level log troubleshooting, Bash scripting
Kubernetes (self-managed): kubeadm, kubelet, containerd, control plane, etcd, CNI, CoreDNS, Ingress, Helm, storage (PV/PVC), scheduling, upgrades
Networking (basic): IP/subnet/gateway, DNS, TCP/UDP, ports, routing/NAT, HTTP/HTTPS, TLS, connectivity troubleshooting
Automation/Monitoring (some of): Ansible, Bash/Python, GitLab CI/CD, Jenkins, Argo CD, Prometheus, Grafana, Alertmanager
Nice to Have
Bare-metal datacenter experience, NVIDIA GPU infra (GPU Operator/CUDA), Ceph/NFS/OpenEBS, private registries, HA Kubernetes, disaster recovery, GitLab admin
Experience
5+ years in DevOps/Linux/SRE/Infrastructure
3+ years hands-on Kubernetes, specifically self-managed/bare-metal
Production support experience
Managed-Kubernetes-only experience is not sufficient
Priority order: Linux → Bare-Metal Kubernetes → Troubleshooting → Automation → Monitoring