About the Role
We are looking for a Mid-Level DevOps / Platform Engineer to help manage and improve our Kubernetes infrastructure and deployment processes. This is a hands-on role focused on reliability, automation, security, monitoring, and operational efficiency.
Responsibilities:
- Manage, maintain, and optimize production Kubernetes clusters.
- Design and implement resource provisioning and capacity planning strategies.
- Develop, maintain, and optimize Helm charts for application deployments.
- Build, maintain, and improve CI/CD pipelines to support efficient software delivery.
- Implement and maintain monitoring, alerting, and observability solutions.
- Manage centralized logging and operational visibility across the platform.
- Design and maintain backup, recovery, and disaster recovery processes.
- Troubleshoot and resolve infrastructure, networking, and deployment issues.
- Optimize infrastructure utilization, performance, and operational costs.
- Maintain internal dependency mirrors, including Docker registries and NPM repositories.
- Create and maintain technical documentation, runbooks, and operational procedures.
- Collaborate with development teams to improve deployment workflows and platform reliability.
Required Skills & Experience
- Experience administering production Kubernetes environments.
- Strong Linux system administration skills.
- Strong networking knowledge, including TCP/IP, DNS, routing, firewalls, load balancing, and troubleshooting.
- Experience creating and maintaining Helm charts.
- Experience building and maintaining CI/CD pipelines.
- Experience implementing monitoring and alerting solutions.
- Experience with backup and disaster recovery planning.
- Understanding of Kubernetes security, RBAC, secrets management, and cluster hardening.
- Ability to troubleshoot infrastructure and application issues independently.
- Strong documentation and communication skills.
- Preferred Qualifications
- Experience with ELK Stack (Elasticsearch, Logstash, Kibana).
- Experience with Grafana and Prometheus.
- Experience managing container registries and package repositories.
- Experience with Infrastructure as Code tools.
- Experience supporting high-availability production environments.
Engagement Details
- Part-time position.
- Flexible working arrangement.
- Ability to work independently and take ownership of platform operations.
- Availability for critical infrastructure incidents when required.