شرکت سفرهای علی‌بابا
شرکت سفرهای علی‌بابا

Senior DevOps Engineer

Tehran/ Kooye Bimeh
Full Time
شنبه تا چهارشنبه
-
-
501 - 1000 employees
Internet Provider / E-commerce / Online Services
Iranian company dealing with Iranian and foreign customers
1393
Privately held
توضیحات بیشتر

key Requirements

5 years experience in similar position
Python - Intermediate
Linux - Advanced
Microsoft Azure Devops / TFS - Intermediate
Docker - Intermediate
Kubernetes - Intermediate
Ansible - Intermediate
Gitlab - Intermediate

Job Description

Alibaba runs the platform behind a fast-growing, multi-brand travel business. We're hiring a Senior DevOps Engineer to join a small, high-ownership DevOps/SRE team responsible for our production and staging environments.

Our platform is fully self-hosted and on-premise — we run our own Kubernetes, GitLab, CI/CD, distributed storage, databases, and messaging on our own infrastructure. This is a hands-on role for an engineer who wants to own systems end-to-end across the full stack, and who is energized by strengthening and modernizing a large, established estate. It is not a managed-cloud position.

What you'll do

  • Operate and evolve our Kubernetes platform — Rancher/RKE and kubespray clusters: version upgrades, scaling, storage (Longhorn), networking, and workload security.
  • Operate our stateful services — PostgreSQL, Redis, MongoDB, Elasticsearch, Kafka, RabbitMQ, and distributed storage: upgrades, backups, failover, capacity planning, and incident response. This is a significant part of the role.
  • Automate provisioning, patching, upgrades, and lifecycle across our Linux estate with Ansible and Terraform/OpenTofu — reducing toil so the team can focus on higher-value work.
  • Lead modernization and lifecycle work — operating-system upgrades, patch management, and standardization across the estate.
  • Own CI/CD and GitOps — GitLab CI pipelines and ArgoCD-based delivery.
  • Strengthen security and access — secrets management (HashiCorp Vault), SSO/IAM (Keycloak), access control, and CIS hardening, all of which we are actively maturing.
  • Own observability — metrics, logs, traces, and alerting for always-up, always-available services.
  • Troubleshoot production issues end-to-end and drive them to resolution with minimal customer impact.
  • Participate in the on-call rotation and partner with development teams to adopt DevOps best practices and move applications toward cloud-native patterns on our platform.
  • Remediate at the root — identify recurring issues and resolve their underlying causes.

Requirements

  • Deep Linux/Unix systems administration — the core of this role. 5–8 years operating, tuning, securing, and troubleshooting Linux servers at scale in production.
  • Strong TCP/IP and networking fundamentals (DNS, HTTP/S, TLS, routing, load balancing).
  • Production expertise with Docker and Kubernetes, including hands-on cluster lifecycle (Rancher, RKE, and/or kubespray).
  • Configuration management and IaC: Ansible, plus Terraform/OpenTofu.
  • CI/CD with GitLab, and solid Git / GitFlow.
  • Strong scripting in at least Bash and Python (Perl or Ruby a bonus).
  • Proven experience operating stateful systems in production — relational (PostgreSQL) and NoSQL/key-value (Redis, MongoDB), plus messaging (RabbitMQ).
  • Web servers and load balancers (Nginx, HAProxy).
  • Observability in production: VictoriaMetrics/Prometheus, Grafana, and log/trace pipelines (Fluentd/EFK or ELK, OpenTelemetry).
  • Working knowledge of at least one programming language (.NET, Node.js, or Go) to partner effectively with application teams.
  • Knowledge of best practices for running an always-up, always-available service.
  • Excellent troubleshooting, a fast learner, and a strong ownership mindset — comfortable operating independently and taking a problem from first report to resolution.

Nice to have

  • GitOps at scale (ArgoCD) and Helm.
  • Distributed storage: Longhorn, Ceph, or MinIO.
  • Security: CIS Benchmarks, OWASP, container image scanning (e.g., Trivy).
  • Service mesh implementation (Istio, Linkerd).
  • Artifact management (JFrog Artifactory) and workflow automation (n8n).
  • Kafka / streaming at scale; the Hadoop ecosystem.
  • Prior production experience with Rancher.

Job Requirements

Gender
Men / Women
Software
Linux| Advanced Docker| Intermediate Kubernetes| Intermediate Gitlab| Intermediate Ansible| Intermediate Microsoft Azure Devops / TFS| Intermediate Python| Intermediate

ثبت مشکل و تخلف آگهی

ارسال رزومه برای شرکت سفرهای علی‌بابا