فناوری اطلاعات و ارتباطات مُهَیمن
فناوری اطلاعات و ارتباطات مُهَیمن

Senior SRE / DevOps Engineer

Tehran/Abas Abad(Beheshti)
Full Time
Saturday to Wednesday
-
-

این فرصت شغلی چقدر برای من مناسب است؟

501 - 1000 employees
IT / Software / Hardware
Iranian company dealing only with Iranian entities
1388
Privately held
توضیحات بیشتر

key Requirements

3 years experience in similar position
Docker - Advanced
Docker Swarm - Intermediate
Kubernetes - Advanced
Helm - Intermediate
Prometheus - Intermediate
Gerafana - Intermediate
Ansible - Intermediate

Job Description

We are building and scaling a platform around AI agents, digital services, and multiple new lines of business. As our systems grow, reliability, scalability, observability, and operational excellence become critical parts of the product itself. We are looking for a Senior SRE / DevOps Engineer who can help us design, operate, and continuously improve the infrastructure and engineering practices behind our services.

This role goes beyond maintaining clusters or deployment pipelines. We are looking for someone who understands large-scale software systems, can identify the right operational and infrastructure patterns, introduce better practices when needed, and work closely with software engineers and external technical teams to make services reliable and production-ready.

What You’ll Do:

  • Design, operate, and continuously improve reliable and scalable production infrastructure.
  • Build and maintain Kubernetes-based environments and containerized workloads.
  • Design and improve GitOps-based deployment workflows using Argo CD and Helm.
  • Create and maintain reusable Helm charts and deployment standards across services.
  • Define and maintain SLIs, SLOs, SLAs, error budgets, and reliability targets for critical services.
  • Help define measurable technical and operational requirements for third-party vendors and development partners.
  • Work with external teams to ensure their services meet agreed standards for availability, latency, monitoring, scalability, and incident response.
  • Design and maintain monitoring, logging, tracing, dashboards, and alerting systems.
  • Improve observability so engineering teams can quickly understand system health and diagnose production issues.
  • Troubleshoot complex production problems across applications, infrastructure, databases, networking, and distributed systems.
  • Perform root-cause analysis and help implement long-term fixes rather than relying on temporary operational workarounds.
  • Improve service resilience through proper timeout, retry, rate-limiting, failover, and scaling strategies.
  • Perform capacity planning and help services scale as traffic and workloads increase.
  • Work closely with software engineers to make new services scalable, observable, resilient, and production-ready.
  • Consult development teams on architecture, infrastructure, databases, caching, messaging, deployment strategies, and operational tooling.
  • Help teams choose technologies and engineering practices based on actual technical and business requirements.
  • Define production-readiness standards and review services before major releases.
  • Build reusable infrastructure and platform capabilities that simplify deployment and operation for development teams.
  • Automate repetitive operational processes and reduce manual intervention.
  • Improve backup, recovery, failover, and disaster-recovery processes where required.

Job Requirements:

Core Requirements:

  • Strong professional experience in Site Reliability Engineering, DevOps, Platform Engineering, or production infrastructure, preferably in large-scale environments.
  • Strong hands-on experience with Kubernetes and containerized production environments.
  • Strong experience with Helm, including designing and maintaining reusable Helm charts.
  • Strong hands-on experience with Argo CD and GitOps-based deployment practices.
  • Good understanding of Kubernetes concepts including deployments and workloads, services and ingress, networking, storage, resource management, autoscaling, high availability, access control, and security.
  • Strong understanding of distributed systems, scalability, availability, fault tolerance, and production architecture.
  • Experience designing and operating high-availability production services.
  • Strong understanding of GitOps and declarative infrastructure and deployment practices.
  • Experience designing deployment strategies such as rolling updates, canary releases, and controlled rollouts.
  • Strong experience with observability technologies such as Prometheus, Grafana, Elasticsearch / OpenSearch, ELK / EFK, Loki, OpenTelemetry, and distributed tracing platforms.
  • Strong understanding of metrics, logs, traces, dashboards, and actionable alerting.
  • Practical experience defining SLIs, SLOs, SLAs, error budgets, availability targets, and reliability requirements.

Preferred Qualifications:

  • Experience in one or more of the following is considered a plus:
  • Large-scale or high-traffic platforms
  • Multi-cluster Kubernetes environments
  • Advanced Helm deployment patterns
  • Argo CD at scale
  • Kubernetes operators and controllers
  • Service mesh technologies
  • API gateways
  • Kafka and event-driven architectures
  • PostgreSQL or other large production databases
  • Redis clusters
  • Elasticsearch / OpenSearch
  • Ceph, MinIO, or other distributed storage platforms
  • Secrets management solutions
  • OpenTelemetry and distributed tracing

Job Requirements

Age
23 - 40 Years Old
Gender
Men / Women
Education
Bachelor| Computer and IT
Software
Kubernetes| Advanced Docker| Advanced Helm| Intermediate Docker Swarm| Intermediate Prometheus| Intermediate Gerafana| Intermediate Ansible| Intermediate

ثبت مشکل و تخلف آگهی

ارسال رزومه برای فناوری اطلاعات و ارتباطات مُهَیمن