تپسی فود
تپسی فود

Senior Site Reliability Engineer (SRE)

Tehran/ Gandi
Full Time
Saturday to Wednesday
-
-
201 - 500 employees
Internet Provider / E-commerce / Online Services
Iranian company dealing only with Iranian entities
1402
توضیحات بیشتر

key Requirements

3 years experience in similar position
Python - Intermediate
Go - Intermediate
Linux - Intermediate
Docker - Intermediate
Kubernetes - Intermediate
Prometheus - Intermediate
Gerafana - Intermediate

Job Description

Job Summary

We are looking for a Senior Site Reliability Engineer (SRE) to join our Platform Engineering team. In this role, you will be responsible for ensuring the reliability, scalability, availability, and performance of our production platform by building resilient systems, automating operations, and improving observability across critical services.

Job Description

  • Define and manage SLOs, SLIs, and error budgets
  • Develop CI/CD and GitOps-based operational workflows
  • Lead incident response, root cause analysis (RCA), and post-incident reviews
  • Improve platform scalability, resilience, and disaster recovery capabilities
  • Design, operate, and optimize highly available production systems
  • Build and improve monitoring, logging, tracing, dashboards, and alerting solutions
  • Automate operational processes and reduce manual effort through Infrastructure as Code
  • Collaborate with engineering teams to enhance application reliability and production readiness
  • Drive reliability best practices and continuously optimize system performance and capacity

Requirements

  • Proficiency in Bash, Python, or Go
  • 3+ years of experience managing large-scale production environments
  • Strong hands-on experience with Kubernetes and containerized applications
  • Solid understanding of Linux, networking, and distributed systems
  • Experience with Infrastructure as Code tools such as Helm or Ansible
  • Experience with CI/CD pipelines and production deployment strategies
  • Experience managing production incidents and performing root cause analysis
  • Hands-on experience with observability tools including Prometheus, Grafana, OpenTelemetry, and ELK
  • Familiarity with Redis, RabbitMQ, PostgreSQL, MySQL, Argo CD, service mesh technologies, or cloud-native security concepts is a plus
  • Strong analytical, automation, and problem-solving skills with a focus on reliability and operational excellence

Job Requirements

Age
20 - 36 Years Old
Gender
Men / Women
Software
Python| Intermediate Go| Intermediate Linux| Intermediate Docker| Intermediate Kubernetes| Intermediate Prometheus| Intermediate Gerafana| Intermediate

ثبت مشکل و تخلف آگهی

ارسال رزومه برای تپسی فود

insight applicant

مقایسه من با سایر متقاضیان