دیجی پی
دیجی پی

SRE Engineer

Tehran/Vanak
Full Time
شنبه تا چهارشنبه
-
-

این فرصت شغلی چقدر برای من مناسب است؟

201 - 500 employees
Finance / Investment
Iranian company dealing only with Iranian entities
1397
Privately held
توضیحات بیشتر

key Requirements

3 years experience in similar position
Linux - Intermediate
Redis - Intermediate
RabbitMQ - Intermediate
Docker - Basic
Kubernetes - Basic
Prometheus - Intermediate
Gerafana - Intermediate

Job Description

Job Summary:

The Production Engineering team at Digipay is responsible for improving the reliability, scalability, performance, and resilience of critical production systems. The role combines software engineering, infrastructure, automation, observability, and performance engineering to build reliable and scalable production environments.

Key Responsibilities:

  • Improve reliability and availability of critical production services.
  • Define and monitor SLIs, SLOs, and Error Budgets.
  • Design and implement automation and self-healing solutions.
  • Build and improve monitoring, observability, logging, and tracing solutions.
  • Conduct performance testing, capacity planning, and bottleneck analysis.
  • Troubleshoot complex production and performance issues.
  • Improve the resilience of applications, databases, messaging systems, and infrastructure.
  • Participate in incident management, root cause analysis, and postmortems.
  • Review system architectures and provide recommendations for reliability and scalability.
  • Develop and maintain production readiness standards, runbooks, and reliability tooling.

Qualifications:

  • Bachelor’s degree in Computer Engineering, Computer Science, or a related field.
  • 3+ years of relevant experience in SRE, Production Engineering, DevOps, Infrastructure, or Performance Engineering.
  • Strong knowledge of Linux and operating system fundamentals.
  • Solid understanding of distributed systems, high availability, fault tolerance, and scalability.
  • Hands-on experience with monitoring and observability tools such as Prometheus and Grafana.
  • Experience with logging and tracing platforms such as ELK and Jaeger.
  • Practical experience with RabbitMQ, Redis, and relational databases.
  • Experience with performance testing, profiling, benchmarking, and capacity planning.
  • Good understanding of database performance, query optimization, caching, and messaging patterns.
  • Experience with CI/CD and infrastructure automation.
  • Programming or scripting experience with Python, Go, Java, or similar languages.
  • Familiarity with containerization technologies such as Docker; Kubernetes experience is a plus.
  • Experience troubleshooting production systems and performing Root Cause Analysis.
  • Good understanding of reliability and resilience patterns such as retries, circuit breakers, rate limiting, idempotency, graceful degradation, and load shedding.

Competencies:

  • Problem Solving: Ability to analyze complex technical problems and identify root causes.
  • Systems Thinking: Ability to understand dependencies and failure modes across distributed systems.
  • Engineering Mindset: Focus on building scalable, maintainable, and automated solutions rather than relying on manual operations.
  • Ownership: Takes responsibility for system reliability and follows issues through to resolution.
  • Analytical Thinking: Ability to use data, metrics, logs, and traces to make technical decisions.
  • Performance Mindset: Strong interest in identifying bottlenecks and continuously improving system performance.
  • Collaboration: Ability to work effectively with Engineering, DevOps, Product, and other technical teams.
  • Incident Management: Ability to remain effective and make sound decisions during production incidents.
  • Continuous Improvement: Proactively identifies opportunities to improve reliability, automation, and operational efficiency.
  • Communication: Ability to clearly communicate technical issues, risks, recommendations, and solutions.
  • Learning Agility: Willingness and ability to learn new technologies, tools, and engineering practices.

Job Requirements

Gender
Men / Women
Education
Bachelor| Computer and IT
Software
Linux| Intermediate Prometheus| Intermediate Gerafana| Intermediate Docker| Basic Kubernetes| Basic RabbitMQ| Intermediate Redis| Intermediate

ثبت مشکل و تخلف آگهی

ارسال رزومه برای دیجی پی