همراه کسب و کارهای هوشمند
همراه کسب و کارهای هوشمند

Senior SRE

Tehran/Vanak
Full Time
Saturday to Wednesday
-
Loan -Bonus -Health insurance -Flexible working hours -Learning stipends -Purchasing coupon -Snacks -In-house Medical doctor -Breakfast -Occasional packages and gifts

این فرصت شغلی چقدر برای من مناسب است؟

201 - 500 employees
Internet Provider / E-commerce / Online Services
Iranian company dealing only with Iranian entities
1399
Privately held
توضیحات بیشتر

key Requirements

6 years experience in similar position
Python - Basic
Linux - Intermediate
Zabbix - Intermediate
Prometheus - Intermediate
Gerafana - Intermediate

Job Description

Key Responsibilities

Service Reliability Engineering

  • Design and implement strategies to improve service reliability, availability, resilience, and operational stability.
  • Monitor and analyze service health across distributed applications, APIs, infrastructure, and supporting components.
  • Define and track SLIs, SLOs, SLAs, and error budgets for critical services.
  • Identify reliability risks, single points of failure, recurring incidents, and systemic weaknesses.
  • Drive initiatives to reduce service degradation, incident frequency, and recovery time.
  • Contribute to resilience, fault tolerance, and operational readiness improvements.
  • Establish reliability standards and best practices across services and platforms.

Observability & Service Visibility

  • Design and enhance observability capabilities across metrics, logs, traces, and events.
  • Develop centralized monitoring and telemetry strategies for critical services.
  • Build and maintain service health dashboards and reliability views.
  • Design effective alerting strategies with appropriate thresholds, prioritization, and noise reduction.
  • Improve visibility across distributed services, APIs, infrastructure, and integrations.
  • Correlate telemetry and operational events to identify service-impacting conditions.

Incident Analysis & Reliability Improvement

  • Perform advanced analysis of service incidents, alerts, and operational events.
  • Support incident response by providing deep service health and telemetry insights.
  • Lead or contribute to Root Cause Analysis (RCA) and post-incident reviews.
  • Identify recurring failure patterns and reliability gaps.
  • Translate incident findings into concrete monitoring, architecture, and operational improvements.
  • Track reliability improvement actions through their implementation and effectiveness.

Performance & Capacity Engineering

  • Analyze service performance, latency, throughput, and resource utilization trends.
  • Identify performance bottlenecks and potential service degradation risks.
  • Support capacity planning and forecasting using operational and monitoring data.
  • Analyze system behavior under normal and abnormal operating conditions.
  • Provide recommendations for performance, scalability, and resource optimization.
  • Develop service reliability, performance, and capacity reports.

Service Health & Operational Readiness

  • Define and maintain Service Health Indicators for critical business and technical services.
  • Establish operational readiness criteria for new and existing services.
  • Assess service dependencies and their potential impact on reliability.
  • Support reliability assessments for major changes, releases, and new service deployments.
  • Ensure monitoring, alerting, recovery, and operational requirements are properly defined.
  • Collaborate with technical teams to improve overall service resilience and operational maturity.

Required Technical Skills

  • Strong hands-on experience in Service Reliability, Production Operations, or Site Reliability Engineering.
  • Strong experience with monitoring and observability platforms such as Prometheus, Zabbix, Grafana, Splunk, and Kibana.
  • Experience with centralized logging and telemetry platforms such as ELK Stack or Splunk.
  • Strong understanding of observability concepts: metrics, logs, traces, events, and telemetry.
  • Solid understanding of distributed systems and service dependencies.
  • Strong Linux administration and performance troubleshooting skills.
  • Good understanding of network and infrastructure monitoring fundamentals.
  • Experience with application, API, and service-level monitoring.
  • Experience with event management, incident management, and ticketing systems.
  • Understanding of service availability, reliability, and operational monitoring practices.

Reliability & Engineering Knowledge

  • Strong understanding of Service Reliability Engineering principles.
  • Practical knowledge of SLO, SLI, SLA, and error budget frameworks.
  • Strong understanding of MTTR, MTBF, availability, resilience, and service continuity.
  • Experience with incident analysis, RCA, and post-incident improvement.
  • Knowledge of alert optimization and signal-to-noise improvement.
  • Ability to identify systemic reliability risks and recurring failure patterns.
  • Understanding of fault tolerance, redundancy, graceful degradation, and resilience.
  • Strong analytical and problem-solving skills.
  • Understanding of service dependency and failure-domain analysis.

Nice to Have

  • Experience with APM and distributed tracing platforms.
  • Experience with OpenTelemetry.
  • Experience with automation and scripting using Python or Bash.
  • Exposure to cloud-native and containerized environments.
  • Familiarity with Kubernetes and microservices architectures.
  • Experience with performance testing and capacity analysis.
  • Familiarity with ITIL / ITSM processes.
  • Experience with service management or enterprise production environments.

Soft Skills

  • Strong analytical and systems-thinking mindset.
  • Calm and effective under incident pressure.
  • Cross-team collaboration skills.
  • Clear reporting and documentation.

Work Conditions

  • Participation in incident scenarios and reliability reviews.
  • Close collaboration with Service Operations teams.
  • Participation in incident reviews, RCA, and reliability improvement initiatives.

Job Requirements

Gender
Men / Women
Military service
Military service must be done
Software
Prometheus| Intermediate Zabbix| Intermediate Gerafana| Intermediate Linux| Intermediate Python| Basic

ثبت مشکل و تخلف آگهی

ارسال رزومه برای همراه کسب و کارهای هوشمند