
این فرصت شغلی چقدر برای من مناسب است؟
At Snappfood, we believe the best ideas emerge when people with different experiences, perspectives, and strengths come together.
We’re looking for people who can bring a fresh perspective, energy, and expertise to help grow our business and create better experiences for millions of users.
If you think you can bring your own unique flavor to our team, we’d love to hear from you.
Responsibilities:
• Drive engineering projects in the SRE team end to end: problem framing, technical design, delivery, rollout, and adoption across engineering teams.
• Own the observability-as-code practice. Build the frameworks and pipelines that let dashboards, alerts, SLOs, and instrumentation be defined as versioned, reviewed, and tested code, with CI/CD pipelines and self-service onboarding for product teams.
• Define and build observability for agentic AI systems: distributed tracing across agent steps and tool calls, token and cost accounting, latency and failure-mode visibility, and signals for output quality and regression detection.
• Set reliability and scalability architecture standards: reference architectures, resilience patterns (retries, circuit breaking, graceful degradation, load shedding), capacity and scaling models, and dependency and failure-domain analysis.
• Act as the technical authority in the team: write and review design documents, review code, set engineering quality standards, and mentor engineers.
• Reduce operational toil through automation and platform work, and use incident and postmortem findings to feed the engineering roadmap.
• Partner with product engineering, platform, and security teams to drive adoption of standards and tooling, including the organizational work of getting buy-in rather than just shipping the tool.
• Keeps up to date with industry trends and best practices in SRE, cloud infrastructure, and distributed systems.
Requirements:
• 4–5+ years of experience in SRE, DevOps, Infrastructure, or Software Engineering.
• Strong understanding of SRE principles, incident management, system reliability, and distributed systems.
• Working knowledge of LLM-based and agentic systems and the monitoring challenges they bring.
• Hands-on experience with Linux, and Kubernetes.
• Experience with CI/CD, Infrastructure as Code, and automation.
• Deep experience with modern observability stacks (OpenTelemetry, Prometheus, Grafana, distributed tracing) and their cost and cardinality trade-offs.
• Programming/scripting experience in Python, Go, Bash, or similar.
• Strong technical leadership, communication, and problem-solving skills.
• Ability to work effectively under pressure and manage shifting priorities.
• Strong ownership and result-oriented mindset.
Benefits:
• Vouchers for vacations, gym, therapy.
• Social Security & Complementary Insurance.
• Educational platform of advanced courses.
• Snappfood’s discount codes.
• Loans.
ثبت مشکل و تخلف آگهی
ارسال رزومه برای اسنپ فود