Description:Pegah is the technology group driving a wide range of digital products and businesses, including Cafe Bazaar, Tapsell, Metrix, Bazaar Pay, Metis AI, Gapify, Athena AI, Bebin TV, Beeptunes, Footballi, and more, serving over 50 million active users.
The Technology Team builds and operates the shared infrastructure, platforms, services, and core capabilities that enable engineering teams across the group to operate reliably at massive scale every single day.
We are looking for a Senior Platform Engineer to join our Technology team and help us build, operate, and evolve shared platforms and managed services used by engineering teams across the group.
Position Summary:As a Senior Platform Engineer, you will build and operate shared runtime platforms that enable engineering teams to deploy and run production workloads reliably, securely, and at scale.
You will work primarily with Kubernetes and cloud-native technologies, improving platform reliability, standardization, and automation while reducing operational complexity for product engineering teams.
This role is a good fit for someone who combines strong systems knowledge with an automation-first mindset, thinks cloud-native by default, and enjoys turning complex infrastructure capabilities into simple, reusable, and reliable platform services.
What You Will Do:- Design, operate, and improve shared Kubernetes clusters and the services that run on top of them.
- Manage Kubernetes cluster lifecycle, configuration, upgrades, capacity, and reliability across multiple production environments.
- Build and maintain shared runtime capabilities such as ingress, gateways, and service mesh technologies.
- Develop standardized, reusable platform components that simplify common engineering workflows.
- Automate platform provisioning, configuration, and lifecycle management using Infrastructure as Code, GitOps, and configuration management practices.
- Build and improve managed platform services, including databases, messaging systems, and other stateful services.
- Troubleshoot complex issues across Kubernetes, networking, workloads, platform services, and underlying infrastructure dependencies.
- Build security, access control, secrets management, and policy enforcement into the platform.
- Ensure platform services are observable with appropriate metrics, logs, and alerts.
- Monitor platform health and identify capacity, performance, availability, and operational risks.
- Participate in incident response and scheduled on-call rotations for critical platform services.
- Improve platform self-service and reduce manual operational work for engineering teams.
- Collaborate across infrastructure, reliability, delivery, security, and product engineering to continuously improve the engineering platform.
What We Expect:- 4+ years of experience in Platform Engineering, Systems Engineering, or a similar production-focused role.
- Strong problem-solving and troubleshooting skills in complex production environments.
- Solid knowledge of Linux systems and core containerization and cloud-native runtime concepts.
- Strong hands-on experience with Kubernetes in production, including a solid understanding of cluster architecture, lifecycle, networking, storage, scheduling, and failure modes.
- Strong understanding of networking and cloud-native networking concepts, including TCP/IP, DNS, service discovery, ingress, east-west traffic, network policies, and mTLS.
- Hands-on experience with service mesh, traffic management, and gateway technologies such as Istio, Linkerd, Envoy, or similar technologies.
- Experience with IaC, configuration management, and GitOps practices using tools such as Terraform, Ansible, Helm, or Argo CD.
- Proficiency in Python or Go, with a strong automation mindset and experience building internal tools, platform automation, or software against infrastructure APIs.
- Experience provisioning and operating highly available managed databases and messaging systems such as PostgreSQL, Redis, Kafka, Cassandra, or similar technologies, with a focus on availability, lifecycle, backup, and operational reliability.
- Proven experience owning and operating highly available distributed systems, production platforms, or shared services with a high degree of autonomy.
- A security-first mindset when designing and operating shared platform services.
- Strong communication and collaboration skills across engineering teams.
- Willingness to participate in scheduled on-call rotations and incident response for platform services.
Nice to Have:- Experience extending Kubernetes through custom controllers, operators, CRDs, or similar platform automation patterns.
- Familiarity with observability systems such as Prometheus, Grafana, VictoriaMetrics, or similar technologies.
- Experience with software delivery tooling and workflows, including GitLab CI/CD, runners, and artifact repositories.
- Experience designing internal platforms, reusable platform APIs, self-service workflows, or developer-facing platform services.
- Practical experience using AI-assisted engineering tools such as Claude, Codex, Gemini, Cursor, etc. to accelerate troubleshooting, automation, documentation, and day-to-day technical workflows.
Benefits:- A dynamic working environment with an open, innovative, and result-oriented culture.
- The opportunity to build shared platforms serving multiple large-scale digital businesses.
- Collaboration with experienced engineers across infrastructure, platform, reliability, data, AI, and product engineering teams.
- Flexible working hours with a hybrid working model.
- Structured on-call rotation program with compensation built into the overall package.
- Supplementary health insurance.
- Competitive compensation package.
- Various on-site entertainment and workplace facilities.