خانه هوش ایران
خانه هوش ایران

Senior DevOps Engineer

Tehran/Meydan Valiasr
Full Time
Saturday to Wednesday from 8 to 17
-
Health insurance -Lunch

این فرصت شغلی چقدر برای من مناسب است؟

11 - 50 employees
Technology and Innovation / VC / Accelerator
توضیحات بیشتر

key Requirements

5 years experience in similar position
GIT - Intermediate
Linux - Advanced
Jenkins - Intermediate
Microsoft Azure Devops / TFS - Intermediate
Docker - Advanced
Kubernetes - Advanced
Ansible - Intermediate

Job Description

Role Summary & Organizational Impact:
We are looking for a Senior DevOps Engineer to design, automate, and maintain the infrastructure of a large-scale platform built on microservices and AI/ML workloads, spanning cloud and hybrid environments. You will define standards for automation, security, observability, and reliability, directly impacting delivery velocity and product quality. This role bridges development, AI, and infrastructure teams, and serves as a technical leader in driving SRE and DevOps best practices.

Key Responsibilities:
1. Design, implement, and continuously improve CI/CD pipelines using GitOps principles for safe and rapid deployments.
2. Manage all infrastructure as code (IaC) and automate provisioning across cloud and on-premises environments.
3. Architect, deploy, and upgrade production-grade Kubernetes clusters, managing the full container lifecycle.
4. Implement and optimize comprehensive monitoring, centralized logging, distributed tracing, and intelligent alerting (Observability).
5. Integrate DevSecOps processes: container image scanning, secrets management, SAST/DAST, and software supply chain security.
6. Lead incident management, conduct blameless root cause analysis (RCA), write postmortems, and drive down MTTR.
7. Continuously optimize infrastructure performance, availability, capacity, and cost (FinOps).
8. Collaborate closely with development teams (.NET, Python AI) on architecture patterns, service communication, and non-functional requirements.
9. Mentor and upskill team members, fostering a culture of automation, measurement, and ownership.
10. Research and evaluate new technologies related to the platform, including GPU infrastructure and AI/ML tooling.

Must-Have Requirements / Essential Technical Skills:
1. Advanced Linux Administration (system-level debugging, kernel tuning, service management, networking, and OS security).
2. Docker & Kubernetes (production-grade): optimized image builds, managing large-scale production clusters, and zero-downtime upgrade strategies.
3. CI/CD pipeline design & implementation (GitLab CI, GitHub Actions, Jenkins): multi-stage pipelines and deployment strategies (Canary, Blue-Green).
4. Scripting & programming (Python, Bash, Go) for automation tools, CI/CD scripts, operators, and API integrations.
5. Infrastructure as Code: Terraform (modules, remote state, policy as code such as Sentinel/OPA).
6. Advanced networking & web serving (DNS, load balancers, reverse proxy, SSL/TLS, firewalls; Nginx/HAProxy).
7. Monitoring & observability (Prometheus, Grafana, ELK or cloud-native equivalents).
8. Cloud services (AWS, GCP, or Azure) across compute, networking, storage, and security services.
9. Database operations (SQL & NoSQL): performance tuning, backup, and disaster recovery (e.g., PostgreSQL, SQL Server, Redis, MongoDB).
10. Security & DevSecOps: container scanning (Trivy, Clair), secrets management (Vault, Sealed Secrets), SAST/DAST, and supply chain security.
11. Microservices & distributed systems architecture (REST, gRPC, message brokers; resilience patterns; service mesh concepts).
12. Kubernetes package management: Helm (production charts, upgrade/rollback strategies).
13. Git & collaborative workflows (Gitflow, trunk-based development, code review).

Behavioral & Leadership Competencies:
1. Automation mindset & continuous improvement.
2. Incident management & grace under pressure (blameless postmortems and corrective actions).
3. Effective communication & cross-team collaboration.
4. Mentoring & team growth.
5. Ownership & outcome-oriented mindset.

Nice-to-Have / Preferred Qualifications:
1. SRE practices (SLIs/SLOs, error budgets).
2. Message brokers & streaming (Kafka, RabbitMQ).
3. Storage & search technologies (Redis, Elasticsearch, MinIO, Qdrant).
4. OpenTelemetry & distributed tracing.
5. GPU infrastructure & AI/ML workloads (e.g., NVIDIA GPU Operator).
6. Advanced GitOps (ArgoCD, Flux).
7. Configuration management (Ansible, Chef, Puppet).
8. FinOps & cloud cost optimization.
9. Professional certifications (CKA, CKAD, AWS Certified DevOps, etc.).
10. Open-source contributions, technical articles, or books.
11. Hybrid cloud & on-prem experience (OpenStack, VMware).

Job Requirements

Gender
Men / Women
Software
Linux| Advanced Docker| Advanced Kubernetes| Advanced Ansible| Intermediate GIT| Intermediate Jenkins| Intermediate Microsoft Azure Devops / TFS| Intermediate

ثبت مشکل و تخلف آگهی

ارسال رزومه برای خانه هوش ایران