خانه هوش ایران
خانه هوش ایران

MLOps Engineer

Tehran/Meydan Valiasr
Full Time
Saturday to Wednesday from 8 to 17
-
Health insurance -Lunch

این فرصت شغلی چقدر برای من مناسب است؟

11 - 50 employees
Technology and Innovation / VC / Accelerator
توضیحات بیشتر

key Requirements

2 years experience in similar position
Python - Intermediate
GIT - Intermediate
Linux - Intermediate
Jenkins - Basic
Docker - Intermediate
Kubernetes - Intermediate
Prometheus - Intermediate

Job Description

Tehran, Iran | Full-time | Salary: Negotiable (Base + Bonus)

About the Role
Bridge GPU/CPU health automation and ML lifecycle management. You’ll automate node health checks and build robust Kubeflow pipelines, ensuring a high-performance environment for AI engineers.

Key Responsibilities
  • GPU/CPU Health Automation (Priority #1): Automate health checks during node provisioning (link stability, ECC, PCIe, thermal) and store results in a Ground Truth Library; collaborate with GPU Performance Engineering on thresholds.
  • Build and maintain end-to-end ML pipelines in Kubeflow (deployment, monitoring).
  • Manage Kubernetes clusters with Docker in production.
  • Automate CI/CD via GitHub Actions (or GitLab CI/Jenkins).
  • Write Python/Bash scripts for Linux systems and infrastructure tasks.
  • Implement monitoring/alerting with Prometheus & Grafana (ELK stack is a plus).
  • Troubleshoot across dev/staging/production; document processes, runbooks, and postmortems.
  • Ensure security, compliance, and governance in cloud/on-prem environments.
Required Skills & Qualifications
  • 2+ years in MLOps, DevOps, SRE, or infrastructure roles.
  • Strong Python and Linux skills.
  • Hands-on Kubernetes & Docker in production.
  • Experience with CI/CD (GitHub Actions, GitLab CI, or Jenkins).
  • Understanding of deep learning pipelines & frameworks (TensorFlow/PyTorch).
  • Prometheus & Grafana experience (ELK familiarity is a plus).
  • Basics of networking, security, system administration, and databases.
  • Hardware-level diagnostic concepts (ECC, PCIe errors, thermal, link stability).
  • Cloud fundamentals (AWS, GCP, or Azure).
Personal Attributes & Soft Skills
  • Detail-oriented
  • Problem-solving & logical thinking
  • Curious & inquisitive
  • Teamwork & collaboration
  • Patience & perseverance
  • Quick learner, flexible
  • Responsibility & discipline
  • Creative & innovative thinking
  • Ownership mindset for reliability and outcomes
  • Resilience under pressure; R&D experimentation mindset
Nice-to-Have
  • SRE practices (SLIs/SLOs, incident response)
  • IaC tools (Terraform, Ansible)
  • Distributed systems & microservices architecture
  • Prior GPU workload or hardware validation exposure
  • Kubeflow/MLflow advanced experience
  • Relevant certifications (CKA/CKAD, cloud certs)
Benefits
  • R&D-driven environment with cutting-edge GPU technologies
  • Competitive salary + performance bonus
  • Continuous learning, skill-building, and career growth
  • Shared individual and organizational development goals

Job Requirements

Gender
Men / Women
Software
Python| Intermediate Linux| Intermediate Kubernetes| Intermediate Docker| Intermediate GIT| Intermediate Jenkins| Basic Prometheus| Intermediate

ثبت مشکل و تخلف آگهی

ارسال رزومه برای خانه هوش ایران