Collaborate with development teams to adopt DevOps best practices and tools in the software development lifecycle across different languages, such as PHP, Python, and Go
Debug and troubleshoot to find acceptable solutions for all running services and infrastructure aspects
Respond quickly to incidents and manage them to ensure service reliability
Improve and develop our monitoring stack based on Prometheus, Grafana, InfluxDB, Sentry, and Graylog
Manage and maintain our scaled and highly available infrastructure
Maintain replicated databases, including MySQL, Elasticsearch, and Redis
Implement and maintain CI/CD pipelines to streamline deployment and delivery processes
Automate infrastructure provisioning, configuration, and deployment using IaC tools like Ansible or Terraform
Ensure system availability by managing monitoring, logging, and alerting setups
Deploy and manage containerized applications using tools like Docker and Kubernetes
Resolve infrastructure issues and performance bottlenecks
Conduct security assessments and implement protective measures for infrastructure and data
Strong problem-solving skills and a willingness to learn new technologies
Be available to work outside regular office hours as required for maintenance, updates, and emergency response
Requirements
Proven experience in a DevOps or SRE role in the web-based online industry for at least 3 years
Good communication skills and a collaborative mindset
Outstanding problem-solving skills and an eye for detail
Proven experience with Linux, Nginx, and Varnish
Knowledge of MySQL, Galera, ProxySQL, Elasticsearch, and Redis administration
Knowledge of containerization with Docker and orchestration with Kubernetes
Familiarity with CI/CD tools (e.g., GitLab CI/CD or Jenkins) and understanding of automated pipelines
Familiarity with monitoring and logging tools (e.g., Prometheus, Grafana, InfluxDB/Telegraf, Graylog)
Knowledge of scripting languages like Bash or Python
Ability to document and report the implementation and resolution of issues