Huawei is a leading global provider of information and communications technology (ICT) infrastructure and smart devices. With integrated solutions across four key domains – telecom networks, IT, smart devices, and cloud services – we are committed to bringing digital to every person, home and organization for a fully connected, intelligent world.
We are looking for a Software Reliability Engineer to ensure the reliability, availability, and scalability of our systems. The role works closely with development teams to improve system resilience, automate operations, and respond to production incidents (software updates, bug fixes, and security patches).
Responsibilities
Monitor system health and assist in identifying and troubleshooting common issues.
Implement and maintain monitoring and alerting dashboards.
Participate in on-call rotations, responding to incidents with guidance.
Contribute to blameless postmortems and recommend systemic improvements.
Assist in capacity planning and basic automation of manual tasks.
Collaborate with development teams on deploying updates and patches.
Good scripting/automation skills (Python, Shell, Go).
Solid experience with container orchestration (Kubernetes/CCE) and service mesh concepts.
Deep understanding of network protocols, load balancing, and security best practices.
Experience with observability stacks and SLI/SLO design.
Ability to work independently and mentor junior engineers.
Basic knowledge in at least one field: network, storage, virtualization, containerization, Linux operating systems, cybersecurity, big data, or databases.
Requirements
Bachelor’s degree in Computer Science, Software Engineering, a related field, or equivalent practical experience.
+2 years of experience with software development or cloud O&M experience.
Excellent communication skills.
Pragmatic, problem-solving attitude.
A naturally curious and proactive approach to learning and problem-solving.
Good skills in writing/editing documentation.
Preferred Familiarity with OpenStack deployment or operations. Familiarity with public cloud deployment or operations.