Technical Leadership & Infrastructure Ownership
- Own the overall architecture, availability, scalability, performance, security, and lifecycle management of the organization's VDI and virtualization platforms.
- Lead the technical direction and continuous evolution of the VDI infrastructure aligned with business objectives and technology roadmap.
- Act as the technical authority and escalation point for all virtualization-related services and critical incidents.
- Drive technical governance, establish engineering standards, operational procedures, and infrastructure best practices.
Enterprise Virtualization Expertise
- Extensive hands-on experience in designing, deploying, configuring, administering, maintaining, optimizing, and troubleshooting VMware vSphere environments, including VMware ESXi and VMware vCenter Server.
- Deep understanding of VMware Cluster Services, including HA, DRS, Storage DRS, vMotion, Storage vMotion, Distributed Switches, Lifecycle Manager, and Distributed Resource Management.
- Strong experience with VMware Horizon, Unified Access Gateway (UAG), App Volumes, Dynamic Environment Manager (DEM), and enterprise VDI architecture.
- Experience with VMware vSAN design, deployment, optimization, and lifecycle management.
- Experience with VMware NSX-T or equivalent Software Defined Networking technologies is highly desirable.
- Experience with VMware Aria Operations (formerly vRealize Operations), Aria Automation, Aria Log Insight, and related VMware ecosystem products.
- Experience with VMware Cloud Foundation or Hybrid VMware Cloud environments is considered an advantage.
Infrastructure Architecture & Engineering
- Design highly available, scalable, resilient, and secure virtualization architectures for enterprise production environments.
- Develop infrastructure standards, reference architectures, technical blueprints, and lifecycle strategies.
- Evaluate emerging technologies and recommend technical improvements aligned with business needs.
- Lead infrastructure modernization, platform transformation, consolidation, and migration initiatives.
- Design infrastructure capable of supporting business continuity, high availability, and future organizational growth.
Capacity Planning & Performance Engineering
- Perform enterprise-level Capacity Planning, Resource Forecasting, and Infrastructure Sizing.
- Analyze infrastructure utilization trends and proactively identify scalability limitations.
- Optimize CPU, Memory, Storage, Network, and Compute resources to maximize infrastructure efficiency.
- Develop long-term infrastructure growth strategies based on business forecasts.
- Establish proactive performance monitoring and predictive capacity planning processes.
Automation & Infrastructure as Code
- Strong experience automating VMware environments using PowerCLI, PowerShell, Python, REST APIs, or similar technologies.
- Experience implementing Infrastructure as Code (IaC) using Terraform, Ansible, or comparable automation platforms.
- Experience designing automated provisioning workflows using VMware Aria Automation or equivalent platforms.
- Promote an Automation-First operational model to minimize manual intervention and improve operational consistency.
- Experience integrating infrastructure automation into CI/CD pipelines is considered an advantage.
Monitoring, Observability & AIOps
- Design and implement enterprise monitoring strategies covering infrastructure health, availability, capacity, and performance.
- Strong knowledge of VMware Aria Operations, enterprise monitoring platforms, and observability best practices.
- Experience implementing intelligent alerting, predictive analytics, and AIOps capabilities.
- Define monitoring standards, dashboards, KPIs, and operational metrics.
-
- Storage, Compute & Data Protection
- Strong experience with enterprise storage technologies including SAN, NAS, Fibre Channel, iSCSI, and NFS.
- Hands-on experience managing Blade Infrastructure including HPE BladeSystem C7000, HPE Synergy, or equivalent enterprise compute platforms.
- Strong understanding of enterprise networking concepts related to virtualization infrastructures.
- Extensive experience with Veeam Backup & Replication, backup validation, SureBackup, replication, and disaster recovery.
- Experience implementing immutable backups and cyber-resilient backup architectures is preferred.
Disaster Recovery & Business Continuity
- Design, implement, maintain, and regularly test Disaster Recovery (DR) and Business Continuity (BCP) strategies.
- Define and manage Recovery Time Objectives (RTO) and Recovery Point Objectives (RPO).
- Plan and execute DR drills, failover/failback procedures, and recovery validation exercises.
- Ensure infrastructure resilience against hardware failures, cyber incidents, and site outages.
Security & Compliance
- Strong understanding of VMware security hardening guidelines and security best practices.
- Experience implementing least privilege access models, RBAC, certificate lifecycle management, and secure administrative practices.
- Collaborate with Cyber Security teams to maintain compliance with organizational security policies.
- Experience supporting vulnerability remediation, patch governance, and infrastructure compliance audits.
- Knowledge of Zero Trust principles and virtualization security architecture is highly desirable.
Operations & Service Management
- Lead operational management of enterprise virtualization services ensuring high availability and service reliability.
- Plan and execute Patch Management, Lifecycle Management, Firmware Upgrades, and Infrastructure Modernization with minimal service disruption.
- Lead Change Management activities while minimizing operational risk.
- Establish operational runbooks, SOPs, technical documentation, and knowledge repositories.
- Continuously improve operational maturity through automation, standardization, and process optimization.
Incident, Problem & Root Cause Management
- Lead Major Incident (P1/P2) response activities across cross-functional technical teams.
- Coordinate infrastructure recovery efforts during business-critical outages.
- Perform Root Cause Analysis (RCA), Post Incident Reviews (PIR), and implement Corrective and Preventive Actions (CAPA).
- Drive continuous operational improvements based on incident trends and service analytics.
- Experience working within ITIL-based Incident, Problem, Change, and Release Management processes.
Project & Technology Leadership
- Lead infrastructure Upgrade, Migration, Expansion, Consolidation, and Transformation projects.
- Develop technical project plans, risk assessments, rollback strategies, and implementation roadmaps.
- Coordinate with architecture, security, networking, storage, application, and business teams throughout project execution.
- Ensure project delivery within agreed scope, timeline, quality, and operational requirements.
Vendor & Stakeholder Management
- Manage relationships with technology vendors, solution providers, and support partners.
- Evaluate technical proposals, product roadmaps, and proof-of-concept solutions.
- Escalate complex technical issues with vendors and ensure SLA compliance.
- Present technical recommendations, risks, and strategic initiatives to executive management.
- Effectively communicate with technical and non-technical stakeholders.
Financial & License Management
- Participate in infrastructure budgeting, procurement planning, and technology investment decisions.
- Optimize licensing models, operational costs, and infrastructure utilization.
- Support Total Cost of Ownership (TCO) and Return on Investment (ROI) analysis for infrastructure initiatives.
Team Leadership & People Management
- Lead, mentor, coach, and develop a high-performing virtualization engineering team.
- Define team objectives, KPIs, performance expectations, and development plans.
- Promote knowledge sharing, technical mentoring, succession planning, and continuous learning.
- Plan resource allocation, workload balancing, and operational coverage.
- Foster a culture of ownership, accountability, collaboration, and operational excellence.
- Conduct technical reviews, performance evaluations, and competency assessments.
Innovation & Continuous Improvement
- Continuously evaluate emerging virtualization, cloud, automation, AI, and infrastructure technologies.
- Recommend innovative solutions to improve service quality, scalability, operational efficiency, and cost optimization.
- Define and maintain a multi-year virtualization technology roadmap.
- Promote engineering excellence through continuous improvement initiatives.
Technical Skills
- VMware vSphere
- VMware Horizon
- VMware vSAN
- VMware NSX-T
- VMware Aria Suite
- VMware Cloud Foundation
- VMware Tanzu (Preferred)
- VMware HCX (Preferred)
- Veeam Backup & Replication
- HPE BladeSystem / Synergy
- Enterprise SAN Storage
- PowerCLI
- PowerShell
- Python
- REST APIs
- Terraform
- Ansible
- Git
- Windows Server
- Active Directory
- DNS
- DHCP
- PKI / Certificate Services
- Enterprise Networking
- Microsoft SQL Server (Basic Administration)
- Monitoring & Observability Platforms
Preferred Certifications
- VMware Certified Professional (VCP)
- VMware Certified Advanced Professional (VCAP)
- VMware Certified Design Expert (VCDX) (Highly Preferred)
- ITIL Foundation
- HPE Accredited Solutions Expert (HPE ASE)
- Veeam Certified Engineer (VMCE)
Technical Proficiency
- Server Virtualization (VMware ESXi / vSphere and related technologies) — Advanced
- Veeam Backup & Replication — Advanced
- Ansible — Intermediate
Leadership Competencies
- Strategic Thinking
- Technical Leadership
- Enterprise Architecture Mindset
- Ownership & Accountability
- Decision Making Under Pressure
- Critical Thinking
- Analytical Problem Solving
- Risk Assessment & Mitigation
- Stakeholder Management
- Executive Communication
- Negotiation Skills
- Coaching & Mentoring
- Conflict Resolution
- Time & Priority Management
- Customer & Service Orientation
- Continuous Learning Mindset
- Innovation & Technology Evaluation