
SME - Kubernetes, Terraform
About the role
Job Summary
As a Subject Matter Expert in Support & Operations, you will play a critical role in ensuring timely resolution of escalated incidents while adhering to quality standards. Your expertise in Kubernetes and Terraform will guide the team in optimizing operations, enhancing customer satisfaction, and maintaining effective communication across business units.
Key Responsibilities
-
Ensure Timely Resolution And Quality Compliance Of Escalated Tickets And Incidents Using Kubernetes And Terraform, Adhering To Agreed Slas.
-
Mentor Team Members And Administrators In Best Practices For Kubernetes And Terraform, Prepare Standard Operating Procedures (Sops), And Maintain Effective Documentation To Enhance Operational Efficiency.
-
Validate Change Order Implementation Plans And Ensure Human Error Compliance While Participating In Capacity Planning Activities Using Aws Cloud Formation And Eks.
-
Actively Participate In Customer Meetings To Understand Issues Faced, Ensuring Positive Customer Feedback And Satisfaction Through Effective Communication And Problem-Solving.
-
Validate Root Cause And Trend Analyses Using Aws Paas Tools And Present Reports To Key Business Stakeholders To Facilitate Performance Improvements In Operations.
Skill Requirements
-
Strong Understanding Of Kubernetes And Terraform With Practical Application In Incident Resolution
-
Proficiency In Aws Cloud Formation, Eks, And Linux Environments
-
Solid Knowledge Of Support Processes And Service Level Agreements (Slas)
-
Excellent Communication And Presentation Skills For Stakeholder Engagement
-
Ability To Perform Root Cause And Trend Analysis Effectively
Other Requirements
-
Extensive experience with Kubernetes, including deploying, managing, and troubleshooting clusters in production.
-
Certified Kubernetes Administrator (CKA) or equivalent certification.
-
Strong Linux expertise, including experience with performance tuning, troubleshooting, and scripting.
-
Networking fundamentals for cloud and telecom environments.
-
Proven ability to create detailed technical documentation, including requirements, HLD, and LLD documents.
-
Strong troubleshooting skills across hardware (servers), OS (Linux), networks, Kubernetes, and cloud platform software.
-
Proven experience in incident management and root cause analysis in production environments.
-
Development expertise in operational efficiency and automation tools (e.g., Ansible, Python).
-
Proven management and communication skills to lead operations as a technical leader, including progress and risk management.
Benefits and perks
•Learning Budget
Required skills
Kubernetes
Terraform
AWS
CloudFormation
EKS
Linux
Incident management
Root cause analysis
About HCL Technologies
Bengaluru
Headquarters