
Senior Site Reliability Engineer Lead
About the role
Job Summary
Role Summary:
The Operations SME focuses on day-to-day availability, reliability, and performance of production systems, ensuring SRE practices are embedded into operational processes.
Key Responsibilities:
-
Maintain and operate production applications, infrastructure, and databases
-
Implement and manage monitoring, alerting, and performance tools
-
Perform incident response, troubleshooting, and RCA
-
Execute hardware and software upgrades
-
Conduct capacity planning and performance analysis
-
Maintain runbooks, SOPs, and operational documentation
-
Support cloud, enterprise, SaaS, COTS, and legacy platforms
Must Have
-
6+ years of experience in production operations
-
Experience managing VMs, servers, networks, and applications
-
Strong experience with monitoring, logging, and alerting
-
Understanding of reliability metrics and operational KPIs
-
Proficiency in scripting (Python, Bash, Groovy, Go Lang)
-
Understanding of container platforms
-
Familiarity with ITSM processes
Good to Have:
-
Application architecture exposure in Java or .NET
-
CI/CD and automation experience
-
Knowledge of Chaos Engineering
-
SRE certification
Key Responsibilities
-
Lead and manage a team of support engineers in resolving incidents, requests, and problems to ensure system uptime and reliability.
-
Collaborate with the engineering and development teams to implement efficient and scalable solutions that enhance system performance.
-
Develop and maintain support documentation, standard operating procedures, and best practices for the support team.
-
Identify opportunities for automation and implement tools to streamline support processes.
-
Monitor system performance and provide recommendations for improvements to optimize system reliability.
-
Participate in on call rotations to address critical incidents and ensure 24/7 system availability.
-
Conduct regular performance evaluations, provide feedback, and mentor team members to promote professional growth.
Skill Requirements
-
In-depth knowledge of site reliability engineering (sre) principles and best practices.
-
Proficiency in system monitoring, incident management, and performance tuning tools.
-
Strong understanding of cloud services, microservices architecture, and containerization technologies.
-
Excellent problem-solving skills and the ability to troubleshoot complex technical issues.
-
Experience with scripting languages (e.g., python, bash) for automation and tool development.
-
Familiarity with agile methodologies and devops practices for continuous integration and delivery.
-
Strong communication and leadership skills to effectively lead a support team and collaborate with cross functional teams.
-
Ability to work under pressure, prioritize tasks, and manage multiple projects simultaneously.
Other Requirements
1.Relevant certifications in Site Reliability Engineering (SRE) or Cloud Services are a plus.
Benefits and perks
•Learning Budget
Required skills
Cloud infrastructure
Reliability engineering
Automation
About HCL Technologies
Bengaluru
Headquarters