HCL Technologies
HCL Technologies

Senior Site Reliability Engineer Lead

RoleDevops
LevelSenior
LocationBengaluru, India
WorkOn-site
TypeFull-time
Posted1 month ago
Apply now

About the role

Job Summary

Role Summary:

The Operations SME focuses on day-to-day availability, reliability, and performance of production systems, ensuring SRE practices are embedded into operational processes.

Key Responsibilities:

  • Maintain and operate production applications, infrastructure, and databases

  • Implement and manage monitoring, alerting, and performance tools

  • Perform incident response, troubleshooting, and RCA

  • Execute hardware and software upgrades

  • Conduct capacity planning and performance analysis

  • Maintain runbooks, SOPs, and operational documentation

  • Support cloud, enterprise, SaaS, COTS, and legacy platforms

Must Have

  • 6+ years of experience in production operations

  • Experience managing VMs, servers, networks, and applications

  • Strong experience with monitoring, logging, and alerting

  • Understanding of reliability metrics and operational KPIs

  • Proficiency in scripting (Python, Bash, Groovy, Go Lang)

  • Understanding of container platforms

  • Familiarity with ITSM processes

Good to Have:

  • Application architecture exposure in Java or .NET

  • CI/CD and automation experience

  • Knowledge of Chaos Engineering

  • SRE certification

Key Responsibilities

  1. Lead and manage a team of support engineers in resolving incidents, requests, and problems to ensure system uptime and reliability.

  2. Collaborate with the engineering and development teams to implement efficient and scalable solutions that enhance system performance.

  3. Develop and maintain support documentation, standard operating procedures, and best practices for the support team.

  4. Identify opportunities for automation and implement tools to streamline support processes.

  5. Monitor system performance and provide recommendations for improvements to optimize system reliability.

  6. Participate in on call rotations to address critical incidents and ensure 24/7 system availability.

  7. Conduct regular performance evaluations, provide feedback, and mentor team members to promote professional growth.

Skill Requirements

  1. In-depth knowledge of site reliability engineering (sre) principles and best practices.

  2. Proficiency in system monitoring, incident management, and performance tuning tools.

  3. Strong understanding of cloud services, microservices architecture, and containerization technologies.

  4. Excellent problem-solving skills and the ability to troubleshoot complex technical issues.

  5. Experience with scripting languages (e.g., python, bash) for automation and tool development.

  6. Familiarity with agile methodologies and devops practices for continuous integration and delivery.

  7. Strong communication and leadership skills to effectively lead a support team and collaborate with cross functional teams.

  8. Ability to work under pressure, prioritize tasks, and manage multiple projects simultaneously.

Other Requirements

1.Relevant certifications in Site Reliability Engineering (SRE) or Cloud Services are a plus.

Benefits and perks

Learning Budget

Required skills

Cloud infrastructure

Reliability engineering

Automation

About HCL Technologies

Bengaluru

Headquarters