HCL Technologies
HCL Technologies

SME - AWS IAC, Terraform,Python

RoleEngineering
LevelMid Level
LocationBengaluru, India
WorkOn-site
TypeFull-time
Posted3 months ago
Apply now

About the role

Job Summary

We are looking for a Systems Storage Site Reliability Engineer (SRE) to support and scale our global storage platforms. This is a contractor position focused on applying SRE principles to storage systems—improving reliability, reducing operational toil, and enabling sustainable growth through automation and observability.\r\n You will work at the intersection of storage engineering and reliability engineering, partnering closely with infrastructure and application teams to operate production systems at scale.\r\n What You’ll Do\r\n Own the reliability, availability, and performance of production NAS and/or Object Storage services.\r\n Apply SRE principles to storage platforms: define reliability goals, improve observability, and reduce manual operational work through automation.\r\n Design and build automation and Infrastructure‑as‑Code to manage storage systems at scale.\r\n Lead troubleshooting and resolution of complex storage incidents; participate in on‑call and incident response.\r\n Perform capacity planning, forecasting, and demand modeling to support business growth.\r\n Partner with engineering teams to support application onboarding, testing, and production readiness.\r\n Contribute to global storage initiatives, including lab and infrastructure deployments.\r\n Create and maintain runbooks, documentation, and operational best practices to improve team efficiency.\r\n What We’re Looking For\r\n8+ years of experience in SRE, infrastructure automation, or platform engineering, with strong storage exposure.\r\n Hands‑on experience operating NAS and/or Object Storage platforms, Luster/Ceph in production.\r\n Strong proficiency with automation and IaC tools (e.g., Ansible, Terraform, Puppet, Salt Stack).\r\n Experience running highly available, scalable systems in 24×7 environments.\r\n Familiarity with containers and orchestration (Docker, Kubernetes).\r\n Experience with CI/CD pipelines, monitoring, logging, and version control systems (Git, Perforce).\r\n Strong incident management, troubleshooting, and communication skills.\r\n Bachelor’s degree in Computer Science, Engineering, or a related field.\r\n Nice to Have\r\n Experience with large‑scale distributed systems.\r\n Strong understanding of SRE concepts such as SLIs, SLOs, error budgets, observability, and logging.\r\n Ability to debug and optimize infrastructure and automate repetitive workflows.\r\n Proven ability to work independently and deliver results as a contractor in a global team environment. About the Role We are looking for a Systems Storage Site Reliability Engineer (SRE) to support and scale our global storage platforms. This is a contractor position focused on applying SRE principles to storage systems—improving reliability, reducing operational toil, and enabling sustainable growth through automation and observability.You will work at the intersection of storage engineering and reliability engineering, partnering closely with infrastructure and application teams to operate production systems at scale.What You’ll Do Own the reliability, availability, and performance of production NAS and/or Object Storage services.Apply SRE principles to storage platforms: define reliability goals, improve observability, and reduce manual operational work through automation.Design and build automation and Infrastructure‑as‑Code to manage storage systems at scale.Lead troubleshooting and resolution of complex storage incidents; participate in on‑call and incident response.Perform capacity

Key Responsibilities

Automation

Skill Requirements

NAS/ Storage Platform/ Luster/ Ceph, CI/CD

Other Requirements

kubernetes, docker

Benefits and perks

Learning Budget

Required skills

DevOps

Cloud infrastructure

Quality assurance

Marketing

Design

Communication

About HCL Technologies

Bengaluru

Headquarters