
Senior Site Reliability Engineer
ポジションについて
We are looking for an experienced Site Reliability Engineer (SRE) / Production Engineer to build and operate highly available, scalable, and resilient production systems. The ideal candidate will have strong expertise in cloud infrastructure, automation, observability, incident management, performance engineering, and DevOps practices.
The role requires close collaboration with software engineering, infrastructure, security, and platform teams to ensure operational excellence and improve system reliability at scale.
Reliability Engineering
Design, build, and maintain highly available and fault-tolerant production systems.
Define and monitor SLIs, SLOs, and SLAs for critical services.
Drive reliability improvements through automation and proactive engineering.
Conduct capacity planning and performance optimization activities.
Production Support & Operations:
Manage production environments and ensure service uptime.
Lead incident response, troubleshooting, and root cause analysis (RCA).
Develop runbooks, operational playbooks, and disaster recovery procedures.
Cloud & Infrastructure:
Deploy and manage cloud-native infrastructure across AWS, Azure, or GCP.
Automate infrastructure provisioning using Infrastructure as Code (IaC).
Implement scalable and secure infrastructure solutions.
Support Kubernetes-based platforms and containerized workloads.
Observability & Monitoring:
Build monitoring, logging, tracing, and alerting solutions.
Implement observability frameworks using industry-standard tools.
Monitor application health, performance metrics, and infrastructure utilization.
Drive continuous improvements in platform visibility and diagnostics.
Automation & DevOps:
Automate deployments, infrastructure management, and operational workflows.
Improve CI/CD pipelines and release processes.
Implement self-healing, auto-scaling, and operational automation solutions.
Promote DevOps and SRE best practices across engineering teams.
Security & Compliance:
Ensure production environments meet security and compliance requirements.
Manage secrets, access controls, and vulnerability remediation.
Partner with security teams to implement security best practices.
Education: MCA,MSc,Bachelor of Engineering,BBA,BCom,BCS
Preferred skills: Technology->DevOps->Site Reliability Engineering(SRE),Domain->Manufacturing->Production Planning
福利厚生
•Learning Budget
必須スキル
Cloud infrastructure
Reliability engineering
Automation
Infosysについて
HYDERABAD
本社所在地