
Senior Site Reliability Engineer
ポジションについて
Reliability Engineering
Design, build, and maintain highly available and fault-tolerant production systems.
Define and monitor SLIs, SLOs, and SLAs for critical services.
Drive reliability improvements through automation and proactive engineering.
Conduct capacity planning and performance optimization activities.
Production Support & Operations:
Manage production environments and ensure service uptime.
Lead incident response, troubleshooting, and root cause analysis (RCA).
Develop runbooks, operational playbooks, and disaster recovery procedures.
Participate in on-call rotations and major incident management processes.
Cloud & Infrastructure:
Deploy and manage cloud-native infrastructure across AWS, Azure, or GCP.
Automate infrastructure provisioning using Infrastructure as Code (IaC).
Implement scalable and secure infrastructure solutions.
Support Kubernetes-based platforms and containerized workloads.
Observability & Monitoring:
Build monitoring, logging, tracing, and alerting solutions.
Implement observability frameworks using industry-standard tools.
Monitor application health, performance metrics, and infrastructure utilization.
Drive continuous improvements in platform visibility and diagnostics.
Automation & DevOps:
Automate deployments, infrastructure management, and operational workflows.
Improve CI/CD pipelines and release processes.
Implement self-healing, auto-scaling, and operational automation solutions.
Promote DevOps and SRE best practices across engineering teams.
Security & Compliance:
Ensure production environments meet security and compliance requirements.
Manage secrets, access controls, and vulnerability remediation.
Partner with security teams to implement security best practices.
Leadership & Mentoring:
Mentor junior SREs, DevOps Engineers, and Production Support Engineers.
Lead incident reviews and reliability improvement initiatives.
Conduct knowledge-sharing sessions and technical workshops.
Drive operational excellence and engineering best practices.
Education: Bachelor of Engineering
Preferred skills: Technology->DevOps->Site Reliability Engineering(SRE)
福利厚生
•Learning Budget
必須スキル
Cloud infrastructure
Reliability engineering
Automation
Infosysについて
BANGALORE
本社所在地