ポジションについて
- Stack Management: Deploy, configure, and maintain the core observability stack using Prometheus, Grafana, Alertmanager, and Loki.
- Dashboarding & Visualization: Collaborate with engineering teams to design and build comprehensive Grafana dashboards for application and infrastructure health monitoring.
- Alerting Strategy: Configure and fine-tune Alertmanager rules to ensure accurate, actionable alerts while minimizing alert fatigue.
- Log Management: Architect and manage centralized logging solutions using Loki to ensure efficient log aggregation and querying.
- System Optimization: Monitor the performance of the observability stack itself, optimizing resource usage and scaling infrastructure as needed.
- Continuous Improvement: Evaluate and integrate modern observability tools (such as Victoria Metrics and Victoria Logs) to enhance system correlation and analysis capabilities.
- Experience: 9+ years of hands-on experience in DevOps, SRE, or Observability roles.
- Core Stack: Deep technical expertise in Prometheus, Grafana, Alertmanager, and Loki (PLG Stack).
- Automation: Proficiency in scripting (Python, Bash) and Infrastructure as Code (e.g., Terraform, Ansible).
- Infrastructure & OS: Strong working knowledge of Linux/Unix administration.
- Containerization: Experience monitoring containerized environments (Docker, Kubernetes).
- Hands-on knowledge or prior exposure to Victoria Metrics (for scalable time-series data) and Victoria Logs.
- Experience with distributed tracing tools (e.g., Jaeger, Tempo, or Victoria Traces).
- Understanding of service level indicators (SLIs) and service level objectives (SLOs).
Location of posting: Chennai, Bangalore, Hyderabad, Pune
Education: Bachelor Of Science,Bachelor of Engineering
Preferred skills: Technology->Server-Virtualization->Open Stack
福利厚生
•Learning Budget
Infosysについて
BANGALORE
本社所在地
