
Incident Operation Engineer
About the role
Incident Operation Engineer (Remote, Romania)
At Capgemini Engineering, the world leader in engineering services, we bring together a global team of engineers, scientists, and architects to help the world's most innovative companies unleash their potential. From autonomous cars to life-saving robots, our digital and software technology experts think outside the box as they provide unique R&D and engineering services across all industries. Join us for a career full of opportunities. Where you can make a difference. Where no two days are the same.
Your role:
We are looking for an experienced Incident Operations Engineer to join our team and play a critical role in managing high-priority incidents across complex production environments. You will serve as a central coordination point during incidents, driving communication, impact assessment, stakeholder alignment, and operational excellence.
- Monitor and respond to real-time alerts, triage incidents, and support incident response activities.
- Coordinate across Engineering, Incident Command, Customer Support, and Operations teams to drive efficient incident resolution.
- Assess incident impact, severity, and customer exposure using monitoring tools and system insights.
- Own customer-facing communications, including incident notifications, status updates, and resolution reports.
- Manage and maintain public status pages, ensuring timely and accurate updates.
- Contribute to post-incident reviews, RCA processes, SLA reporting, and operational improvements.
- Drive automation and process optimization initiatives using technologies such as Python or Kotlin.
- Support enhancements to monitoring, observability, escalation processes, and operational tooling.
Your Profile:
- 7+ years of experience in Incident Management, Site Reliability Engineering (SRE), Technical Operations, Production Operations, or a similar role.
- Experience working in on-call and SLA-driven environments.
- Strong understanding of distributed systems, production environments, and service reliability.
- Hands-on experience with monitoring tools such as Datadog, Grafana, Prometheus, or similar platforms.
- Experience with incident management tools such as Pager Duty, Opsgenie, Service Now, or Rootly.
- Programming experience with Python or Kotlin.
- Strong communication skills with the ability to manage high-pressure situations and multiple priorities simultaneously.
About Capgemini
Timisoara
Headquarters