Capgemini Engineering

Incident Operation Engineer

Capgemini Engineering

Cluj-Napoca, Iasi, Timisoara, Bucharest, Brasov, Romania (Remote)

Job Type Permanent
Experience Experienced Professionals
Posted 2 weeks 1 day ago
Remote Friendly Experienced Professionals Level

At Capgemini Engineering, the world leader in engineering services, we bring together a global team of engineers, scientists, and architects to help the world's most innovative companies unleash their potential. From autonomous cars to life-saving robots, our digital and software technology experts provide unique R&D and engineering services across all industries.

Your role

We are looking for an experienced Incident Operations Engineer to join our team and play a critical role in managing high-priority incidents across complex production environments. You will serve as a central coordination point during incidents, driving communication, impact assessment, stakeholder alignment, and operational excellence.

  • Monitor and respond to real-time alerts, triage incidents, and support incident response activities.
  • Coordinate across Engineering, Incident Command, Customer Support, and Operations teams to drive efficient incident resolution.
  • Assess incident impact, severity, and customer exposure using monitoring tools and system insights.
  • Own customer-facing communications, including incident notifications, status updates, and resolution reports.
  • Manage and maintain public status pages, ensuring timely and accurate updates.
  • Contribute to post-incident reviews, RCA processes, SLA reporting, and operational improvements.
  • Drive automation and process optimization initiatives using technologies such as Python or Kotlin.
  • Support enhancements to monitoring, observability, escalation processes, and operational tooling.

Your Profile

  • 7+ years of experience in Incident Management, Site Reliability Engineering (SRE), Technical Operations, Production Operations, or a similar role.
  • Experience working in on-call and SLA-driven environments.
  • Strong understanding of distributed systems, production environments, and service reliability.
  • Hands-on experience with monitoring tools such as Datadog, Grafana, Prometheus, or similar platforms.
  • Experience with incident management tools such as PagerDuty, Opsgenie, ServiceNow, or Rootly.
  • Programming experience with Python or Kotlin.
  • Strong communication skills with the ability to manage high-pressure situations and multiple priorities simultaneously.

Similar Job Openings