The largest job portal in the Middle East
Apply now

Job Description

Maintain high reliability of cloud-native systems on Kubernetes & OCI.

Monitor, triage & resolve complex issues across APIs and MQs.

Lead incident response, RCA, and postmortem.

Build an observability stack (logs, metrics, tracing).

Analyze Java/Node.js issues and support development teams.

Automate operational tasks and implement self-healing.

Mentor junior engineers and enhance SRE best practices.