Maintain high reliability of cloud-native systems on Kubernetes & OCI.
Monitor, triage & resolve complex issues across APIs and MQs.
Lead incident response, RCA, and postmortem.
Build an observability stack (logs, metrics, tracing).
Analyze Java/Node.js issues and support development teams.
Automate operational tasks and implement self-healing.
Mentor junior engineers and enhance SRE best practices.