Incident & Problem Management: Own the RCA process for production incidents – diagnose, resolve, and put preventive measures in place so issues don’t recur
Production Monitoring & Support: Continuously monitor service health, detect anomalies early, and act before they become incidents
Deployment Execution: Design, implement, and maintain CI/CD pipelines using GitHub Actions and related tooling to automate build, test, security scanning, and deployment processes.
Environment Oversight: Keep Pre-Production and Production environments stable and aligned — not building them from scratch, but ensuring they behave as expected day to day
Runbook & Knowledge Management: Document operational procedures, known issues, and resolution steps to build a reliable knowledge base for the team
Cross-team Collaboration: Work shoulder-to-shoulder with development and platform teams to triage issues, clarify operational requirements, and close the feedback loop between prod and dev