Senior Platform Reliability & AI Operations Engineer
Most reliability teams react to incidents. We engineer them out of existence.
We operate an AI-native community and social engagement platform that underpins customer relationships for Fortune 100 brands. When our platform goes down, it isn't a blip on an internal dashboard — it's a reputational crisis for some of the most recognized companies on earth. That context shapes everything about how we work, what we build, and who we hire.
We're looking for a senior reliability engineer who can carry production on their shoulders and build the autonomous systems that progressively carry it for them. You'll join a team where AI agents are first-class operational teammates — triaging alerts, validating changes, drafting root-cause analyses, and applying remediations within defined guardrails. Your mission is to make that surface area grow every single week.
Your Day-to-Day
Carry the pager and own the outcome. You're the first responder on your shift window. When production degrades, you command the incident — diagnose, mitigate,
escalate when blast radius demands it, and restore service. You treat every customer-impacting minute as personal accountability.
Engineer autonomous operational workflows. Build, deploy, and refine the AI agents that handle pre-triage, change-gate validation, auto-healing, RCA drafting, and preventive-fix tracking. The agents are the product; your operational expertise is the training data.
Ship safe production changes. Every deploy, config update, and cost-optimization action flows through quality gates with a validated rollback plan. You abort without hesitation the moment telemetry deviates from the expected path.
Investigate to true root cause — then close the loop. Separate symptom from cause with disciplined analysis, then go further: identify the systemic prevention, build it, and track it to production. Unshipped RCA action items are unfinished work.
Generalize every manual intervent