עדיין מחפשים עבודה במנועי חיפוש? הגיע הזמן להשתדרג!
במקום לעבור לבד על אלפי מודעות, Jobify מנתחת את קורות החיים שלך ומציגה לך רק משרות שבאמת מתאימות לך.
מעל 80,000 משרות • 4,000 חדשות ביום
חינם. בלי פרסומות. בלי אותיות קטנות.
Senior Platform Reliability & AI Operations Engineer
Most reliability teams react to incidents. We engineer them out of existence.
We operate an AI-native community and social engagement platform that underpins customer relationships for Fortune 100 brands. When our platform goes down, it isn't a blip on an internal dashboard — it's a reputational crisis for some of the most recognized companies on earth. That context shapes everything about how we work, what we build, and who we hire.
We're looking for a senior reliability engineer who can carry production on their shoulders and build the autonomous systems that progressively carry it for them. You'll join a team where AI agents are first-class operational teammates — triaging alerts, validating changes, drafting root-cause analyses, and applying remediations within defined guardrails. Your mission is to make that surface area grow every single week.
Your Day-to-Day
- Carry the pager and own the outcome. You're the first responder on your shift window. When production degrades, you command the incident — diagnose, mitigate, escalate when blast radius demands it, and restore service. You treat every customer-impacting minute as personal accountability.
- Engineer autonomous operational workflows. Build, deploy, and refine the AI agents that handle pre-triage, change-gate validation, auto-healing, RCA drafting, and preventive-fix tracking. The agents are the product; your operational expertise is the training data.
- Ship safe production changes. Every deploy, config update, and cost-optimization action flows through quality gates with a validated rollback plan. You abort without hesitation the moment telemetry deviates from the expected path.
- Investigate to true root cause — then close the loop. Separate symptom from cause with disciplined analysis, then go further: identify the systemic prevention, build it, and track it to production. Unshipped RCA action items are unfinished work.
- Generalize every manual intervention. A one-off fix restores service; encoding it into an agent, runbook, or guardrail prevents recurrence. You're measured by how much the autonomous layer can handle — not by how many tickets pass through your hands.
- Multiply team knowledge. Encode procedures, context, and decision logic so agents can retrieve it and the next responder never starts from scratch. In a distributed, async organization, undocumented expertise doesn't count.
- 5+ years of hands-on production operations in SRE, Platform Engineering, DevOps, or Cloud Infrastructure at SaaS scale — with real first-responder incident experience, not adjacent project work.
- Battle-tested on AWS — multi-AZ, multi-account environments, infrastructure-as-code, production incident management, and change control with gates and rollbacks. You've weathered significant outages and carry the operational intuition that only comes from living through them.
- Radically self-directed. You run your shift like a founder runs a company. You identify gaps, prioritize ruthlessly, ship solutions, and raise the bar — without waiting for direction. If a standard is wrong, you challenge it openly; you never quietly ignore it.
- AI-native in practice, not in theory. You routinely delegate substantive operational work to agents, critically evaluate their output, and iterate on the underlying capabilities when they fall short. Experienced with agentic tooling — Claude Code, Codex, Warp, custom agent frameworks — and compulsively curious about new models and techniques as they emerge.
- AWS Solutions Architect – Associate or higher — or a production track record that renders the certification a formality.
- Fluent, precise English — in incident-bridge communication and in long-form writing alike.
- Committed to shift-based coverage. On-call rotations and your designated shift window are foundational to the role, with the time-zone overlap your window requires.
- OFAC-clear country of residence.
- Original contributions to the agentic operations or AIOps space — open-source tooling, technical writing, conference talks, or shipped internal platforms.
- Production experience with multi-tenant B2B SaaS — community platforms, social tools, customer-experience products, or observability systems.
- Working knowledge of Grafana, Prometheus, Datadog, PagerDuty, or OpsGenie. Azure exposure alongside your AWS depth.
- Evidence of deep, sustained obsession with a hard problem — professional or personal. Depth of curiosity matters more than breadth of résumé.
You won't just read about the future of AI-driven reliability — you'll be one of the engineers who built it, on a platform that Fortune 100 brands depend on daily. The agent-ops skills, incident patterns, and architectural instincts you develop here are things the broader industry is still struggling to define, let alone hire for.
How We Work
- Enterprise clients, startup velocity. Fortune 100 contractual stakes with a small-team cadence — weekly delivery cycles, fast decision-making, and a playbook that evolves as the field does.
- Uncapped tooling and compute. The agent harness is the product. If the right answer is a bigger model, more infrastructure, or a tool we haven't adopted yet, we invest.
- Fully remote. Global team.
במקום לעבור לבד על אלפי מודעות, Jobify מנתחת את קורות החיים שלך ומציגה לך רק משרות שבאמת מתאימות לך.
מעל 80,000 משרות • 4,000 חדשות ביום
חינם. בלי פרסומות. בלי אותיות קטנות.