עדיין מחפשים עבודה במנועי חיפוש? הגיע הזמן להשתדרג!
במקום לעבור לבד על אלפי מודעות, Jobify מנתחת את קורות החיים שלך ומציגה לך רק משרות שבאמת מתאימות לך.
מעל 80,000 משרות • 4,000 חדשות ביום
חינם. בלי פרסומות. בלי אותיות קטנות.
DeepRec.ai is partnered with a global enterprise software company that is rethinking how reliability engineering operates in the age of AI. Rather than scaling teams through headcount, they're building an AI-native operating model where autonomous agents perform much of the operational workload, allowing engineers to focus on governance, architecture, resilience, and customer outcomes.
This is a rare opportunity to lead the evolution of AI-driven operations at enterprise scale.
As the senior leader responsible for reliability, platform operations, and incident management, you will own uptime, operational excellence, and customer trust across a large-scale SaaS platform serving enterprise customers globally.
You'll lead a highly experienced distributed team of SREs, Platform Engineers, and DevOps specialists while remaining deeply hands-on yourself. This is not a role for someone who wants to manage through dashboards and status meetings. The successful candidate will be actively involved in major incidents, reliability strategy, automation initiatives, and AI agent development.
A significant part of the role involves designing and governing the AI systems that support operational workflows including:
- Incident triage and diagnosis
- Intelligent alerting and monitoring
- Change validation and deployment controls
- Root cause analysis generation
- Auto-remediation and self-healing workflows
- Continuous operational learning and improvement
Key Responsibilities
- Own reliability, availability, and operational performance across a large-scale multi-tenant SaaS platform.
- Lead executive-level customer escalations and critical incidents when required.
- Define and evolve an AI-first operational model focused on automation, resilience, and efficiency.
- Establish governance, quality standards, and guardrails for AI-driven operational workflows.
- Drive measurable improvements in uptime, MTTR, operational efficiency, and customer satisfaction.
- Recruit, develop, and retain a high-calibre team of senior engineers.
- Partner closely with Engineering, Product, and Customer Success leadership to ensure operational excellence remains a competitive advantage.
- Remain hands-on with incident response, root cause analysis, architecture reviews, and automation initiatives.
What We're Looking For
- 10+ years operating and scaling complex SaaS environments.
- Experience leading Site Reliability Engineering, Platform Engineering, Infrastructure, or Production Operations functions.
- Proven success managing senior engineering teams at VP, SVP, or Head of level.
- Deep AWS expertise across large-scale production environments.
- Strong background in incident management, observability, automation, and distributed systems.
- Demonstrated experience implementing AI-driven operations, AIOps, agentic workflows, or automated remediation systems.
- Comfortable writing code, reviewing architecture, and contributing technically when needed.
- Track record of improving reliability and operational outcomes without simply increasing team size.
- Excellent communication skills with the ability to engage both executive stakeholders and enterprise customers.
Ideal Background
We're particularly interested in leaders who have experience within:
- Enterprise SaaS
- Developer Platforms
- Observability & Monitoring
- Cloud Infrastructure
- Customer Experience Platforms
- Collaboration Software
- AI Infrastructure & Automation
- High-availability B2B software environments
Why This Role?
- Shape one of the most advanced AI-native operations functions currently being built.
- Work with a highly experienced, globally distributed engineering team.
- Operate at meaningful enterprise scale with significant autonomy.
- Influence the future direction of AI-powered reliability engineering.
- Fully remote environment with global hiring scope and overlap with US business hours.
Fixed $400,000 USD base salary globally. There is no bonus or equity component.
If you're passionate about the intersection of AI, automation, platform reliability, and operational leadership - and want to help define what modern SRE looks like over the next decade - we'd love to speak with you.
במקום לעבור לבד על אלפי מודעות, Jobify מנתחת את קורות החיים שלך ומציגה לך רק משרות שבאמת מתאימות לך.
מעל 80,000 משרות • 4,000 חדשות ביום
חינם. בלי פרסומות. בלי אותיות קטנות.