jobify_logo ×
  • מִשׁתַמֵשׁ
  • התחברות/הרשמה111
  • עמוד הבית
  • מי אנחנו
  • מעסיקים מובילים
  • פרסום משרה חינם
  • צרו קשר
  • תנאי שימוש
  • מדיניות פרטיות
  • הצהרת נגישות
קרן עזריאלי טקסט בעברית עם סמל אינסוף social_security the_israeli_employment_service work_office המקום
jobify_logo
  • מי אנחנו
  • מעסיקים מובילים
  • פרסום משרה חינם
  • צרו קשר
דילוג לתוכן

עדיין מחפשים עבודה במנועי חיפוש? הגיע הזמן להשתדרג!

במקום לעבור לבד על אלפי מודעות, Jobify מנתחת את קורות החיים שלך ומציגה לך רק משרות שבאמת מתאימות לך.

מעל 80,000 משרות • 4,000 חדשות ביום
חינם. בלי פרסומות. בלי אותיות קטנות.

Staff Site Reliability Engineer

Jobgether

Jobgether Jobgether

  • תל אביב - יפו
  • LinkedIn
LinkedIn

Staff Site Reliability Engineer

Jobgether

Jobgether Jobgether

  • תל אביב - יפו
  • bag_icon מלאה, עבודה מהבית
  • coins_icon 35,000-55,000 ₪ (הערכה מבוססת AI)
    הערכה מבוססת AI ולא שכר של המעסיק
  • LinkedIn
LinkedIn


This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff Site Reliability Engineer based in Israel.

This is a high-impact reliability leadership role within a fully remote engineering organization operating globally.

You will be the first dedicated SRE, helping establish reliability practices across multiple engineering teams and critical production systems.

The role combines hands-on engineering with organization-wide influence, covering observability, incident response, operational readiness, and resilience.

You will work closely with engineering leadership, infrastructure specialists, architects, and product teams to make reliability measurable and actionable.

A major focus will be embedding SRE principles into engineering culture rather than simply owning individual services.

You will also help shape how AI is used for incident investigation, operational tooling, observability, and safe system operations.

The position offers substantial autonomy to define standards, coach engineers, and build practices that scale with the organization.

Accountabilities

  • Define and implement SLIs and SLOs for critical production request paths, ensuring reliability objectives are visible, measurable, reviewed, and connected to engineering decisions.
  • Introduce and champion error budgets as a practical framework for balancing reliability investments with product and feature delivery.
  • Establish and maintain the reliability metrics used by engineering leadership to evaluate progress and identify areas requiring investment.
  • Strengthen the complete incident management lifecycle, including detection, response, communication, escalation, postmortems, and follow-up actions.
  • Improve alert quality, anomaly detection, escalation processes, and shared operational tooling in collaboration with infrastructure teams.
  • Lead reliability assessments for high-risk changes and new services, covering production readiness, capacity, failure modes, rollback strategies, and operational risks.
  • Introduce deliberate failure testing, game days, and chaos exercises to identify weaknesses and validate safe operational limits before incidents occur.
  • Work directly with engineering teams on complex reliability challenges through focused engagements, leaving behind stronger practices and clear ownership.
  • Coach Staff and Lead engineers to become reliability advocates within their respective teams and help establish distributed SRE ownership.
  • Develop lightweight, repeatable operational standards covering production readiness, on-call practices, runbooks, change safety, and service operability.
  • Partner with architects and technical leads to ensure reliability and failure tolerance are incorporated into system design rather than addressed after deployment.
  • Remain hands-on during production incidents and investigations, building tooling, dashboards, automation, and reference implementations where appropriate.
  • Promote effective use of AI for incident investigation, telemetry analysis, postmortem development, runbook creation, observability, and reliability tooling.
  • Help structure operational data, alerts, dashboards, and runbooks so that both engineers and AI agents can safely interpret and act on production signals.
  • Contribute production fixes and improvements directly through code and infrastructure changes rather than limiting the role to recommendations and reviews.

Requirements

  • 10+ years of engineering experience, including at least 3 years in SRE, production engineering, or a reliability-focused Staff Engineer role operating across multiple teams.
  • Demonstrated experience owning reliability at a platform or organizational level rather than only for an individual service.
  • Deep practical experience designing and implementing SLIs, SLOs, and error budgets, including successfully driving adoption across product and engineering teams.
  • Strong incident leadership experience, including managing high-severity, customer-facing incidents and leading effective postmortems that result in measurable improvements.
  • Advanced understanding of distributed-system failure modes, including database and cache saturation, cascading failures, retry storms, capacity constraints, graceful degradation, and load shedding.
  • Strong hands-on experience with Kubernetes, AWS, and modern observability platforms such as Datadog or comparable technologies.
  • Ability to read and write production code in Go, TypeScript, or a similar language, as well as work with infrastructure as code.
  • Demonstrated ability to influence teams without direct authority and successfully change engineering practices across an organization.
  • Strong coaching and mentoring skills, with evidence of developing engineers into effective reliability owners.
  • Exceptional written and verbal communication skills, with the ability to clearly communicate incidents, risks, technical trade-offs, and reliability priorities to both engineers and executives.
  • Strong preference for asynchronous, documented decision-making and clear technical communication.
  • Practical experience using AI tools for incident investigation, telemetry analysis, runbook and postmortem development, and engineering tooling.
  • Understanding of how operational data, alerts, dashboards, and runbooks should be structured to support safe AI-assisted diagnosis and operations.
  • Pragmatic approach to reliability, with the ability to balance operational risk, engineering investment, delivery speed, and business priorities.
  • Experience in fraud detection, identity, payments, or other real-time and adversarial environments is an asset.
  • Experience with multi-region architectures, cell-based architectures, or failure-isolation strategies is a plus.
  • Experience operating Elasticsearch, Redis, DynamoDB, or Kafka at scale and understanding their failure modes is beneficial.
  • Familiarity with FinOps and cloud infrastructure cost-versus-reliability trade-offs is an advantage.
  • Must be authorized to work from the hiring location; visa sponsorship is not provided.

Benefits

  • Fully remote working environment.
  • Opportunity to become the first dedicated Site Reliability Engineer and establish organization-wide reliability practices.
  • High level of autonomy and direct influence over engineering standards, operational practices, and platform reliability.
  • Opportunity to work across multiple engineering teams and critical production systems.
  • Close collaboration with engineering leadership, architects, infrastructure teams, and technical leads.
  • Opportunity to shape AI-assisted reliability practices and the future of production operations.
  • Strong focus on professional growth, technical leadership, coaching, and knowledge sharing.
  • Inclusive, globally distributed engineering environment that values diverse perspectives and backgrounds.
  • For US-based employees, the stated cash compensation range is $177,000–$240,000 USD, with actual offers varying according to factors such as experience, skills, education, certifications, and market conditions. Compensation may differ for other hiring locations.
  • Remote work eligibility is subject to applicable regulatory and security requirements in the candidate's location.

How Jobgether Works

We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.

We appreciate your interest and wish you the best!

Why Apply Through Jobgether?

Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.


במקום לעבור לבד על אלפי מודעות, Jobify מנתחת את קורות החיים שלך ומציגה לך רק משרות שבאמת מתאימות לך.

מעל 80,000 משרות • 4,000 חדשות ביום
חינם. בלי פרסומות. בלי אותיות קטנות.

שאלות ותשובות עבור משרת Staff Site Reliability Engineer

התפקיד המרכזי של מהנדס/ת אמינות אתר (Staff Site Reliability Engineer) ב-Jobgether הוא להוביל את נוהלי האמינות ברחבי הארגון, תוך שילוב עבודה מעשית עם השפעה רחבה על צוותי ההנדסה ומערכות הייצור הקריטיות. התפקיד כולל הגדרת מדדי אמינות (SLIs ו-SLOs), שיפור מחזור ניהול האירועים, והטמעת עקרונות SRE בתרבות ההנדסית, תוך שימוש בבינה מלאכותית לשיפור תהליכי חקירה ותפעול.

תפקיד Staff Site Reliability Engineer תורם לשיפור אמינות המערכות ב-Jobgether על ידי הגדרת יעדי אמינות מדידים, הטמעת תקציבי שגיאות, חיזוק תהליכי ניהול אירועים ותגובה, וביצוע הערכות אמינות לשינויים ושירותים חדשים. בנוסף, התפקיד כולל עבודה ישירה עם צוותי הנדסה על אתגרי אמינות מורכבים, אימון מהנדסים אחרים, וקידום שימוש יעיל בבינה מלאכותית לניתוח טלמטריה וחקירת אירועים.

כדי להצליח בתפקיד Staff Site Reliability Engineer ב-Jobgether, נדרשות לפחות 10 שנות ניסיון בהנדסה, מתוכן 3 שנים בתפקיד SRE או תפקיד דומה המתמקד באמינות ברמת פלטפורמה או ארגון. נדרש ניסיון מעשי עמוק בתכנון והטמעת SLIs, SLOs ותקציבי שגיאות, מנהיגות חזקה בניהול אירועים, הבנה מתקדמת של כשלים במערכות מבוזרות, וניסיון עם Kubernetes, AWS ופלטפורמות ניטור מודרניות. כמו כן, נדרשת יכולת כתיבת קוד ב-Go, TypeScript או שפה דומה, ויכולת השפעה על צוותים ללא סמכות ישירה.

משרות נוספות מומלצות עבורך
  • רשימת משאלות

    Staff Site Reliability Engineer

    • map_icon תל אביב - יפו
    Tapingo Ltd

    Tapingo Ltd

  • רשימת משאלות

    Staff Site Reliability Engineer

    • map_icon תל אביב - יפו
    Grubhub

    Grubhub

לכל המשרות של Staff Site Reliability Engineer

הכשרות רלוונטיות

ג׳ון ברייס

ג׳ון ברייס

קורס DevOps

  • ערב
ג׳ון ברייס

ג׳ון ברייס

קורס DevOps

  • ערב
NAYA College

NAYA College

מהנדס DevOps בסביבת הענן – Cloud DevOps Engineer – test only

ג׳ון ברייס

ג׳ון ברייס

קורס DevOps

  • בוקר

ניתן לצפות במשרות שסימנת בכל שלב תחת התפריט הראשי בקטגוריית 'משרות שאהבתי'

המקום קרן עזריאלי טקסט בעברית עם סמל אינסוף
  • מי אנחנו
  • מעסיקים מובילים
  • צרו קשר
  • תנאי שימוש
  • מדיניות פרטיות
  • הצהרת נגישות

2026 Ⓒ ג'וביפיי - כל הזכויות שמורות

קרן עזריאלי טקסט בעברית עם סמל אינסוף social_security the_israeli_employment_service israel_innovation_authority work_office המקום
המערכת בונה את הפרופיל התעסוקתי שלך

עוד רגע...

המערכת זיהתה ששינית את הנתונים באזור האישי ומעדכנת את ההמלצות על תפקידים ומשרות בהתאם.

מצטערים, לא הצלחנו לנתח בהצלחה את הנתונים שהזנת.
אתם מוזמנים לנסות להזין שוב או להעלות קובץ קורות חיים במידה ויש לכם.
בהצלחה

הגעת להגבלה היומית של שלושה עדכונים בפרופיל האישי ביום

loader

הבקשה שלך נשלחה בהצלחה!

יש באפשרותך לשלוח בקשה לקבלת ייעוץ אישי ללא עלות מיועצת קריירה.

באפשרותך לשלוח בקשה לקבלת ייעוץ אישי ללא עלות

  • בעיה טכנית

  • סיוע בכתיבת קורות חיים או בהכנה לראיון עבודה

  • התאמה של משרות

  • אחר:

פנייתך נשלחה בהצלחה. נציג מטעם ארגון נכי צהל ייצור איתך קשר בהקדם