עדיין מחפשים עבודה במנועי חיפוש? הגיע הזמן להשתדרג!
במקום לעבור לבד על אלפי מודעות, Jobify מנתחת את קורות החיים שלך ומציגה לך רק משרות שבאמת מתאימות לך.
מעל 80,000 משרות • 4,000 חדשות ביום
חינם. בלי פרסומות. בלי אותיות קטנות.
This is a hands-on engineering role, not a research role. You will partner with product, platform, data, security, support, and incident response teams to make production systems and AI-powered experiences more resilient. You will use software engineering, infrastructure as code, SLOs, telemetry, provider observability, and automation as your main tools, and you will apply AI where it creates measurable reliability value rather than novelty.
This position is based out of our new Tel Aviv office.
What You'll Do:
Support AI-based application solutions where reliability matters. Partner with the development teams building AI-powered travel experiences to support the development and production operation of their solution.
Work with AI solutions, providers, and APIs. Partner with teams integrating AI capabilities and providers, with attention to API reliability, authentication, quotas, rate limits, latency and provider-specific operational constraints.
Troubleshoot AI tools and provider issues. Diagnose failures across AI-powered workflows, provider APIs, configuration, permission errors, degraded responses and related areas.
Operate reliable production platforms. implement and run cloud infrastructure,and help product teams move quickly without compromising reliability.
Improve observability. Build dashboards, alerts, traces, logs, and runbooks that make service health clear, actionable, and tied to SLOs and customer impact.
Apply AI to SRE workflows. Prototype and productionize AI-assisted systems that create effective and efficient operations
Automate operational toil. Create tools, workflows, and automation that remove repetitive manual work and make operational knowledge easier to use.
5+ years of experience as a Senior SRE, Infrastructure Software Engineer, Production Engineer, or DevOps Engineer.
3+ years of experience operating production, 24x7 customer-facing systems.
Hands-on experience delivering production infrastructure, platform tooling, and automation used by engineering teams.
Strong software engineering skills in Python, Go, Java, or a similar language, with a bias toward production-quality code, tests, monitoring, and documentation.
Experience with cloud infrastructure, container orchestration, Linux systems, networking, CI/CD, and infrastructure as code such as Terraform or CloudFormation.
Experience building, tuning, and automating observability systems such as Grafana, Prometheus, New Relic, Datadog, Splunk, or similar tools.
Familiarity with SLOs, incident response, on-call practices, root cause analysis, and blameless postmortems.
Practical experience or strong interest in AI solutions, AI providers, agents, AI APIs, provider integrations, or AI-assisted internal tools.
Ability to troubleshoot AI tools and provider/API issues, including rate limits, quota, auth, permission errors, latency, SDK or API contract changes, content quality issues, and service degradations.
Excellent communication skills and the ability to work with stakeholders and domain experts across the company.
במקום לעבור לבד על אלפי מודעות, Jobify מנתחת את קורות החיים שלך ומציגה לך רק משרות שבאמת מתאימות לך.
מעל 80,000 משרות • 4,000 חדשות ביום
חינם. בלי פרסומות. בלי אותיות קטנות.