עדיין מחפשים עבודה במנועי חיפוש? הגיע הזמן להשתדרג!
במקום לעבור לבד על אלפי מודעות, Jobify מנתחת את קורות החיים שלך ומציגה לך רק משרות שבאמת מתאימות לך.
מעל 80,000 משרות • 4,000 חדשות ביום
חינם. בלי פרסומות. בלי אותיות קטנות.
our company's Cybersecurity business unit is building an AI-powered security platform for the world's largest service provider networks. The platform combines a cloud-native microservices backend with self-hosted LLM inference, delivered into customers' own Kubernetes environments - often air-gapped, with no data leaving their premises.
Responsibilities
Own the platform's Helm charts and make them portable across any Kubernetes distribution, including OpenShift (non-root, restricted SCCs, NetworkPolicies).
Operate the self-hosted operator stack: CloudNativePG, Strimzi, ECK, MinIO, Redis, and Temporal.
Run the Nx monorepo pipelines and drive the ArgoCD GitOps rollout.
Build zero-downtime releases using blue/green and canary deployments, expand/contract migrations, and automated rollback.
Manage autoscaling with KEDA/HPA, backup/restore for stateful services, and load-validation of stamp sizing tiers.
Provide on-call support for dev/staging environments and customer stamps.
Build and maintain observability using Grafana Alloy, OpenTelemetry, and Prometheus/Grafana - dashboards and alerts across services, Kafka, databases, and GPUs.
Improve developer experience through a shared dev cluster with local-debug traffic interception (Telepresence / mirrord), targeting sub-15-minute onboarding.
Technical Skills
5+ years of DevOps / platform engineering experience in a product company.
Deep hands-on experience with Kubernetes and Helm, including shipping software onto clusters you don't control.
Production experience with GitOps (ArgoCD/Flux), infrastructure as code, and zero-downtime release engineering.
Experience running stateful workloads on Kubernetes via operators (Kafka, PostgreSQL, Elasticsearch).
Solid grasp of observability fundamentals: Prometheus, Grafana, OpenTelemetry.
Comfortable operating in air-gapped environments without managed cloud services.
Soft Skills
Strong cross-functional collaboration - works daily with AI/ML Ops, backend, research, and solutions teams on customer onboarding.
Comfortable owning production reliability and on-call responsibilities.
Nice to Have / Advantage
Experience with OpenShift, KEDA, and Temporal.
Hands-on AI inference infrastructure experience: deploying and operating LLM inference on Kubernetes with GPUs (vLLM, TGI, Triton, or similar).
GPU scheduling and node readiness, model packaging for offline delivery.
Experience with LLM gateways (LiteLLM) and LLM observability (Langfuse).
במקום לעבור לבד על אלפי מודעות, Jobify מנתחת את קורות החיים שלך ומציגה לך רק משרות שבאמת מתאימות לך.
מעל 80,000 משרות • 4,000 חדשות ביום
חינם. בלי פרסומות. בלי אותיות קטנות.
משרות נוספות מומלצות עבורך
-
Senior DevOps Engineer
-
תל אביב - יפו
Jobgether
-
-
Senior DevOps Engineer, AIOps
-
רעננה
Nvidia
-
-
Senior DevOps Engineer (Secrets Manager)
-
פתח תקווה
Cyber Ark Software Ltd
-
-
Principal DevOps Engineer (Cortex Agentix Endpoint Security)
-
תל אביב - יפו
Cyber Ark Software Ltd
-
-
Senior DevOps- 243492
-
יהוד-מונוסון
Experis Israel
-
-
Senior DevOps Engineer, AIOps
-
תל אביב - יפו
NVIDIA
-
ערב
באר שבע