עדיין מחפשים עבודה במנועי חיפוש? הגיע הזמן להשתדרג!
במקום לעבור לבד על אלפי מודעות, Jobify מנתחת את קורות החיים שלך ומציגה לך רק משרות שבאמת מתאימות לך.
מעל 80,000 משרות • 4,000 חדשות ביום
חינם. בלי פרסומות. בלי אותיות קטנות.
Ramat Gan, Israel · Full-time · Intermediate
About The Position
Coralogix is a modern, full-stack observability platform transforming how businesses process and understand their data. Our unique architecture powers in-stream analytics without reliance on expensive indexing or hot storage. We specialize in comprehensive monitoring of logs, metrics, trace, and security events with features such as APM, RUM, SIEM, Kubernetes monitoring, and more, all enhancing operational efficiency and reducing observability spending by up to 70%.
Coralogix is seeking a Site Reliability Engineer to join our team and help build the next generation of our stream-based observability platform. We deliver real-time data analytics at scale to some of the world’s leading tech companies
The SRE team is responsible for the foundational infrastructure that powers Coralogix:
Kubernetes Infrastructure: Managing over 10,000 nodes across multiple cloud providers and regions. Coralogix production is 100% Kubernetes based
Support service owners in running over 1,000 instances of multiple datastore types, cloud-based and self-hosted (on Kubernetes).
Maintaining critical, large-scale clusters processing billions of events per second.
Automation & Operators: Building and maintaining both open-source and custom Kubernetes operators to manage complex stateful workloads like Kafka. Our tech stack is constantly evolving. It includes: Kubernetes, Go (Golang), AWS, GCP, Kafka, Istio, and more.
Write and maintain Kubernetes Controllers using frameworks like controller-runtime and KubeBuilder
Responsibilities
Act as a hands-on technical leader with deep expertise in relational DBs or other distributed datastores
Serve as a go-to person in the team — leading through influence, not hierarchy.
Collaborate cross-functionally to refine requirements and propose innovative, scalable solutions.
Drive long-term, high-impact infrastructure projects across multiple teams, from design to implementation, within defined timelines.
Contribute to improving system reliability, performance, and cost-efficiency at scale.
Requirements
5+ years of experience in DevOps, SRE, platform engineering, or infrastructure roles.
Experience with maintaining datastores in high-scale environments, whether relational DBs or other distributed datastores (Mongo, Cassandra, ClickHouse, Redis, OpenSearch).
Experience with running stateful workloads in Kubernetes.
Proven experience managing large-scale cloud infrastructure (AWS, GCP, etc.).
Experience in incident response and troubleshooting complex distributed systems.
Some software engineering experience, preferably in Golang.
Passion for automation, performance tuning, and operational excellence.
Cultural Fit
We’re seeking candidates who are hungry, humble, and smart. Coralogix fosters a culture of innovation and continuous learning, where team members are encouraged to challenge the status quo and contribute to our shared mission. If you thrive in dynamic environments and are eager to shape the future of observability solutions, we’d love to hear from you.
Coralogix is an equal opportunity employer and encourages applicants from all backgrounds to apply.
במקום לעבור לבד על אלפי מודעות, Jobify מנתחת את קורות החיים שלך ומציגה לך רק משרות שבאמת מתאימות לך.
מעל 80,000 משרות • 4,000 חדשות ביום
חינם. בלי פרסומות. בלי אותיות קטנות.
שאלות ותשובות עבור משרת Site Reliability Engineer – Datastores
מהנדס/ת אמינות אתר (SRE) ב-Coralogix מצטרף/ת לצוות האחראי על תשתית הליבה המניעה את פלטפורמת ה-Observability של החברה. התפקיד כולל ניהול תשתית Kubernetes עם למעלה מ-10,000 צמתים, תמיכה בבעלי שירותים המפעילים למעלה מ-1,000 מופעים של סוגי מאגרי נתונים שונים (מבוססי ענן ובאירוח עצמי), ותחזוקת אשכולות קריטיים בקנה מידה גדול המעבדים מיליארדי אירועים בשנייה. כמו כן, התפקיד כולל בנייה ותחזוקה של אופרטורים של Kubernetes לניהול עומסי עבודה מורכבים בעלי מצב (stateful workloads) כמו Kafka.
משרות נוספות מומלצות עבורך
-
Site Reliability Engineer (SRE)
-
תל אביב - יפו
QualiGate
-
-
Site Reliability Engineer - Datastores
-
רמת גן
Coralogix
-
-
Production Opertaions Engineer
-
תל אביב - יפו
Torq
-
-
Senior DevOps SRE Engineer
-
תל אביב - יפו
Axon
-
-
Site Reliability Engineer (SRE)
-
תל אביב - יפו
Paragon
-
-
Solutions / Site Reliability Engineer
-
תל אביב - יפו
APERIO
-
ערב
באר שבע