עדיין מחפשים עבודה במנועי חיפוש? הגיע הזמן להשתדרג!
במקום לעבור לבד על אלפי מודעות, Jobify מנתחת את קורות החיים שלך ומציגה לך רק משרות שבאמת מתאימות לך.
מעל 80,000 משרות • 4,000 חדשות ביום
חינם. בלי פרסומות. בלי אותיות קטנות.
Job Title: Site Reliability Engineer (SRE)
Duration: 6+ months
Location -Haifa, Haifa Disctrict, Israel. (Hybrid 2-3 days in a week)
5 years’ experience
On-prem infrastructure management
- Manage on-prem infrastructure. Maintain uptime, reliability and readiness of on-prem engineering cloud spread across multiple data centers.
Guard SLAs
- Guard service level agreements (SLAs) for critical engineering services. Implement monitoring, alerting, and incident response procedures to ensure adherence to defined performance targets. Perform root cause analysis and post-mortems of incidents for any threshold breaches.
Observability
- Set up and manage monitoring and logging tools such as Prometheus, Grafana, or the ELK Stack to oversee system health and performance. Maintain KPI pipelines using Jenkins, Python and ELK.
- Improve monitoring systems by adding custom alerts based on business needs.
Automation & Optimization
- Help in capacity planning, optimization and better utilization efforts.
Day-to-Day Support
- Support user reported issues & issues. Monitor alerts and take necessary action.
- Actively participate in WAR room for critical issues
Collaboration & Documentation
- Create and maintain documentation for operational procedures, configurations, and troubleshooting guides.
Tech stack
- Baremetal data center machine management tools like IPMI, Redfish, KVM etc.
- Automation using Jenkins, Python, Go, Bash.
- Infrastructure tools like Kubernetes, MySQL, Prometheus, Grafana and ELK.
- Any familiarity with hardware like GPU & Tegras is a plus
במקום לעבור לבד על אלפי מודעות, Jobify מנתחת את קורות החיים שלך ומציגה לך רק משרות שבאמת מתאימות לך.
מעל 80,000 משרות • 4,000 חדשות ביום
חינם. בלי פרסומות. בלי אותיות קטנות.
שאלות ותשובות עבור משרת Site Reliability Engineer
מהנדס/ת Site Reliability Engineer ב-Tekgence Inc. יהיה/תהיה אחראי/ת על ניהול תשתית ה-On-prem, שמירה על זמינות, אמינות ומוכנות של ענן ההנדסה הפרוס על פני מספר מרכזי נתונים, וכן על אבטחת הסכמי רמת שירות (SLAs) לשירותי הנדסה קריטיים.
משרות נוספות מומלצות עבורך
-
Site Reliability Engineer (SRE) On-Prem
-
תל אביב - יפו
Dream
-
-
Sr. SRE AI Engineer
-
תל אביב - יפו
Navan
-
-
Service Reliability Engineer MI
-
אור יהודה
AudioCodes
-
-
Agentic AI Developer
-
רמת גן
Empire Media Network
-
-
Agent Operations Engineer
-
רמת גן
Empire Media Network
-
-
Site Reliability Engineer
-
תל אביב - יפו
BMC Software
-
ערב
באר שבע