עדיין מחפשים עבודה במנועי חיפוש? הגיע הזמן להשתדרג!
במקום לעבור לבד על אלפי מודעות, Jobify מנתחת את קורות החיים שלך ומציגה לך רק משרות שבאמת מתאימות לך.
מעל 80,000 משרות • 4,000 חדשות ביום
חינם. בלי פרסומות. בלי אותיות קטנות.
We are seeking a highly skilled Principal Staff SRE to join our dynamic Core Infrastructure team. This role will lead initiatives that improve the efficiency, performance, reliability, and evolution of infrastructure services across on-premises and cloud environments. Widely considered to be one of the technology world’s most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. As you plan your future, see what we can offer to you and your family
What You'll Be Doing
- Lead initiatives that transform the IT Compute Core architecture and build new infrastructure service offerings across on-premises and cloud environments.
- Design, scale, and deploy core infrastructure services including DNS, NTP/PTP, DHCP, and LDAP, with responsibility for performance, reliability, automation, monitoring, high availability, capacity planning, and lifecycle management at global scale.
- Define and implement service-efficiency metrics, and drive improvements through software and hardware optimization, including SR-IOV and DPU capabilities where appropriate.
- Apply technologies such as eBPF and XDP to improve observability, performance analysis, and DDoS-mitigation capabilities.
- Collect and analyze system and capacity data; develop enterprise-wide capacity plans and coordinate with management to implement appropriate changes.
- Develop and maintain tools for collecting, analyzing, and visualizing data for reporting, alerting, and monitoring.
- Collaborate with NVIDIA leadership, senior engineers, program managers, and product managers to develop compelling IT products and services that meet customer needs.
- Bachelor's degree in Engineering, Computer Science, Mathematics, or a related field, or equivalent experience.
- 12+ years of demonstrable experience in compute platform engineering, with a strong focus on automation and technical leadership in large-scale environments.
- Experience designing and deploying containerization architectures and distributed-systems infrastructure.
- Demonstrable experience evaluating existing application architectures and seeing opportunities for containerization that improve scalability, reliability, and efficiency.
- Strong analytical skills, including the ability to define and track key performance metrics.
- Experience developing tools for data analysis and performance profiling, including Terraform and configuration-management tools.
- Proficiency in Go, Python, or similar programming languages.
- Linux OS proficiency, including kernel internals.
- Experience operating large-scale environments with bare-metal build infrastructure.
- Understanding of network protocols and architectures, including VLAN, VXLAN, SDN, BGP, and Anycast.
- Deep understanding of complementary infrastructure components, including DNS, LDAP, and security tools.
- Hands-on experience with containers and their implementation.
- Experience deploying and managing services such as DNS and LDAP at scale.
- Solid understanding of microservices architecture, infrastructure as code, and configuration-management tools.
, , JR2024251
במקום לעבור לבד על אלפי מודעות, Jobify מנתחת את קורות החיים שלך ומציגה לך רק משרות שבאמת מתאימות לך.
מעל 80,000 משרות • 4,000 חדשות ביום
חינם. בלי פרסומות. בלי אותיות קטנות.
שאלות ותשובות עבור משרת Senior Site Reliability Engineer
מהנדס/ת אמינות אתרים בכיר/ה ב-NVIDIA מוביל/ה יוזמות לשיפור היעילות, הביצועים והאמינות של שירותי תשתית. התפקיד כולל תכנון, הרחבה ופריסה של שירותי תשתית ליבה כמו DNS, NTP/PTP, DHCP ו-LDAP, תוך אחריות על ביצועים, אמינות, אוטומציה, ניטור, זמינות גבוהה ותכנון קיבולת בקנה מידה גלובלי. העבודה מתבצעת בסביבות מקומיות (on-premises) ובענן.
משרות נוספות מומלצות עבורך
-
Senior Site Reliability Engineer
-
יקנעם עילית
NVIDIA
-
-
Senior Site Reliability Engineering - Storage
-
יקנעם עילית
NVIDIA
-
-
Senior Site Reliability Engineer (Cortex)
-
תל אביב - יפו
Cyber Ark Software Ltd
-
-
Senior Site Reliability Engineer (SRE)
-
תל אביב - יפו
Taboola
-
-
Senior Site Reliability Engineer
-
פתח תקווה
Oracle
-
-
Senior Platform Reliability Engineer
-
תל אביב - יפו
IgniteTech
-
ערב
באר שבע