עדיין מחפשים עבודה במנועי חיפוש? הגיע הזמן להשתדרג!
במקום לעבור לבד על אלפי מודעות, Jobify מנתחת את קורות החיים שלך ומציגה לך רק משרות שבאמת מתאימות לך.
מעל 80,000 משרות • 4,000 חדשות ביום
חינם. בלי פרסומות. בלי אותיות קטנות.
What you'll be doing:
Drive end-to-end performance strategy, characterization, test plans, and optimization for next-generation our AI GPU clusters, focusing on large-scale distributed training and inference workloads.
Deeply evaluate and optimize our Networking core technologies performance, including RDMA/PRDMA, networking protocols, collective communication (NCCL), congestion control, and load-balancing algorithms.
Work on performance research and analysis of NVIDIA DPUs and storage technologies in North-South (N-S) use cases and deployment scenarios to maximize performance and efficiency for AI inference jobs.
Drive the strategy for performance observability and dashboards across next-generation NVIDIA data center solutions and supercomputers by leveraging scalable, streamlined telemetry pipelines to build performance dashboards and automated analytics based on real-time performance metrics across NICs, Switches, GPUs, and NVLink boundaries.
Perform deep root-cause analysis (RCA) on complex multi-node performance bottlenecks, driving actionable mitigation plans across hardware, firmware, and software teams.
What we need to see:
B.Sc. or M.Sc. in Computer Science, Computer Engineering, Software Engineering, or equivalent technical experience.
8+ overall years of experience and deep expertise in High Performance Networking, RDMA, and Systems level performance.
3+ years of experience as an engineering team manager leading technical performance or R&D teams.
Hands-on experience analyzing and optimizing collective communication (e.g., NCCL, MPI) and network traffic patterns for large-scale distributed AI workloads (LLM training and inference).
Hands-on experience designing, deploying, and customizing Grafana dashboards for cluster monitoring, alerting, and data visualization.
Exceptional cross-team leadership, analytical thinking, and communication skills to drive alignment across hardware, software, and architecture groups.
Ways to stand out from the crowd:
Proven track record of optimizing NCCL, RDMA/RoCEv2, and custom collective algorithms specifically tailored for multi-thousand GPU deployments running LLMs or Mixture-of-Experts (MoE) architectures.
Deep experience tuning advanced network traffic mechanisms such as adaptive routing, PFC/ECN congestion control, and packet-spraying technologies.
Experience building autonomous performance-driven tools, AI-assisted root cause analysis agents, or automated regression frameworks for continuous cluster-level performance evaluation.
Hands-on experience developing custom Grafana plugins, complex dashboard panels, or integrated alert management workflows using PromQL/LogQL for hyperscale or HPC environments.
במקום לעבור לבד על אלפי מודעות, Jobify מנתחת את קורות החיים שלך ומציגה לך רק משרות שבאמת מתאימות לך.
מעל 80,000 משרות • 4,000 חדשות ביום
חינם. בלי פרסומות. בלי אותיות קטנות.
משרות נוספות מומלצות עבורך
-
Host and Systems Performance Senior Manager
-
יקנעם עילית
NVIDIA
-
-
Host and Systems Performance Senior Manager
-
יקנעם עילית
Nvidia
-
-
Manager, Performance Research and Analysis
-
תל אביב - יפו
NVIDIA AI
-
-
Manager, Performance Research and Analysis
-
יקנעם עילית
NVIDIA
-
-
Manager, Performance Research and Analysis
-
תל אביב - יפו
NVIDIA
-
-
Performance Engineering Manager
-
תל אביב - יפו
Check Point Software
-