עדיין מחפשים עבודה במנועי חיפוש? הגיע הזמן להשתדרג!
במקום לעבור לבד על אלפי מודעות, Jobify מנתחת את קורות החיים שלך ומציגה לך רק משרות שבאמת מתאימות לך.
מעל 80,000 משרות • 4,000 חדשות ביום
חינם. בלי פרסומות. בלי אותיות קטנות.
In addition, this role will play a key leadership position in building scalable telemetry frameworks, performance dashboards, and job-level monitoring solutions to enable continuous performance tracking and root cause analysis across our supercomputing environments. The position also includes deep ownership of competitive benchmarking and performance analysis. You will work closely with a wide range of our hardware and software platforms, including HCAs, DPUs, switches, CPUs, GPUs, and full system architectures, across multiple networking stacks and performance-critical software layers.
What you'll be doing:
Lead performance research and evaluation of advanced networking technologies supporting AI workloads, including LLM training and inference at supercomputing scale.
Define end-to-end performance test plans and methodology for next-generation Networking HW and networking technologies, including performance expectations and target KPIs.
Drive benchmarking, profiling, reporting, and deep performance characterization of networking workloads and offload features.
Collaborate closely with simulation, architecture, chip-design, firmware, and software teams to assess performance tradeoffs and identify bottlenecks.
Perform deep root cause analysis (RCA) for performance gaps and stability issues, and drive cross-team mitigation plans.
Develop and enhance performance analysis tools, automation frameworks, and scalable methodologies for cluster-level performance evaluation.
Own performance observability efforts, including telemetry pipelines, dashboards, and job-level performance analytics.
B.Sc in Computer Science or Software Engineering
5+ years of experience with high-performance Networking technologies (RDMA, Storage, Security, OVS, MPI)
3+ years as an engineering team manager
Demonstrated Performance Analysis skills and methodologies.
Experience with Cluster level performance, Telemetry, NIC, DPUs, Switches, and GPUs.
Fast and self-learning capabilities with strong analytical and problem solving skills
Programming Languages: Python, Bash and C/C++ languages
Experience with Linux OS distros
Team player and a leader with good communication and interpersonal skills
Ways to stand out from the crowd:
Deep system-level architecture knowledge (Intel / AMD / ARM CPUs, NVIDIA GPUs, HCA/DPU architecture, memory subsystems, PCIe, storage, NVLink).
Strong expertise in RDMA networking performance and AI communication stacks (e.g., NCCL).
Proven experience analysing AI workload communication patterns and benchmarking distributed LLM training workloads at scale.
Experience designing telemetry frameworks, monitoring pipelines, and performance dashboards for large clusters.
Familiarity with modern AI tooling including performance-driven agents, automation pipelines, and RAG-based applications.
במקום לעבור לבד על אלפי מודעות, Jobify מנתחת את קורות החיים שלך ומציגה לך רק משרות שבאמת מתאימות לך.
מעל 80,000 משרות • 4,000 חדשות ביום
חינם. בלי פרסומות. בלי אותיות קטנות.