עדיין מחפשים עבודה במנועי חיפוש? הגיע הזמן להשתדרג!
במקום לעבור לבד על אלפי מודעות, Jobify מנתחת את קורות החיים שלך ומציגה לך רק משרות שבאמת מתאימות לך.
מעל 80,000 משרות • 4,000 חדשות ביום
חינם. בלי פרסומות. בלי אותיות קטנות.
Requirements
- B.Sc. or higher in Computer Science, Computer Engineering, Electrical Engineering, or a closely related field
- 8+ years of experience in systems software, distributed computing, or AI infrastructure, with 3+ years in a leadership or team lead role
- Deep expertise in large-scale communication systems: collective communication, RDMA, network topology-aware routing, and bandwidth optimization
- Hands-on experience building software infrastructure for distributed training on custom accelerators or heterogeneous hardware (GPU, NPU, TPU)
- Strong knowledge of runtime systems: scheduling, execution graphs, kernel dispatch, synchronization primitives, and pipeline management
- Experience with memory management at scale: activation checkpointing, tensor offloading, rematerialization, KV cache management
- Proficiency in C/C++ and Python, with a focus on high-performance, production-quality code in Linux environments
- Proven ability to define technical vision, lead multi-person projects end-to-end, and deliver results under research and engineering timelines
- Excellent communication skills in English — confident presenting to international audiences, writing technical reports, and driving cross-team alignment
- Strong collaborative mindset and experience working in globally distributed, multicultural teams
Ways to Stand Out From the Crowd
M.Sc. or Ph.D. in a relevant field, with a strong publication record at systems or ML venues (EuroSys, OSDI, SC, NeurIPS, MLSys, ISCA)
- Hands-on experience with communication frameworks such as NCCL, MPI, HCCL, or UCX
- Experience with compiler and graph optimization for AI workloads (XLA, TVM, Triton, or custom operator fusion)
- Background in mixed-precision training, model parallelism (Tensor Parallelism, Pipeline Parallelism, Expert Parallelism), and large model co-design
- Experience profiling and debugging performance bottlenecks on heterogeneous clusters using tools like Chrome tracing, nsight, or custom profilers
במקום לעבור לבד על אלפי מודעות, Jobify מנתחת את קורות החיים שלך ומציגה לך רק משרות שבאמת מתאימות לך.
מעל 80,000 משרות • 4,000 חדשות ביום
חינם. בלי פרסומות. בלי אותיות קטנות.