עדיין מחפשים עבודה במנועי חיפוש? הגיע הזמן להשתדרג!
במקום לעבור לבד על אלפי מודעות, Jobify מנתחת את קורות החיים שלך ומציגה לך רק משרות שבאמת מתאימות לך.
מעל 80,000 משרות • 4,000 חדשות ביום
חינם. בלי פרסומות. בלי אותיות קטנות.
We seek a versatile Senior Software Engineer who is passionate about performance optimization and generative AI. Our team brings the latest research in LLM inference - from novel decoding strategies to quantization schemes - into production across our hardware lineup, from large data center servers to powerful edge devices. We work on the most advanced architectures in the field, with a focus on NVIDIA's own.
What you'll be doing:
Implement and optimize inference algorithms for LLM and omnimodal architectures, including hybrid Mamba-Transformer and mixture-of-experts models.
Profile inference pipelines using NVIDIA's profiling and simulation tools. Correlate simulation predictions against real hardware across data center and edge devices.
Write and tune GPU kernels (CUDA, Triton) for operators like fused MoE layers, SSM state updates, and quantized GEMMs.
Solve distributed inference problems: expert parallelism, communication-compute overlap, collective tuning, multi-node deploym
דרישות:
What we need to see:
B.Sc., M.Sc., or equivalent experience in Computer Science or Computer Engineering.
5+ years of hands-on software engineering experience in performance-critical systems.
Solid understanding of deep learning architectures (Transformers, SSMs, MoE, ).
Experience with systems where hardware constraints matter: GPU programming, memory hierarchy, networking, or distributed computing.
Strong software engineering fundamentals: clean design, extensibility, testability. Good judgment about when complexity is warranted.
Effective communicator who works well across teams and time zones.
Experience optimizing deep learning workloads on our GPUs using roofline models, Nsight/PyTorch profilers and end-to-end traces.
Ways to stand out from the crowd:
Contributions to open-source inference runtimes and libraries - vLLM, SGLang, FlashInfer, Dynamo or similar.
Hands-on work with LLM quantization (FP8, NVFP4, MXFP8, mixed-precision) and practical understanding of numerica
במקום לעבור לבד על אלפי מודעות, Jobify מנתחת את קורות החיים שלך ומציגה לך רק משרות שבאמת מתאימות לך.
מעל 80,000 משרות • 4,000 חדשות ביום
חינם. בלי פרסומות. בלי אותיות קטנות.
אונליין
אונליין