עדיין מחפשים עבודה במנועי חיפוש? הגיע הזמן להשתדרג!
במקום לעבור לבד על אלפי מודעות, Jobify מנתחת את קורות החיים שלך ומציגה לך רק משרות שבאמת מתאימות לך.
מעל 80,000 משרות • 4,000 חדשות ביום
חינם. בלי פרסומות. בלי אותיות קטנות.
We seek a versatile Senior Software Engineer who is passionate about performance optimization and generative AI. Our team brings the latest research in LLM inference - from novel decoding strategies to quantization schemes - into production across our hardware lineup, from large data center servers to powerful edge devices. We work on the most advanced architectures in the field, with a focus on NVIDIA's own.
What you'll be doing:
Implement and optimize inference algorithms for LLM and omnimodal architectures, including hybrid Mamba-Transformer and mixture-of-experts models.
Profile inference pipelines using NVIDIA's profiling and simulation tools. Correlate simulation predictions against real hardware across data center and edge devices.
Write and tune GPU kernels (CUDA, Triton) for operators like fused MoE layers, SSM state updates, and quantized GEMMs.
Solve distributed inference problems: expert parallelism, communication-compute overlap, collective tuning, multi-node deploym
דרישות:
What we need to see:
B.Sc., M.Sc., or equivalent experience in Computer Science or Computer Engineering.
5+ years of hands-on software engineering experience in performance-critical systems.
Solid understanding of deep learning architectures (Transformers, SSMs, MoE, ).
Experience with systems where hardware constraints matter: GPU programming, memory hierarchy, networking, or distributed computing.
Strong software engineering fundamentals: clean design, extensibility, testability. Good judgment about when complexity is warranted.
Effective communicator who works well across teams and time zones.
Experience optimizing deep learning workloads on our GPUs using roofline models, Nsight/PyTorch profilers and end-to-end traces.
Ways to stand out from the crowd:
Contributions to open-source inference runtimes and libraries - vLLM, SGLang, FlashInfer, Dynamo or similar.
Hands-on work with LLM quantization (FP8, NVFP4, MXFP8, mixed-precision) and practical understanding of numerica
במקום לעבור לבד על אלפי מודעות, Jobify מנתחת את קורות החיים שלך ומציגה לך רק משרות שבאמת מתאימות לך.
מעל 80,000 משרות • 4,000 חדשות ביום
חינם. בלי פרסומות. בלי אותיות קטנות.
שאלות ותשובות עבור משרת Senior Software Engineer, LLM Inference
כמהנדס תוכנה בכיר בהסקה של מודלי שפה גדולים ב-Nvidia, תהיה אחראי על יישום ואופטימיזציה של אלגוריתמי הסקה עבור ארכיטקטורות LLM ואומנימודליות, כולל מודלים היברידיים של Mamba-Transformer ומודלים של Mixture-of-Experts. התפקיד כולל גם פרופיל צינורות הסקה באמצעות כלי פרופיל וסימולציה של Nvidia, כתיבה וכוונון של ליבות GPU (CUDA, Triton), ופתרון בעיות הסקה מבוזרות על פני מגוון רחב של חומרות, משרתי דאטה סנטר ועד התקני קצה.
משרות נוספות מומלצות עבורך
-
Software Engineer (AI Training)
-
תל אביב - יפו
Alignerr
-
-
מפתח/ת תוכנה
-
חיבת ציון
Bluebird Aero Systems
-
-
Software Engineer III, Database Migration, Cloud
-
תל אביב - יפו
Google
-
-
Backend Software Engineer II and Senior Software Engineer
-
הרצליה
מיקרוסופט ישראל
-
-
Platform Enablement & AI Engineer
-
הרצליה
Payoneer
-
-
מפתח /ת תוכנת מבדקים מנוסה? זה הזמן לאתגר הבא שלך!
-
באר יעקב
Experis
-
אונליין
אונליין