jobify_logo ×
  • מִשׁתַמֵשׁ
  • התחברות/הרשמה
  • עמוד הבית
  • מי אנחנו
  • מעסיקים מובילים
  • פרסום משרה חינם
  • צרו קשר
  • תנאי שימוש
  • מדיניות פרטיות
  • הצהרת נגישות
קרן עזריאלי טקסט בעברית עם סמל אינסוף social_security the_israeli_employment_service work_office המקום
jobify_logo
  • מי אנחנו
  • מעסיקים מובילים
  • פרסום משרה חינם
  • צרו קשר
דילוג לתוכן

עדיין מחפשים עבודה במנועי חיפוש? הגיע הזמן להשתדרג!

במקום לעבור לבד על אלפי מודעות, Jobify מנתחת את קורות החיים שלך ומציגה לך רק משרות שבאמת מתאימות לך.

מעל 80,000 משרות • 4,000 חדשות ביום
חינם. בלי פרסומות. בלי אותיות קטנות.

AI Performance Engineer JB-5254

Recruitx

Recruitx Recruitx

  • תל אביב - יפו
  • LinkedIn
LinkedIn

AI Performance Engineer JB-5254

Recruitx

Recruitx Recruitx

  • תל אביב - יפו
  • bag_icon מלאה
  • coins_icon 25,000-35,000 ₪ (הערכה מבוססת AI)
    הערכה מבוססת AI ולא שכר של המעסיק
  • LinkedIn
LinkedIn


• We are looking for an AI Performance Engineer to help optimize large language model inference across the full systems stack. You will work on improving the performance, efficiency, scalability, and reliability of LLM serving infrastructure, from low-level GPU kernel optimization to single-node runtime performance and distributed cluster-wide inference optimization.

• This role sits at the intersection of systems engineering, GPU computing, and AI infrastructure. You will work closely with researchers, ML engineers, infrastructure engineers, and product teams to make state-of-the-art AI models faster, more efficient, and easier to serve at scale.

• What You’ll Do

• Optimize LLM inference performance across the full stack, including kernels, runtimes, model execution, networking, scheduling, and distributed serving.

• Analyze and improve GPU utilization, memory bandwidth, latency, throughput, and cost efficiency.

• Develop, tune, or integrate high-performance GPU kernels using technologies such as CUDA, Triton, CUTLASS, or similar frameworks.

• Improve single-node inference performance through runtime optimization, memory management, batching, quantization, parallelism, and profiling.

• Optimize distributed inference across clusters, including tensor parallelism, pipeline parallelism, expert parallelism, communication patterns, load balancing, and scheduling.

• Build tools, benchmarks, and profiling workflows to identify bottlenecks and measure performance improvements.

• Collaborate with ML, infrastructure, and product teams to translate performance improvements into production impact.

• Stay current with advances in LLM serving, GPU architectures, compiler/runtime systems, and AI infrastructure.

• What We’re Looking For We are looking for candidates with strong experience in at least one of the following areas:

• Performance engineering and low-level software optimization.

• GPU kernel development and optimization.

• AI/ML systems optimization, especially for inference or distributed training/serving.

• Ideal Qualifications

• Strong programming skills in C++, CUDA, Python, or similar systems-oriented languages.

• Experience profiling and optimizing software for latency, throughput, memory usage, or hardware utilization.

• Familiarity with GPU architectures and performance characteristics.

• Experience with one or more of CUDA, Triton, CUTLASS, ROCm, NCCL, TensorRT, XLA, TVM, vLLM, TensorRT-LLM, PyTorch, or similar technologies.

• Understanding of LLM inference techniques such as batching, KV cache management, quantization, speculative decoding, parallelism strategies, and distributed serving.

• Experience working with large-scale distributed systems, high-performance computing, networking, or cluster scheduling is a plus.

• Ability to reason from first principles, use profiling data effectively, and drive measurable performance improvements.

• Strong communication skills and ability to collaborate across research, engineering, and infrastructure teams.

• Nice to Have:

• Experience optimizing transformer-based models or production LLM inference systems.

• Experience with multi-GPU or multi-node inference.

• Familiarity with compiler-level optimization, graph optimization, or model runtime internals.

• Experience with observability, benchmarking, and performance regression testing.

• Contributions to open-source AI infrastructure, GPU computing, or systems performance projects.

• Why Join Us:

• You will work on some of the most important performance challenges in modern AI infrastructure. Your work will directly improve the speed, efficiency, and scalability of LLM systems, enabling better user experiences and more cost-effective AI deployment at scale.



במקום לעבור לבד על אלפי מודעות, Jobify מנתחת את קורות החיים שלך ומציגה לך רק משרות שבאמת מתאימות לך.

מעל 80,000 משרות • 4,000 חדשות ביום
חינם. בלי פרסומות. בלי אותיות קטנות.

שאלות ותשובות עבור משרת AI Performance Engineer JB-5254

התפקיד המרכזי של מהנדס ביצועי AI ב-Recruitx הוא לייעל את הסקת המסקנות של מודלי שפה גדולים (LLM) בכל ערימת המערכת. זה כולל שיפור ביצועים, יעילות, מדרגיות ואמינות של תשתית ה-LLM, החל מאופטימיזציה של ליבות GPU ברמה נמוכה ועד לביצועי זמן ריצה של צומת בודד ואופטימיזציה של הסקת מסקנות מבוזרת ברמת האשכול. העבודה תשפר ישירות את המהירות, היעילות והמדרגיות של מערכות LLM, ותאפשר חוויות משתמש טובות יותר ופריסת AI חסכונית יותר בקנה מידה גדול.

לתפקיד מהנדס ביצועי AI המתמקד באופטימיזציית LLM, נדרשים כישורי תכנות חזקים ב-C++, CUDA, Python או שפות דומות מוכוונות מערכות. כמו כן, נדרש ניסיון בפרופיל ואופטימיזציה של תוכנה עבור זמן אחזור, תפוקה, שימוש בזיכרון או ניצול חומרה. היכרות עם ארכיטקטורות GPU ומאפייני ביצועים, וכן ניסיון עם טכנולוגיות כמו CUDA, Triton, CUTLASS, ROCm, NCCL, TensorRT, XLA, TVM, vLLM, TensorRT-LLM, PyTorch או דומות, הם חיוניים. הבנה של טכניקות הסקת מסקנות של LLM כגון אצווה, ניהול מטמון KV, קוונטיזציה ואסטרטגיות מקביליות היא יתרון משמעותי.

מהנדס ביצועי AI תורם לשיפור ביצועי הסקת מסקנות מבוזרת של LLM באשכולות על ידי אופטימיזציה של מקביליות טנסורים, מקביליות צינורות, מקביליות מומחים, דפוסי תקשורת, איזון עומסים ותזמון. התפקיד כולל גם בניית כלים, מדדי ביצועים ותהליכי פרופיל כדי לזהות צווארי בקבוק ולמדוד שיפורים בביצועים, ובכך להבטיח שהמודלים המתקדמים ביותר של AI יהיו מהירים, יעילים וקלים יותר להגשה בקנה מידה גדול.

לכל המשרות של AI Systems Performance Engineer

ניתן לצפות במשרות שסימנת בכל שלב תחת התפריט הראשי בקטגוריית 'משרות שאהבתי'

המקום קרן עזריאלי טקסט בעברית עם סמל אינסוף
  • מי אנחנו
  • מעסיקים מובילים
  • צרו קשר
  • תנאי שימוש
  • מדיניות פרטיות
  • הצהרת נגישות

2026 Ⓒ ג'וביפיי - כל הזכויות שמורות

קרן עזריאלי טקסט בעברית עם סמל אינסוף social_security the_israeli_employment_service israel_innovation_authority work_office המקום
המערכת בונה את הפרופיל התעסוקתי שלך

עוד רגע...

המערכת זיהתה ששינית את הנתונים באזור האישי ומעדכנת את ההמלצות על תפקידים ומשרות בהתאם.

מצטערים, לא הצלחנו לנתח בהצלחה את הנתונים שהזנת.
אתם מוזמנים לנסות להזין שוב או להעלות קובץ קורות חיים במידה ויש לכם.
בהצלחה

הגעת להגבלה היומית של שלושה עדכונים בפרופיל האישי ביום

loader

הבקשה שלך נשלחה בהצלחה!

יש באפשרותך לשלוח בקשה לקבלת ייעוץ אישי ללא עלות מיועצת קריירה.

באפשרותך לשלוח בקשה לקבלת ייעוץ אישי ללא עלות

  • בעיה טכנית

  • סיוע בכתיבת קורות חיים או בהכנה לראיון עבודה

  • התאמה של משרות

  • אחר:

פנייתך נשלחה בהצלחה. נציג מטעם ארגון נכי צהל ייצור איתך קשר בהקדם