עדיין מחפשים עבודה במנועי חיפוש? הגיע הזמן להשתדרג!
במקום לעבור לבד על אלפי מודעות, Jobify מנתחת את קורות החיים שלך ומציגה לך רק משרות שבאמת מתאימות לך.
מעל 80,000 משרות • 4,000 חדשות ביום
חינם. בלי פרסומות. בלי אותיות קטנות.
The team is new and small, which means broad scope, direct ownership, and real influence over the technical direction we take. If you enjoy digging into performance bottlenecks and turning analysis into measurable wins, this role is for you.
Key job responsibilities
Design and implement high-performance compute kernels for ML operations, leveraging the Neuron architecture and programming models.
Profile ML workloads end-to-end to identify bottlenecks - memory, compute, or communication - and drive optimizations through to a measured improvement.
Enhance the programming model and tooling that kernel and model developers rely on, improving usability and debugging workflows.
Identify and drive optimization opportunities across the Neuron software stack (compiler, runtime, frameworks).
Document software designs, operational runbooks, and performance findings so the broader team can build on your work.
A day in the life
You might start your morning reviewing profiling data from a customer's large diffusion model training job, tracing a utilization gap back to a specific kernel. After a design discussion with compiler engineers about a new operator fusion strategy, you spend the afternoon writing and benchmarking a kernel prototype. Later, you review a teammate's pull request for a runtime optimization and share your findings in a short write-up for the broader Neuron organization. Your work directly translates into faster model execution and lower cost for AWS customers running ML workloads at scale.
Basic Qualifications
- 3+ years of non-internship professional software development experience.
- Knowledge of Python and/or C++ programming.
- Knowledge of computer architecture, operating systems, and parallel computing.
- Experience with PyTorch, TensorFlow, and/or JAX.
Preferred Qualifications
- Master's degree in Computer Science, Engineering, Mathematics, or a related field.
- Experience optimizing performance for LLM, Vision, or other deep-learning models.
- Experience with kernel writing or parallel programming (CUDA, Triton, CUTLASS, Pallas, Mojo, SIMD, MPI).
- Experience with compiler optimization or hardware-software co-design.
במקום לעבור לבד על אלפי מודעות, Jobify מנתחת את קורות החיים שלך ומציגה לך רק משרות שבאמת מתאימות לך.
מעל 80,000 משרות • 4,000 חדשות ביום
חינם. בלי פרסומות. בלי אותיות קטנות.
ערב