עדיין מחפשים עבודה במנועי חיפוש? הגיע הזמן להשתדרג!
במקום לעבור לבד על אלפי מודעות, Jobify מנתחת את קורות החיים שלך ומציגה לך רק משרות שבאמת מתאימות לך.
מעל 80,000 משרות • 4,000 חדשות ביום
חינם. בלי פרסומות. בלי אותיות קטנות.
GSI Technology is a leading international company publicly traded on NASDAQ, specializing in the development of the Gemini® processor — a cutting-edge Associative Processing Unit (APU) designed for computer-in-memory acceleration.
We are seeking a highly skilled and self-driven Senior Software Engineer to lead the development and optimization of large language model (LLM) implementations on a proprietary Associative Processing Unit (APU). This role combines high-level machine learning understanding with low-level system and performance engineering, primarily using C and C++. The ideal candidate will possess deep knowledge of transformer architecture, be capable of writing reference implementations, debugging Python pipelines, and optimizing memory and performance across the full stack.
Key Responsibilities:
- AI Library Development:
- Build and optimize software libraries for Large Language Models (LLMs) and Large Vision Models (LVMs).
- System Architecture:
- Design end-to-end system flows integrating AI models, particularly focusing on computer vision applications.
- Hardware Optimization:
- Improve performance under hardware constraints (e.g., latency, memory, compute limitations).
- Cross-Stack Integration:
- Ensure seamless interaction between AI models, software infrastructure, and underlying hardware.
Required Skills:
At least 5 years of experience in:
✔️ LLM/LVM Development: Practical experience with transformer architectures and vision-language models
✔️ Computer Vision: Strong background in CV pipelines and multimodal systems
✔️ System Design: Proven ability to architect complex software systems from concept to deployment.
Technical Specializations
✔️ Hardware-aware optimization techniques (quantization, pruning, kernel fusion)
✔️ Performance profiling tools (e.g., PyTorch Profiler, NVIDIA Nsight)
✔️ Low-level optimization (CUDA, OpenCL, or hardware-specific SDKs)
Preferred Qualifications
- Experience with AI accelerator architectures (GPUs, TPUs, NPUs)
- Familiarity with model compression techniques and distributed training
- Background in ML compiler frameworks (TVM, MLIR, or XLA)
Our Privacy Policy: Your resume and information will be kept confidential.
במקום לעבור לבד על אלפי מודעות, Jobify מנתחת את קורות החיים שלך ומציגה לך רק משרות שבאמת מתאימות לך.
מעל 80,000 משרות • 4,000 חדשות ביום
חינם. בלי פרסומות. בלי אותיות קטנות.