עדיין מחפשים עבודה במנועי חיפוש? הגיע הזמן להשתדרג!
במקום לעבור לבד על אלפי מודעות, Jobify מנתחת את קורות החיים שלך ומציגה לך רק משרות שבאמת מתאימות לך.
מעל 80,000 משרות • 4,000 חדשות ביום
חינם. בלי פרסומות. בלי אותיות קטנות.
NeuReality is seeking a Lead System Architect to join our system architecture team and help define NR-NEXUS, our next-generation AI inference platform.
Responsibilitie
- sLead the software architecture and technical roadmap for NeuReality’s NR-Nexu
- sWrite system specifications for NR-Nexus produc
- tResearch AI infrastructure, SaaS platforms, model serving, and inference trend
- sWork with engineering to translate technical capabilities into product valu
- eWork closely with engineering teams to optimize performance, scalability, and feature delivery
- .Define performance goals and lead profiling, benchmarking, and optimization efforts for GenAI and distributed AI workloads
- .Collaborate with customers, partners, and open-source communities to ensure ecosystem compatibility and adoption
- .Mentor software engineers and provide technical leadershi
p
Requirement
- s:7+ years of software engineering experience, including 3+ years in software architecture or technical leadershi
- p.Strong experience with Kubernetes-based platforms and cloud-native architectur
- e.Deep understanding of Gen AI/LLM infrastructure and distributed workloa
- dsExperience designing management software or SaaS platforms for production system
- s.Strong background in distributed systems, microservices, APIs, and automatio
- n.Hands-on experience with observability stacks, monitoring, logging, alerting, and SLA/SLO trackin
- g.Experience with CI/CD, deployment automation, upgrades, and rollback mechanism
- s.Good understanding of security, authentication, authorization, and integration with customer data center environment
s.
Nice to h
- aveDeep understanding of GenAI / LLM inference infrastructure, including model serving, scaling, batching, latency, throughput, and resource utilizati
- on.Experience with production AI inference clusters using GPUs, AI accelerators, or other specialized compute infrastructu
- re.Understanding of how distributed inference systems operate, including scheduling, load balancing, autoscaling, failover, and cluster-level observabili
- ty.Experience with LLM serving frameworks such as vLLM, Triton Inference Server, TensorRT-LLM, or simil
- ar.Familiarity with GPU/accelerator orchestration, device plugins, resource scheduling, and cluster capacity planni
- ng.Familiarity with GPU communication technologies such as GPUDirect RDMA, NCCL, NVLink, or UALi
- nk.Experience optimizing communication for distributed AI/ML workloa
- ds.Knowledge of Prometheus, Grafana, OpenTelemetry, Helm, Argo CD, Istio, KServe, Kubeflow, or similar too
- ls.Experience deploying software in on-prem, edge, private cloud, or hybrid environmen
במקום לעבור לבד על אלפי מודעות, Jobify מנתחת את קורות החיים שלך ומציגה לך רק משרות שבאמת מתאימות לך.
מעל 80,000 משרות • 4,000 חדשות ביום
חינם. בלי פרסומות. בלי אותיות קטנות.