עדיין מחפשים עבודה במנועי חיפוש? הגיע הזמן להשתדרג!
במקום לעבור לבד על אלפי מודעות, Jobify מנתחת את קורות החיים שלך ומציגה לך רק משרות שבאמת מתאימות לך.
מעל 80,000 משרות • 4,000 חדשות ביום
חינם. בלי פרסומות. בלי אותיות קטנות.
The role
We build the inference layer that powers F5's AI security products: the systems that run large language models fast, cheaply, and reliably at production scale so we can inspect, secure, and govern enterprise giveGenAI traffic in real time.
You'll own how models are served — squeezing maximum throughput out of every GPU, cutting tail latency, and keeping the serving stack observable and self-scaling under real customer load. If you think in tokens-per-second, prefill latency, and GPU memory budgets, this is your seat.
What you'll do
• Optimise LLM inference pipelines (LLaMA, GLM, GPT-OSS and similar model classes) for throughput and latency — multi-GPU inference, prefix caching, and memory-efficient serving — targeting order-of-magnitude gains in requests-per-second (RPS).
• Own end-to-end model serving in production: deployment, low-latency inference, multi gpu parallelism, and high-throughput serving across cloud and on-prem environments.
• Build and maintain OpenAI-compatible serving APIs (e.g. /v1/chat, /v1/responses) that support reliable tool calling across multi-step agentic workflows.
• Tune and operate a modern serving stack (vLLM or equivalent) — continuous batching, KV-cache management — balancing throughput, latency, generation quality, and memory footprint.
• Maximise GPU across chip architectures like Ada Lovelace, Hopper, Blackwell, etc; profile, benchmark, and eliminate bottlenecks.
• Instrument the serving layer with logging, telemetry, and metrics (prefill latency, tokens-per-batch capacity, preemption count, queue depth) to drive observability and metric-based autoscaling.
• Ship on Kubernetes: containerised deployments (Docker, Helm), CI/CD, and low-risk, version-controlled rollouts across staging and production.
• Benchmark rigorously against recognised standards and build tooling that automates performance characterisation.
What you'll bring (must-have)
• Hands-on production experience serving LLMs at scale, with measurable throughput/latency wins you can walk through end to end.
• Deep familiarity with a modern inference/serving framework (vLLM, TensorRT-LLM, TGI, or similar), including batching and speculative decoding.
• Strong grip on GPU performance: memory management, model/tensor parallelism, and hardware-aware optimisation.
• Solid software engineering in Python OR GoLang, plus Docker + Kubernetes for production deployment.
• A benchmarking mindset — you measure rather than guess, and you can defend the trade-offs.
Nice to have
• Building OpenAI-compatible endpoints and agentic / tool-calling serving paths.
• Distributed training exposure (large-model fine-tuning / pretraining on multi-node GPU clusters).
• MLPerf submission experience or embedded / ARM inference optimisation.
• Release-management / CI-CD ownership across production and staging.
• Interest in AI security — adversarial robustness, model scanning, or securing GenAI in production.
Tech you'll work with
vLLM · NVIDIA H100 / H200 / B200 · · Kubernetes · Docker · Helm · CI/CD · GenAI-Perf
במקום לעבור לבד על אלפי מודעות, Jobify מנתחת את קורות החיים שלך ומציגה לך רק משרות שבאמת מתאימות לך.
מעל 80,000 משרות • 4,000 חדשות ביום
חינם. בלי פרסומות. בלי אותיות קטנות.
ערב