עדיין מחפשים עבודה במנועי חיפוש? הגיע הזמן להשתדרג!
במקום לעבור לבד על אלפי מודעות, Jobify מנתחת את קורות החיים שלך ומציגה לך רק משרות שבאמת מתאימות לך.
מעל 80,000 משרות • 4,000 חדשות ביום
חינם. בלי פרסומות. בלי אותיות קטנות.
Scandinavian Capital Markets is expanding its technology capabilities to build a next-generation AI automation platform. We are seeking a Senior AI / LLM Engineer to spearhead our model infrastructure, hosting, and system architecture centered around open-source Large Language Models.
In this role, you will be responsible for deploying, serving, and fine-tuning open-source foundation models, as well as building the core orchestration layer that powers complex, real-time user automations and workflow engines. You will bridge the gap between cutting-edge AI model hosting and reliable enterprise software architecture.
Responsibilities
- Deploy, host, and optimize open-source foundation models (Qwen, Llama, Mistral) for low-latency, high-concurrency production inference.
- Design and build the agentic orchestration layer, function-calling pipelines, and structured workflow engines that enable users to build custom automated logic.
- Implement pipelines for model fine-tuning (LoRA/QLoRA), quantization (AWQ, GGUF), and domain adaptation.
- Build robust, event-driven backends and APIs to seamlessly connect LLM inference engines with real-time data feeds and user inputs.
- Manage GPU cluster orchestration, inference throughput, and cost-efficient scaling across cloud providers.
Key Requirements
- 3+ years of hands-on experience in software engineering with a strong focus on production AI/ML systems and Large Language Models.
- Demonstrated experience hosting and serving open-source models using inference frameworks such as vLLM, TensorRT-LLM, Hugging Face TGI, or Ollama.
- Experience building function-calling pipelines, custom tool integration, and structured output systems (e.g., using Pydantic, LangGraph, or custom frameworks).
- Hands-on experience with cloud GPU orchestration (AWS, RunPod, Modal, Lambda Labs), Docker, and Kubernetes.
Preferred Qualifications
- Experience working with high-throughput, low-latency data pipelines or real-time streaming architectures.
- Background in building workflow engines, rule-based automation tools, or developer-facing platforms.
- Familiarity with evaluate/eval frameworks for monitoring LLM accuracy and performance in production.
What We Offer
- The opportunity to lead the AI architecture for a brand-new, high-impact initiative from the ground up.
- Collaborative culture backed by an established financial institution.
- Competitive salary, modern hardware setup, and flexible working arrangements.
במקום לעבור לבד על אלפי מודעות, Jobify מנתחת את קורות החיים שלך ומציגה לך רק משרות שבאמת מתאימות לך.
מעל 80,000 משרות • 4,000 חדשות ביום
חינם. בלי פרסומות. בלי אותיות קטנות.
ערב