jobify_logo ×
  • מִשׁתַמֵשׁ
  • התחברות/הרשמה
  • עמוד הבית
  • מי אנחנו
  • מעסיקים מובילים
  • פרסום משרה חינם
  • צרו קשר
  • תנאי שימוש
  • מדיניות פרטיות
  • הצהרת נגישות
קרן עזריאלי טקסט בעברית עם סמל אינסוף social_security the_israeli_employment_service work_office המקום
jobify_logo
  • מי אנחנו
  • מעסיקים מובילים
  • פרסום משרה חינם
  • צרו קשר
דילוג לתוכן

עדיין מחפשים עבודה במנועי חיפוש? הגיע הזמן להשתדרג!

במקום לעבור לבד על אלפי מודעות, Jobify מנתחת את קורות החיים שלך ומציגה לך רק משרות שבאמת מתאימות לך.

מעל 80,000 משרות • 4,000 חדשות ביום
חינם. בלי פרסומות. בלי אותיות קטנות.

AI Evaluation & Reliability Engineer

abra R&D

abra R&D abra R&D

  • תל אביב - יפו

AI Evaluation & Reliability Engineer

abra R&D

abra R&D abra R&D

  • תל אביב - יפו
  • bag_icon מלאה
  • coins_icon 25,000-35,000 ₪ הערכה מבוססת AI ולא שכר שהתקבל מהמעסיק
    הערכה מבוססת AI ולא שכר של המעסיק

Description:

abra R&D is looking for a Reliability Engineer!

abra R&D is looking for a Reliability Engineer who will take part in building the next-generation agentic analytics platform, the first real-time database optimized for AI agents at scale.

We’re looking for a Senior AI Evaluation & Reliability Engineer to define and build how AI agents are measured, validated, monitored, and improved in production. This role sits at the intersection of LLM systems, evaluation research, and production-grade engineering.

You will design evaluation methodologies, build LLM-as-a-judge systems, and develop agent-based testing frameworks to ensure correctness, robustness, and reliability of complex multi-agent workflows operating on real-time data.

What You’ll Do:

  • Design and implement evaluation frameworks for AI agents and multi-agent systems
  • Build LLM-as-a-judge pipelines to assess correctness, reasoning quality, and output quality
  • Develop agent-based evaluation systems (agents evaluating agents) for scalable testing
  • Define metrics, benchmarks, scorecards, and methodologies for agent reliability and performance
  • Build data-driven evaluation pipelines using synthetic and real-world datasets
  • Identify and analyze failure modes, edge cases, and non-deterministic behaviors
  • Improve agent robustness, consistency, and reliability in production environments
  • Work with tools such as Google ADK, Opik, and related evaluation frameworks
  • Collaborate closely with AI, platform, and database teams to shape agent–data interaction quality


Requirements:

Must have:

  • 4–8+ years of experience in software engineering, AI systems, or evaluation/QA engineering
  • Strong programming skills in Python
  • Hands-on experience working with LLMs in production environments
  • Experience building evaluation systems, automation frameworks, or testing infrastructure
  • Strong understanding of prompt engineering, tool use, and agent behavior
  • Ability to think in terms of metrics, correctness, and system reliability

Nice to have:

  • Experience with LLM evaluation frameworks (Opik, LangSmith, etc.)
  • Experience with Google ADK / agent frameworks
  • Experience implementing LLM-as-a-judge or ranking systems
  • Background in data systems, analytics, or real-time pipelines
  • Experience with multi-agent systems
  • Familiarity with statistical evaluation methods or experimentation (A/B testing, scoring systems)



במקום לעבור לבד על אלפי מודעות, Jobify מנתחת את קורות החיים שלך ומציגה לך רק משרות שבאמת מתאימות לך.

מעל 80,000 משרות • 4,000 חדשות ביום
חינם. בלי פרסומות. בלי אותיות קטנות.

שאלות ותשובות עבור משרת AI Evaluation & Reliability Engineer

מהנדס/ת AI Evaluation & Reliability ב-abra R&D אחראי/ת על הגדרה ובנייה של אופן המדידה, אימות, ניטור ושיפור של סוכני AI בסביבת ייצור. התפקיד כולל תכנון מתודולוגיות הערכה, בניית מערכות LLM-as-a-judge ופיתוח מסגרות בדיקה מבוססות סוכנים כדי להבטיח נכונות, חוסן ואמינות של תהליכי עבודה מורכבים מרובי סוכנים הפועלים על נתונים בזמן אמת, כחלק מפיתוח פלטפורמת הניתוח הסוכנתית מהדור הבא.

לתפקיד מהנדס/ת AI Evaluation & Reliability בכיר/ה ב-abra R&D נדרשות 4-8 שנות ניסיון בהנדסת תוכנה, מערכות AI או הנדסת הערכה/QA, מיומנויות תכנות חזקות בפייתון, וניסיון מעשי עם LLMs בסביבות ייצור. כמו כן, נדרש ניסיון בבניית מערכות הערכה, מסגרות אוטומציה או תשתית בדיקות, והבנה חזקה של הנדסת פרומפטים, שימוש בכלים והתנהגות סוכנים. יכולת חשיבה במונחים של מדדים, נכונות ואמינות מערכת היא חיונית.

מהנדס/ת AI Evaluation & Reliability תורם/ת לשיפור אמינות סוכני AI בפלטפורמת הניתוח של abra R&D על ידי זיהוי וניתוח מצבי כשל, מקרי קצה והתנהגויות לא דטרמיניסטיות. התפקיד כולל שיפור חוסן, עקביות ואמינות של סוכנים בסביבות ייצור, עבודה עם כלים כמו Google ADK ו-Opik, ושיתוף פעולה הדוק עם צוותי AI, פלטפורמה ובסיסי נתונים כדי לעצב את איכות האינטראקציה בין סוכנים לנתונים.

לכל המשרות של AI Reliability Engineer

ניתן לצפות במשרות שסימנת בכל שלב תחת התפריט הראשי בקטגוריית 'משרות שאהבתי'

המקום קרן עזריאלי טקסט בעברית עם סמל אינסוף
  • מי אנחנו
  • מעסיקים מובילים
  • צרו קשר
  • תנאי שימוש
  • מדיניות פרטיות
  • הצהרת נגישות

2026 Ⓒ ג'וביפיי - כל הזכויות שמורות

קרן עזריאלי טקסט בעברית עם סמל אינסוף social_security the_israeli_employment_service israel_innovation_authority work_office המקום
המערכת בונה את הפרופיל התעסוקתי שלך

עוד רגע...

המערכת זיהתה ששינית את הנתונים באזור האישי ומעדכנת את ההמלצות על תפקידים ומשרות בהתאם.

מצטערים, לא הצלחנו לנתח בהצלחה את הנתונים שהזנת.
אתם מוזמנים לנסות להזין שוב או להעלות קובץ קורות חיים במידה ויש לכם.
בהצלחה

הגעת להגבלה היומית של שלושה עדכונים בפרופיל האישי ביום

loader

הבקשה שלך נשלחה בהצלחה!

יש באפשרותך לשלוח בקשה לקבלת ייעוץ אישי ללא עלות מיועצת קריירה.

באפשרותך לשלוח בקשה לקבלת ייעוץ אישי ללא עלות

  • בעיה טכנית

  • סיוע בכתיבת קורות חיים או בהכנה לראיון עבודה

  • התאמה של משרות

  • אחר:

פנייתך נשלחה בהצלחה. נציג מטעם ארגון נכי צהל ייצור איתך קשר בהקדם