עדיין מחפשים עבודה במנועי חיפוש? הגיע הזמן להשתדרג!
במקום לעבור לבד על אלפי מודעות, Jobify מנתחת את קורות החיים שלך ומציגה לך רק משרות שבאמת מתאימות לך.
מעל 80,000 משרות • 4,000 חדשות ביום
חינם. בלי פרסומות. בלי אותיות קטנות.
Senior AI Quality Engineer
About HiO
HiO is an AI business partner for small and medium business owners, helping them grow their business through every customer conversation, acting as a proactive partner that surfaces opportunities and flags what actually needs the owner's attention. We've raised a $16M seed round backed by investors behind OpenAI and Anthropic, alongside Wix.
We're a high-talent, fast-moving team, and we're just getting started.
About the role
You'll be the guardian of HiO's quality. The person who holds the line on what great looks like as we grow. As more of the business runs through HiO, the trust owners place in it comes down to the quality of every reply, insight, and decision it makes on their behalf. You'll shape how we think about that quality, raise the bar over time, and make sure HiO stays something owners are proud to put their name on.
Quality is also our speed. Teams with great evals ship faster because they can tell, reliably and quickly, whether a change actually improved the product. You'll own that ability for HiO: every model, prompt, and context change moves at the pace your measurement allows.
This is a senior individual contributor role. You lead the craft: the methodology, the standard everyone at HiO holds AI output to, and our connection to the fast-evolving world of evaluating AI products. We need someone sharp, current, and hands-on who can own this domain end to end.
What you'll own
- Define what quality means for HiO's AI, from factual accuracy to owner-voice fidelity to the judgment calls HiO makes on the owner's behalf, including when it should act on its own versus hand something back.
- Build and run the eval system, offline and online: test sets, LLM-as-judge graders validated against human review, regression suites, and production monitoring that catches drift before customers do.
- Run the human-review loop: sample production output, name failure modes precisely, and turn findings into concrete changes to prompts, context, and product.
- Own the before/after measurement on every model, prompt, and context change, in close partnership with product and engineering. Your call on whether it shipped an improvement or a regression.
- Set the quality agenda: follow how AI evaluation is evolving as a field, bring in the methods that matter, and be the internal source of truth on where HiO is strong, where it's weak, and what it takes to close the gap.
What we're looking for
Required
- 2-3 years in AI / LLM quality, evaluation, or a closely related role.
- A deep understanding of how modern AI products actually work: how context is assembled, consumed, and used for reasoning, and how that shows up in the output. You evaluate a response by understanding why the model produced it, not just by checking it against a spec.
- You know how to read AI output. "Matches the spec, tone is fine, says the right things" is the surface. You go deeper: what the model used and what it ignored, where the reasoning bent, why a technically correct reply still fails, and what a great one would have done instead.
- Hands-on eval experience with agentic and multi-step AI systems, where quality lives in the full interaction rather than a single response. You can walk us through an eval you designed: the metric, the test set and why it looked the way it did, and a failure mode it caught in production. We'll ask.
- Fluency with modern eval practice: LLM-as-judge, human review and labeling, regression testing, and production monitoring, with a clear sense of when each one applies.
- Fluent AI operator, not just an evaluator. Strong prompting craft and a natural habit of tight build-test-refine loops, using AI daily to move faster.
- Solid testing fundamentals. Comfortable with traditional functional QA across web and mobile, so quality holds up at the product level, not just the AI layer.
- Genuinely detail-obsessed, with sharp written judgment about voice and tone. You notice the off-key word, the reply that's technically correct but lands badly, and you can tell the difference between something a real business owner would write and something an AI would, and explain the gap.
- Excellent English, written and spoken. This job is judging language at a high level, so yours has to be at one.
Preferred
- Experience with the SMB ecosystem. You understand how these owners talk to their customers.
במקום לעבור לבד על אלפי מודעות, Jobify מנתחת את קורות החיים שלך ומציגה לך רק משרות שבאמת מתאימות לך.
מעל 80,000 משרות • 4,000 חדשות ביום
חינם. בלי פרסומות. בלי אותיות קטנות.