Talent.com
Zafin
AI Evaluation EngineerZafin • Toronto, ON, Canada
AI Evaluation Engineer

AI Evaluation Engineer

Zafin • Toronto, ON, Canada
15 days ago
Salary
CA$80,000.00 yearly
Job type
  • Full-time
Job description

Zafin is an AI platform company helping regulated institutions modernize how critical work is designed, governed, and delivered. Our technology enables organizations to move faster while maintaining the governance, accountability, and control required in highly regulated environments.

Our portfolio includes Zafin AIOS, an agent orchestration platform for governed AI work; the Zafin Banking Platform, which helps banks modernize product, pricing, offers, billing, loyalty, and relationship management; and Zafin IO, an integration platform that connects data, systems, and workflows across complex enterprise environments.

Headquartered in Toronto, Canada, Zafin partners with leading financial institutions across North America, Europe, the Middle East, Africa, and Asia-Pacific. As AI transforms the future of financial services, we're building the platforms that help regulated organizations adopt AI responsibly and at scale.

What’s the Opportunity?

The AI Evaluation Engineer ensures AI agent solutions are accurate, reliable, safe, and production-ready within regulated banking environments. The role provides objective, evidence-based evaluation of AI agent behaviour against defined business, quality, risk, performance, and regulatory criteria, helping ensure AI solutions deliver consistent, trusted outcomes in production.

Working within the Reliability Testing phase of the AIOS (Zafin's AI Operating System) delivery lifecycle, the role designs, executes, and leads evaluation activities that validate AI agent behaviour across the full development lifecycle. This includes developing meaningful evaluation scenarios, identifying defects and failure modes, monitoring quality across releases, and providing actionable feedback that continuously improves AI agent reliability. The role helps operationalize AIOS's principle of reliability first, velocity second through disciplined evaluation and objective production-readiness decisions.

The AI Evaluation Engineer works closely with Agent Engineers, Industry Consultants, Agent Architects, AI Knowledge & Governance, Product, and Delivery teams to ensure evaluation reflects intended business logic, regulatory requirements, technical standards, and real-world operating conditions. Depending on experience and level, the role may also lead evaluation activities, coach other evaluation engineers, improve evaluation practices, and drive quality improvements across multiple AI agent initiatives.

Ultimately, the role helps ensure AI agent capabilities earn and maintain customer trust by delivering consistent, reliable, and explainable outcomes in production.

What Will You Do?

  • Design and execute structured evaluation scenarios that validate AI agent accuracy, reliability, safety, compliance, and business outcomes.
  • Validate AI agent behaviour against approved business rules, policies, technical requirements, source knowledge, and expected outcomes.
  • Conduct regression evaluation across releases and monitor behavioural drift, performance degradation, and newly introduced failure modes.
  • Identify, document, prioritize, and track quality issues, defects, and production risks.
  • Support production-readiness decisions through objective evaluation evidence and recommendations.
  • Analyze evaluation results to identify root causes, recurring quality trends, and opportunities to improve prompts, workflows, knowledge, integrations, and engineering practices.
  • Maintain reusable evaluation scenarios, benchmark datasets, expected outcomes, regression suites, and supporting evidence.
  • Contribute to continuous improvement of evaluation methodologies, automation, tooling, and engineering feedback loops.
  • Support investigation of production issues and validate corrective actions.
  • Partner with Agent Engineering, Industry Consultants, Agent Architects, AI Knowledge & Governance, Product, and Delivery teams to ensure evaluation reflects business requirements and production expectations.
  • Participate in production-readiness reviews, release planning, and AI agent optimization activities.
  • Communicate evaluation findings, quality risks, and recommendations clearly to technical and business stakeholders.
  • Depending on experience and level, you may also:
    • Lead evaluation activities across one or more AI agent initiatives or squads.
    • Review evaluation approaches, production-readiness recommendations, and quality evidence produced by other Evaluation Engineers.
    • Coach and mentor less experienced Evaluation Engineers, supporting technical growth and consistent evaluation practices.
    • Coordinate evaluation priorities across multiple concurrent initiatives and support delivery planning.
    • Drive improvements in evaluation tooling, automation, benchmark management, CI/CD quality integration, and operational effectiveness.
    • Analyze systemic quality trends and lead continuous improvement initiatives across multiple AI agent capabilities.
    • Partner with Delivery and Engineering leadership to improve quality outcomes, operational consistency, and AI agent reliability.

What Do You Need to Succeed?

Must Haves

  • Typically3–10+ years of relevant experience in software quality engineering, AI evaluation, AI quality engineering, machine learning evaluation, software testing, or related disciplines.
  • Level and scope of responsibility will be determined based on demonstrated technical capability, evaluation expertise, leadership experience, independence, and ability to influence quality outcomes.
  • Experience leading evaluation activities, mentoring technical professionals, or coordinating quality initiatives is advantageous for more senior levels.
  • Degree in Computer Science, Software Engineering, Data Science, Artificial Intelligence, or related discipline, or equivalent practical experience.
  • Strong understanding of AI evaluation, large language model behaviour, reasoning quality, hallucination detection, safety, instruction adherence, factual accuracy, and business correctness.
  • Experience with structured software testing, regression evaluation, production-readiness assessment, and quality engineering.
  • Working knowledge of SDLC, CI/CD, automated evaluation, AI observability, and engineering delivery practices.
  • Familiarity with benchmark management, evaluation tooling, quality automation, and AI engineering workflows.
  • Understanding of privacy, security, governance, and regulatory considerations relevant to enterprise AI.
  • Proficiency with Python, SQL, or similar tools supporting evaluation and analysis.

Nice to Have

  • Experience evaluating LLMs, RAG systems, AI agents, or agentic AI platforms.
  • Experience with AI evaluation platforms such as LangSmith, OpenAI Evals, or comparable tools.
  • Experience integrating automated evaluation into CI/CD or MLOps workflows.
  • Experience with model observability, behavioural-drift detection, or AI production monitoring.
  • Banking, financial services, or other regulated industry experience.
  • Experience leading technical teams, quality initiatives, or engineering improvement programs.

Additional Job Details

  • Expected Salary Range: $80,000 - $180,000; we hire into multiple career levels for this role based on a candidate's experience, skills and demonstrated capabilities.
  • Vacancy Status: Open Position to be filled
  • Mode of Work: Hybrid
  • Use of AI: Zafin may use Artificial Intelligence (AI) and/or other forms of automated technology to screen and/or assess applicants for this position. Zafin will not utilize AI for conducting interviews and/or making hiring decisions.

What’s in it for you

Joining our team means being part of a culture that values diversity, teamwork, and high-quality work. We offer competitive salaries, annual bonus potential, generous paid time off, paid volunteering days, wellness benefits, and robust opportunities for professional growth and career advancement. Want to learn more about what you can look forward to during your career with us? Visit our careers site and our openings:zafin.com/careers

Zafin welcomes and encourages applications from people with disabilities. Accommodations are available on request for candidates taking part in all aspects of the selection process.

Zafin is committed to protecting the privacy and security of the personal information collected from all applicants throughout the recruitment process. The methods by which Zafin contains uses, stores, handles, retains, or discloses applicant information can be accessed by reviewing Zafin’s privacy policy at https://zafin.com/privacy-notice/. By submitting a job application, you confirm that you agree to the processing of your personal data by Zafin described in the candidate privacy notice.

#J-18808-Ljbffr
Create a job alert for this search

AI Evaluation Engineer • Toronto, ON, Canada

Similar jobs

Principal AI Solutions Developer

OpenTextrichmond hill, york region, Canada
Full-time

Become a Principal AI Solutions Developer at OpenText, leading the charge in creating sophisticated agentic AI systems.Use your expertise to enhance IT Operations security and observability.OpenTex... Show more

 • Promoted

Generative AI Implementation Expert

SiaToronto, ON, CA
Full-time

Drive the future of AI solutions as a Generative AI Implementation Expert.Utilize your skills to design and optimize generative AI systems for widespread impact across various industries.You will p... Show more

 • Promoted

AI Platform Engineer: Build & Evaluate AI Systems

Armilla AItoronto, on, Canada
Full-time

A leading AI-focused startup in Toronto is seeking an experienced AI Engineer to shape the future of AI risk management.In this pivotal role, you will be responsible for building core AI tools and ... Show more

 • Promoted

AI Evaluation Engineer

ZafinToronto, ON, CA
Full-time

Zafin is an AI platform company helping regulated institutions modernize how critical work is designed, governed, and delivered.Our technology enables organizations to move faster while maintaining... Show more

 • Promoted

Enterprise AI Engineer

Porter Airlines Inc.toronto, on, Canada
Full-time

Reporting to the Managing Director, Artificial Intelligence, the Enterprise AI Engineer will play a foundational role in establishing Porter’s enterprise AI capability.As one of the first dedicated... Show more

 • Promoted

Senior, Applied AI Engineer

Headstart AItoronto, on, Canada
Full-time

This range is provided by Headstart AI.Your actual pay will be based on your skills and experience — talk with your recruiter to learn more.At Headstart our mission is to bring companies into the a... Show more

 • Promoted

ML Engineer (Speech-to-Speech) — Subject Matter Expert

Vosyntoronto, on, Canada
Full-time

ML Engineer (Speech-to-Speech) — Subject Matter Expert.At Vosyn, we embrace the exciting, game-changing world of Artificial Intelligence, driving innovation and pioneering impactful projects across... Show more

 • Promoted

Intern Researcher - AI Agent Evaluation

Huawei CanadaMarkham, York Region, CA
Full-time

Huawei Canada has an immediate 4-12 month internship opening for an Intern Researcher.Established in 2014, the Distributed Scheduling and Data Engine Lab is Huawei Cloud's technical innovation cent... Show more

 • Promoted

DevOps Engineer - AI Model Evaluator

MercorToronto, Ontario, Canada
CA$85.00 hourly
Remote
Part-time
Quick Apply

Headquartered in San Francisco, our investors include.DevOps / SRE / Cloud Engineer (Coding Agent Experience).Review model-generated implementations involving.Identify bugs, edge cases, reliability... Show more

Applied AI Engineer

Nexxa.AIToronto, ON, CA
Full-time

Our mission is to translate deep technical breakthroughs into operational reality, solving some of the hardest systems-level problems in industry.We’re looking for a Applied AI Engineer to work as ... Show more

 • Promoted

Remote Ai Data Quality & Evaluation Engineer - C$85,000 - C$225,000 A Year - Remote

Cloud Technology CompanyNorth York, Canada
Remote
Full-time

AI Data Quality & Evaluation Engineer needed to validate and ensure the reliability of AI Agents using expertise in data quality, prompt engineering, and Python. Show more

 • Promoted

Demo Engineer - AI [32469]

Stealth AI Startuptoronto, on, Canada
Full-time

Get AI-powered advice on this job and more exclusive features.This range is provided by Stealth AI Startup.Your actual pay will be based on your skills and experience — talk with your recruiter to ... Show more

 • Promoted

AI Systems Engineer – AI Model (Training & Inference)

AMDmarkham, on, Canada
Full-time

At AMD, our mission is to build great products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems.Grounded in a culture of innovatio... Show more

 • Promoted

AI Engineer – Generative AI Expertise

Synechrontoronto, on, Canada
Full-time

Join Synechron as an AI Engineer specializing in Generative AI.Leverage your advanced skills in machine learning to innovate solutions for capital markets.As a vital contributor, you will focus on ... Show more

 • Promoted

Principal Engineer for AI Inference Systems

Mediumtoronto, on, Canada
Full-time

Lead the development of advanced AI inference systems as a Principal Engineer.Design and implement robust platforms for agentic functionality across multiple data modalities.In this pivotal role, y... Show more

 • Promoted

AI Engineer

BioRenderToronto, ON, CA
Full-time

At BioRender, we don’t settle for average, and neither should you.We’re building tools that transform how scientists communicate complex ideas, and we need someone who’s ready to turn ambitious ide... Show more

 • Promoted

Remote AI Data Quality & Evaluation Engineer

Veeva Systemstoronto, on, Canada
Remote
Full-time

A leading cloud technology company in Toronto is seeking an experienced professional to validate and ensure the reliability of Veeva AI Agents.The ideal candidate should have expertise in data qual... Show more

 • Promoted

Thomson Reuters Lead Engineer in AI Innovation

PowerToFlytoronto, on, Canada
Full-time

Thrive as a Lead Research Engineer at Thomson Reuters, where you will harness AI and Deep Learning to create transformative solutions for clients.Drive innovative projects in a dynamic hybrid work ... Show more

 • Promoted

AI Engineer for Reliability Solutions

Mosaic.techtoronto, on, Canada
Full-time

Join Rootly as an AI Engineer specializing in developing reliable AI-driven solutions for incident management.Use your LLM experience to shape the future of business reliability.As a Senior AI Engi... Show more

 • Promoted

Machine Learning Engineer Focused on AI Systems

HR Tech Jobtoronto, on, Canada
Full-time

Step up as a Machine Learning Engineer at Upwork, where you'll develop cutting-edge AI and reinforcement learning systems.Drive transformative changes with advanced technologies across our platform... Show more