GenAI Testing: LLM, RAG, Agentic AI, API and Test Automation
What are the top 3 skills required for this role?
1. GenAI, LLM, RAG and agentic AI testing
2. Python-based API and test automation, including PyTest or Robot Framework
3. LLM evaluation, prompt-injection and hallucination testing, CI/CD and observability
Additional Information:
Team size, direct reports, key deliverables, unique selling points, additional qualifications, team culture etc. Key Responsibilities
Develop and execute comprehensive test strategies for GenAI applications, including LLM-powered solutions, RAG pipelines, chatbots, copilots and agentic AI workflows.
Validate functional and non-functional requirements across prompts, model responses, APIs, user interfaces, integrations and end-to-end business workflows.
Create reusable evaluation datasets, test cases and scoring rubrics for relevance, groundedness, factual accuracy, completeness, coherence, instruction adherence and response consistency.
Test RAG solutions by evaluating retrieval quality, context precision and recall, source attribution, answer faithfulness and resistance to unsupported or hallucinated responses.
Validate agentic AI behaviour, including tool selection, function calling, reasoning flows, multi-step task completion, memory handling, fallback paths and failure recovery.
Perform adversarial, security and safety testing for prompt injection, jailbreaks, sensitive-data exposure, toxic or prohibited content, authorization bypass and insecure tool use.
Conduct performance, scalability and reliability testing for response latency, throughput, token usage, cost, concurrency, rate limits and degradation under load.
Build automated regression suites in Python and integrate GenAI evaluations, API tests and quality gates into CI/CD pipelines.
Monitor production quality through traces, evaluation metrics, user feedback and drift indicators; identify defects, analyse root causes and support remediation.
Collaborate with product, development, data science, security and business teams to define acceptance criteria, test coverage and release-readiness standards.
Required Qualifications and Skills:
Strong experience in software quality engineering, test strategy, test design, defect management and automation for enterprise applications.
Hands-on knowledge of GenAI concepts, LLM behaviour, prompt engineering, embeddings, vector databases, RAG architectures and agentic AI systems.
Proficiency in Python and experience with automation frameworks such as PyTest or Robot Framework, together with API testing tools and libraries.
Experience evaluating LLM outputs using deterministic checks, semantic similarity, model-based evaluation and human-review workflows.
Practical experience testing hallucinations, bias, robustness, non-deterministic behaviour, prompt injection, guardrails and content-safety controls.
Working knowledge of GenAI evaluation and observability tools such as Ragas, DeepEval, Promptfoo, LangSmith, Arize Phoenix or equivalent platforms.
Strong SQL skills and experience validating source data, document ingestion, chunking, metadata, embeddings and retrieval results.
Experience with GitHub, GitHub Actions or comparable CI/CD tools; familiarity with Docker, Kubernetes, AWS or Azure is desirable.
Ability to define measurable acceptance thresholds, analyse evaluation results and communicate quality risks clearly to technical and business stakeholders.
Strong analytical, investigative and problem-solving skills, with attention to responsible AI, privacy, security and regulatory expectations.
Preferred Qualifications:
Bachelor's or Master's degree in Computer Science, Software Engineering, Data Science or a related field, with at least five years of experience in quality engineering, test automation, data testing, AI/ML testing or a related discipline.
Experience testing applications in banking or another regulated environment, including model governance, audit evidence and responsible-AI controls, is an asset.