Design and execute end-to-end testing strategies specifically tailored for Machine Learning models, Generative AI systems, RAG architectures, and Autonomous Agents.
Validate model accuracy, fairness, bias detection, explainability, robustness, and performance across diverse and edge-case datasets.
Execute adversarial testing, prompt-injection, jailbreaking, and red-teaming to evaluate prompt robustness and behavioral variations under stress.
Validate agentic workflows, including multi-step reasoning paths, state transitions, tool execution, and fallback behaviors during service failures.
Evaluate LLM outputs for correctness, grounding, factuality, consistency, safety, and hallucination reduction.
Assess vector store behavior, document chunking logic, retriever configurations, and semantic search accuracy.
Conduct API, performance, latency, throughput, and concurrency testing on AI inference endpoints and data pipelines.
Ensure compliance with AI ethics, data privacy laws, business rules, and insurance regulatory guidelines, maintaining audit-ready test evidence and behavioral reports.
Define AI quality KPIs, establish test governance, and build automated testing frameworks integrated into CI/CD pipelines.
Collaborate closely with Data Scientists, ML Engineers, SMEs, and DevOps teams while mentoring junior QA engineers and creating reusable test accelerators.
Lead AI Quality Engineer & Test Automation Architect
Descripción del puesto / Funciones
Requisitos mínimos
Experience & Specialization: Proven senior/lead expertise in software quality engineering with a dedicated focus on AI/ML systems and GenAI applications.
Programming & Automation: Advanced proficiency in Python for test automation, data validation, and custom AI testing scripts.
GenAI & RAG Ecosystems: Hands-on experience with GenAI frameworks, vector databases, chunking strategies, and retrieval evaluation.
Model Evaluation & Metrics: Deep understanding of data validation, model evaluation metrics, fairness/bias testing, and drift detection (data and concept drift).
API Testing: Expertise in testing AI services and model endpoints using tools such as Postman, REST Assured, or Python REST clients.
DevOps, Cloud & Infrastructure:
Experience with CI/CD pipelines for continuous testing integration.
Exposure to cloud platforms hosting AI deployments.
Working knowledge of containerization and orchestration environments (e.g., Docker, Kubernetes).
Familiarity with Big Data ecosystems for large-scale AI testing.
Security & Governance: Experience in AI ethics, compliance testing, observability tools, and security testing for data pipelines and model-serving endpoints.
Requisitos valorables
Advanced Red Teaming: Hands-on experience building automated adversarial test suites and automated synthetic data generation for rare edge cases.
Framework Automation: Direct implementation of specialized LLM evaluation frameworks (e.g., Ragas, DeepEval, TruLens).
Observability Setup: Advanced configuration of AI monitoring dashboards and automated regression testing workflows for retrained models.
