QA Engineer for AI Products/Solutions
Job Description
Hi,
\n
\n
Please review the below JD:
\n
\n
QA Engineer for AI Products/Solutions
\n
Location - Bellevue WA
\n
Job Description
\n
• Design and execute test plans for AI/ML-driven features, including model outputs, prompts, and integrated application behavior
\n
• Build and maintain automated test suites covering functional, regression, integration, and API testing
\n
• Evaluate model outputs for accuracy, consistency, bias, hallucination, and edge-case failures
\n
• Develop evaluation frameworks and golden datasets/test cases to benchmark model performance over time
\n
• Test prompt engineering changes, model version upgrades, and fine-tuning outputs for regressions
\n
• Perform adversarial and red-team style testing to surface safety, security, and robustness issues
\n
• Validate data pipelines feeding into AI models (data quality, schema, drift detection)
\n
• Collaborate with data scientists/ML engineers to define acceptance criteria and quality metrics for models
\n
• Test latency, scalability, and reliability of AI services under load
\n
• Contribute to CI/CD pipelines, integrating automated and model-evaluation tests
\n
• Hands-on experience testing LLM-based products (chatbots, copilots, RAG systems, AI agents) — designing test cases for non-deterministic, generative outputs
\n
• Practical experience with AI/LLM evaluation frameworks (e.g., Ragas, DeepEval, LangSmith, Promptfoo, OpenAI Evals, TruLens) — building eval suites, scoring rubrics, and golden datasets
\n
• Working knowledge of eval metrics for generative AI: hallucination rate, faithfulness/groundedness, relevance, answer correctness, toxicity/bias scoring, BLEU/ROUGE/semantic similarity where applicable
\n
• Experience with prompt regression testing — validating prompt changes and model/version upgrades against baseline eval sets
\n
• Strong proficiency in Python for writing eval scripts, test harnesses, and data validation logic
\n
• Familiarity with SQL and data validation techniques
\n
\n
The environment is primarily Microsoft Azure-based, with Snowflake also playing a key role. Their GenAI initiatives largely leverage OpenAI and Claude models.
