Wallex is a FinTech company providing a secure cryptocurrency exchange platform and digital asset payment solutions.
With a vision to become a leading global FinTech enterprise, Wallex is recognized as one of the largest cryptocurrency exchange platforms in Iran, supporting the online trading of over 130 cryptocurrencies.
We are looking for a Senior AI QA Engineer to own how we measure, evaluate, and defend the quality and safety of our generative AI systems.
This is an engineering-focused quality role rather than traditional manual QA. You will build the evaluation infrastructure that determines whether AI-generated content is ready to ship, including automated evaluation systems, quality gates, regression suites, adversarial testing, defect registries, and measurement instrumentation.
You will work closely with AI, Creative, Product, and Safety teams to translate quality standards into reliable systems that run automatically, produce measurable and trustworthy results, and prevent unacceptable content from reaching users. As a senior member of a small team, you will have significant ownership over quality methodology, tooling, and release decisions.
Responsibilities:
- Own the evaluation infrastructure end to end, including evaluation harnesses, scoring pipelines, datasets, dashboards, and internal tooling.
- Build and maintain automated evaluators that assess generated content across quality, age-appropriateness, safety, consistency, and alignment with intended outcomes.
- Translate quality rubrics defined by Creative, AI, and domain experts into executable, reproducible, and version-controlled evaluation systems.
- Design and operate multi-model evaluation systems and judge panels for assessing non-deterministic AI outputs.
- Monitor evaluator consistency and investigate significant changes in evaluation results.
- Manage evaluator selection, calibration, and ongoing validation.
- Build calibration and exemplar datasets that provide consistent reference points for evaluation and reduce reliance on subjective internal judgment.
- Implement deterministic quality and safety checks that identify structural defects before probabilistic evaluation.
- Build and maintain a defect registry that turns newly discovered failure modes into permanent regression checks.
- Integrate quality and safety gates directly into CI/CD and release workflows.
- Define clear routing for failed outputs, including rejection, regeneration, escalation, and human review.
- Maintain rigorous evaluation methodology, including controlled comparisons, consistent test conditions, and complete regression runs where appropriate.
- Monitor release-level quality signals and provide clear go/no-go recommendations based on evidence.
- Escalate and hold releases when quality or safety thresholds are not met.
- Build and maintain automated adversarial and edge-case test suites for model, prompt, and pipeline changes.
- Work with Safety and AI teams to translate safety policies into automated, measurable, and auditable enforcement mechanisms.
- Design coverage for edge cases across different ages, languages, names, family structures, cultural contexts, and emotionally sensitive subjects.
- Test model and pipeline behavior under unexpected, adversarial, and ambiguous inputs.
- Incorporate findings from expert and human reviewers back into automated evaluation and regression systems.
- Ensure recurring classes of failures become detectable automatically rather than being handled only as individual incidents.
- Turn quality and safety into measurable, trendable data through baselines, dashboards, regression alerts, and release reporting.
- Track evaluation performance and identify meaningful changes across models, prompts, datasets, and pipeline versions.
- Instrument generation cost and end-to-end latency as continuously monitored metrics.
- Reconcile offline evaluation results with live product behavior to ensure evaluation reflects the actual user experience.
- Design experiments and analyses that distinguish meaningful improvements from measurement noise.
- Identify sources of bias and instability in evaluation methodology and improve measurement reliability over time.
- Work closely with AI and Creative teams to diagnose failure modes and determine whether issues should be addressed in generation, prompting, system design, or evaluation.
- Provide structured and machine-readable failure reasons so quality issues can be analyzed and addressed systematically.
- Partner with Product and Engineering teams to ensure quality requirements are incorporated early in development rather than only before release.
- Contribute to quality standards, documentation, evaluation methodology, and engineering practices as the platform expands.
- Communicate quality findings and release risks clearly to both technical and non-technical stakeholders.
Requirements:
- 5+ years of professional experience in QA Engineering, Software Engineering, ML/AI testing, or a related engineering discipline, with significant experience evaluating ML- or LLM-based systems.
- Strong proficiency in Python and the ability to build production-grade evaluation harnesses, pipelines, automation, and internal tooling.
- Strong understanding of software testing principles and experience applying them to non-deterministic systems where traditional exact-match assertions are insufficient.
- Strong statistical literacy, including the ability to distinguish signal from noise, understand inter-rater agreement, design appropriate samples, and identify sources of bias in measurement.
- Hands-on experience with LLM-as-a-Judge or similar model-based evaluation approaches, including rubric design, concrete evaluation criteria, and evaluator calibration.
- Experience building evaluation datasets, regression suites, and automated testing systems for generative AI or other probabilistic systems.
- Strong CI/CD knowledge and experience integrating automated quality gates directly into release pipelines.
- Strong analytical and problem-solving skills, with a skeptical and evidence-based approach to quality measurement.
- Ability to investigate whether an observed improvement represents a genuine system improvement or a change in the measurement process.
- Strong ownership mindset and the ability to defend quality decisions with evidence, including under schedule and delivery pressure.
- Excellent written and verbal communication skills.
- Fluency in English and Persian.
Nice to Have:
- Experience with red-teaming or adversarial testing of generative AI systems.
- Familiarity with content safety layers and multi-layer safety architectures.
- Experience with LLM observability, evaluation, and monitoring tools.
- Experience tracking per-request generation cost, latency, and quality metrics.
- Background in research methodology, cognitive science, educational measurement, psychometrics, linguistics, or a related field.
- Experience evaluating creative, narrative, or multimodal AI outputs rather than only factual accuracy or classification performance.
- Experience working with audio, image, or multimodal generative AI systems.
- Familiarity with children's safety and privacy requirements such as COPPA and GDPR-K.
- Experience building quality systems for consumer-facing AI products.
- Strong literary or creative sensibility and the ability to evaluate the quality, coherence, and appropriateness of generated narratives.
We build, not watch — This is where you belong!