Supervisr.ai logo

AI QA Specialist

Remote first, optional Montreal office days · Full time · Engineering

AI agents are now talking to real customers over the phone, quoting prices, handling complaints, answering questions about coverage. Most companies have no idea what those agents are actually saying, and even fewer have anyone whose job is to prove the agents behave.

Supervisr.ai is the supervision layer that sits on top of those agents. We listen to every conversation, flag critical compliance issues in real time so they can be stopped mid-call, and review everything post-call to catch drift, off-script behavior and gaps in policy before a customer or regulator does. Voice is where we started; chat is next.

You'll be the person who decides whether an AI agent is ready to talk to the public.


Quality for AI is a different job

Testing a conventional app means checking that the same input gives the same output. Testing an AI agent means checking that a thousand different conversations all stay inside the rules, and that they still do after the model, the prompt or a dependency changes underneath you. Outcomes matter more than screens.

Supervisr is managed in plain English: a rule, an expected behavior or a test scenario is described in a sentence and the platform turns it into structured configuration and test cases. You will use that same machinery to generate, run and judge scenarios at scale, and you will help decide what the Compliance Engine has to support for every client.

What you'll do

You'll own quality validation across the Supervisr.ai platform, with a strong focus on compliance outcomes and real-world AI behavior. That means:

  • Design, run and maintain test strategies covering functional, integration, regression and outcome-based scenarios, and work with engineering on unit coverage so the platform stays stable as code evolves
  • Validate AI agent configurations across scripts, intents, escalation paths and compliance logic so behavior is predictable and auditable
  • Contribute to the definition of compliance rules, tags and scenarios from real client use cases, shaping what the Compliance Engine must support for every customer
  • Validate platform behavior when new LLM versions or core dependencies ship: spot changes in outputs, reasoning and edge-case handling, and define mitigations with engineering
  • Reproduce real-world scenarios seen in production, including through ongoing review of alerts and supervision outputs, to surface defects, edge cases and failure modes, and drive root-cause analysis with engineering
  • Validate alerting, supervision logic and audit outputs against regulatory and audit expectations
  • Own quality assurance for outbound campaigns: phone number provisioning workflows, number reputation tracking and validation of automation so campaigns stay deliverable and compliant
  • Partner with internal teams to turn customer workflows, risk profiles and compliance requirements into test scenarios, validation strategies and platform requirements
  • Take part in pilots and controlled production evaluations to validate AI behavior against regulatory requirements, documented outcomes and quality benchmarks
  • Bring a quality and risk lens to sprint planning, reviews and retrospectives, with clear, actionable feedback to product and engineering

You'll love this role if

  • You're fascinated by how LLMs fail, not just how they succeed
  • You think "it worked in the demo" is where testing starts, not where it ends
  • You like turning a vague regulatory expectation into a concrete, repeatable scenario
  • You want AI in regulated industries to actually be held accountable

Our stack

Python, Java and TypeScript with an Angular UI, built cloud native on Google Cloud. Multiple LLMs through Vertex AI, MCP servers connecting the pieces, and a compliance engine that turns plain-language rules into enforceable checks.

What we're looking for

Must have

  • Experience in software QA, including manual and automated testing
  • Strong understanding of modern web applications, APIs and distributed systems
  • Experience working closely with engineering teams in an agile environment
  • Strong analytical mindset and attention to edge cases and failure modes
  • Ability to clearly document issues and communicate their impact

Nice to have

  • Experience testing AI-driven systems, LLM-based applications or conversational agents
  • Exposure to compliance-driven or regulated environments such as finance, insurance or healthcare
  • Experience validating behavior across model upgrades or third-party dependency changes
  • Familiarity with CI pipelines and test automation frameworks
  • Voice, telephony or contact center experience

Why Supervisr.ai?

  • Work on the front line of AI agent supervision, a problem that is only getting bigger
  • Small team. Big impact. Lots of ownership
  • Flexible work and remote culture
  • Competitive compensation

Interested in this role?

Send us your application.

Apply