Senior AI Engineer, AI Agent Supervision
Remote first, optional Montreal office days · Full time · Engineering
AI agents are now talking to real customers over the phone, quoting prices, handling complaints, answering questions about coverage. Most companies have no idea what those agents are actually saying.
Supervisr.ai is the supervision layer that sits on top of those agents. We listen to every conversation, flag critical compliance issues in real time so they can be stopped mid-call, and review everything post-call to catch drift, off-script behavior and gaps in policy before a customer or regulator does. Voice is where we started; chat is next.
You'll be building the system that keeps other AIs honest.
AI at the core, not bolted on
We build our own product the way we tell our clients to build theirs. Supervisr is managed in plain English. A compliance officer describes a rule, an expected behavior or a test scenario in a sentence, and the platform turns it into the structured configuration, JSON and test cases behind the scenes.
We use LLMs, MCP servers and disciplined prompt and context engineering inside the product itself, not as a demo feature. If you've been frustrated working on "AI features" stapled onto a conventional app, this is the opposite.
What you'll do
You'll own the delivery of features that make AI agents trustworthy in production. That means:
- Design real-time detection for voice AI agents that flags critical compliance breaches mid-call: unauthorized commitments, prohibited statements, hallucinated facts, tone and disclosure failures
- Build the intervention layer: constraining, correcting or handing off to a human in real time when it's critical, and surfacing everything else for mid-call or post-call review
- Build the post-call analysis pipeline that reviews every conversation for drift, loops, off-script behavior and rule effectiveness
- Turn plain-language rules into enforceable, testable guardrails, generating the underlying structured config and JSON automatically
- Generate realistic and adversarial test scenarios with LLMs to stress-test agents before and after they go live
- Extend the platform to chat agents using the same supervision model
- Build the audit trail and dashboards that let a compliance officer replay exactly what an agent did and why
- Platform work: scalable backend services and APIs, event-driven architecture, agent orchestration via MCP, and CI/CD so we ship fast with confidence
You'll love this role if
- You're fascinated by how LLMs fail, not just how they succeed
- You'd rather build the referee than another player
- You think plain English should be the interface, and structured config should be the machine's job
- You want AI in regulated industries to actually be held accountable
Our stack
Python, Java and TypeScript with an Angular UI, built cloud native on Google Cloud. We orchestrate compliance workflows at scale and integrate multiple LLMs through Vertex AI, with MCP servers connecting the pieces.
What we're looking for
Must have
- 5+ years building production software
- Strong Python
- Cloud experience (Google Cloud Platform or similar) and container-based deployments
- Hands-on experience evaluating or observing LLM behavior in production: evals, guardrails, observability
- Comfort with streaming or real-time data
- You write clean, testable, maintainable code
- Strong communicator and mentor
Nice to have
- Voice, telephony or speech-to-text experience
- MCP or agent tooling
- Angular (or React) and modern frontend frameworks
- Compliance or governance workflows
- Contact center or regulated industry background (insurance, financial services)
- Infrastructure as code, Terraform, CI/CD best practices
Why Supervisr.ai?
- Work on the front line of AI agent supervision, a problem that is only getting bigger
- Small team. Big impact. Lots of ownership
- Flexible work and remote culture
- Competitive compensation
Interested in this role?
Send us your application.
