Contra Machina
Human-in-the-loop ops for AI companies

Your AI handles 99% of the work. We handle the hard 1%.

A dedicated, trained team that works your AI's exceptions and checks its output, in your tools or custom tooling we build for you, priced per item. You grade us on your own work before we touch anything live.

Straight answers on who we are, security and price. No sales sequence.

  • You approve every person with access
  • Shadow-mode start
  • Priced per item
  • Month-to-month
your-queue · live24/7
Your AI · 99%Contra Machina · 1%
12,840 items130 items
  • PRM-0118Permit application
    Auto-resolved
    AI
  • REC-5520Bank reconciliation
    Mismatch fixed, logged to schema
    HUMAN ✓
  • TX-0932Sales-tax filing prep
    Auto-resolved
    AI
  • SUB-7714Loss-run extraction
    Auto-resolved
    AI
  • LD-2291Load #2291
    POD recovered from carrier
    HUMAN ✓
  • INV-48213Invoice extraction
    Auto-resolved
    AI

Straight answers first

Who are your people, and where?

100+ reviewers in the US and Rwanda, employed and managed by Contra Machina, working in dedicated teams on your schedule and in your time zone.

Who can see our data?

Only people you approve by name, inside your tools, with accounts you control. Revoke anyone instantly. We keep no copies.

Will you pass our security review?

SOC 2 audit in progress, report pending. We complete your questionnaire before any pilot. If your contracts bar offshore access, we'll say so on the first call.

What does it cost?

Per item, from $1.50. Monthly minimum: $1,500. Exact quote after your blind test.

Contract?

Month-to-month. Leave with 30 days' notice and keep everything we built.

The problem

The hard 1% is where AI companies stall.

Your model resolves most cases on its own. The rest (low-confidence outputs, edge cases, escalations, missing documents) lands in a queue a person has to work, and it grows with every customer you sign.

You're hiring instead of building.

Every new customer means another "Operations Specialist" posting, another round of interviews, another onboarding.

Your founders are the night shift.

When the queue backs up, founders and engineers work it at midnight.

Every surge is a fire drill.

A big customer onboards or a seasonal spike hits, and the queue doubles overnight.

Your model gets evals. Your reviewers get spot checks.

You can quote your model's accuracy to three decimal places. You can't say how many errors your reviewers miss.

Your margins look like a services company.

Every hire adds cost to every job, and investors notice.

0open roles

We read the job boards of 175 AI startups in September 2026. 89 were hiring for operations or review roles: 218 open positions.

HIRINGOperations SpecialistHIRINGStaff AccountantHIRINGPermit CoordinatorHIRINGTrack-and-Trace RepHIRINGBilling SpecialistHIRINGIntake SpecialistHIRINGUnderwriting AssistantHIRINGQA AnalystHIRINGEscalation SpecialistHIRINGTakeoff SpecialistHIRINGFiling SpecialistHIRINGAI TrainerHIRINGOperations SpecialistHIRINGStaff AccountantHIRINGPermit CoordinatorHIRINGTrack-and-Trace RepHIRINGBilling SpecialistHIRINGIntake SpecialistHIRINGUnderwriting AssistantHIRINGQA AnalystHIRINGEscalation SpecialistHIRINGTakeoff SpecialistHIRINGFiling SpecialistHIRINGAI Trainer
What we do

One team for everything your AI can't finish.

Resolve

We work exceptions until they're closed: chase missing documents, fix mismatches, dispute charges, and email and call customers, vendors and carriers on your behalf. We follow your SOP and work from your systems.

Missing documentsMismatchesDisputesEscalations

Verify

We check outputs against their sources before they reach your customers: every output, or only the low-confidence ones you route to us.

ExtractionsFilingsReconciliationsSummaries

Improve

We log every field we check, including the ones we confirm are correct, in your schema and error taxonomy, straight into your warehouse, bucket or API.

Your schemaYour taxonomyJSONL / CSVWebhook

Your tools, or ours.

We work inside the tools you already have. No review tool yet, or one that slows reviewers down? We build custom tooling for you: review queues, side-by-side review screens, QA dashboards and integrations with your stack. You keep it.

Review queuesReview screensQA dashboardsIntegrations

Faster turnaround. Work keeps moving overnight, so items your team would pick up tomorrow are done by morning.

Priced per item. You only pay for work that actually needs a person, and your cost falls as your automation rate rises.

How we start

Live in 14 days, without risking a single customer.

  1. Days 1–5
    Map

    We start from your golden set and SOP, or help you build them from past edge cases.

    Plan on a few hours of your team's time in week one, more if your definitions are complex.

  2. Days 6–13
    Train and certify

    Reviewers train on your examples, then pass a certification on a held-out set they've never seen, including errors you plant.

    Certification decides when we go live, not the calendar.

  3. Day 14
    Shadow mode

    We work your live queue in parallel. Nothing reaches your customers or partners unless your team approves it.

    Overnight drafts wait for you in a morning batch. Urgent items follow your escalation rules.

You decide when we graduate.

Category by category, starting with routine follow-ups.

Who's on your queue

The same trained people on your queue, day and night.

A small dedicated team, not a pool.

The same reviewers learn your SOP and work your queue, with a lead who knows it. Anyone covering your nights is certified on your golden set too.

Where they are.

100+ reviewers in the US and Rwanda, working your schedule in your time zone. You see every name and location before anyone gets access.

How we hire.

Graduates who pass a paid work test with planted errors, matched to your work by background.

Why quality doesn't walk out the door.

Your know-how lives in your SOP and golden set, not in one person's head. Replacements certify before they touch your queue, and team continuity is in every weekly report.

24/7
on your queue
Night shiftUS dayEvening

Two ways to work with us.

Either way, Contra Machina handles management and HR.

You send the work

Managed team

Route items to us and our employees work them inside your tools, with a team lead, QA and weekly reporting. You manage outcomes, not people.

Best for exception queues and verification work you want off your plate entirely.

Embedded in your team

Forward-deployed reviewers

Our human-in-the-loop specialists join your team as contingent workers: in your Slack, your standups, your tools. You direct the work day to day. We handle management and HR.

Best for teams that want dedicated people who feel in-house, without hiring them.

Quality

Every quality number, recomputable by you.

Outsourcing goes wrong when the vendor grades itself. So you hold the answer key.

You plant the errors.

Seed errors of your own types into the live queue anytime. Test items look exactly like real ones.

You adjudicate the re-review.

Every week, a random sample of our work goes back to your team or your ground truth. You rule on it, not us.

You get the raw logs.

Per-field records, including fields we confirmed correct and fields we couldn't verify, so you can recompute every metric.

What we report every week

At field level and by error type, plus SLA, turnaround and team continuity.

Catch rate
your seeded errors caught ÷ seeded errors planted
False-correction rate
correct fields we changed ÷ correct fields reviewed
Residual error rate
errors your team finds in the re-review sample ÷ fields in the sample
Exception outcomes
share closed without escalation, time to close, and the outcomes you define

A week's seeded sample is small, so we show counts rather than decimals and trend them over time.

Weekly quality report
Monday 08:00 · your inbox
Example · illustrative
Items worked
0
Exceptions · outputs verified
1,203 · 3,609
Returned within SLA
0 of 4,812
Your seeded errors caught
0 of 96
False corrections
3 of 41,650 correct fields
Residual errors, your re-review
2 of 2,880 fields
Exceptions closed without you
0 of 1,203
Median time to close
7h 40m
Same team as last week
6 of 6 reviewers
Seeded-error catch rate · 8 weeks
Top error type
Wrong amount extracted · 31%
Suggested fix
Cross-check line totals against the invoice total

Every field, in your schema.

Delivered to your warehouse, bucket, webhook or API as JSONL or CSV, using your error taxonomy. We keep no copy. Your evals get real data, and your queue gets smaller.

corrections.jsonl → your warehousecorrectedconfirmedcant_verify
1{
2 "item_id": "inv_48213",
3 "field": "line_items[3].amount",
4 "status": "corrected",
5 "model_value": 1240.00,
6 "corrected_value": 12400.00,
7 "error_type": "magnitude_error",
8 "evidence": { "doc": "invoice_48213.pdf", "page": 2 },
9 "model_version": "extractor-v14",
10 "reviewer": "rv_014",
11 "reviewed_at": "2026-10-05T03:12:44Z"
12}
Security

You control who sees what, and from where.

You approve every person.

Before anyone gets access, you see their name, location and background-check status. You can revoke anyone, instantly.

We work inside your tools.

Reviewers use your review UI, admin panel or virtual desktop with accounts you create. Logs and corrections are written into your systems, and we keep no copies. Custom tooling we build for you can be deployed in your own environment.

Locked-down access.

SSO or 2FA, least-privilege permissions, company-managed devices, no downloads, and access logs you can audit.

Paperwork first.

We sign your NDA and DPA before any data moves, share our subprocessor list, and complete your security questionnaire before the pilot. SOC 2 audit in progress, report pending.

Our limits, stated plainly.

We don't handle PHI today, and we're not a fit if your customer contracts bar offshore access.

People with access to your systemsapproved by you
  • AM
    Reviewer · team lead
    location verified · Queue, review UI
  • JK
    Reviewer
    location verified · Queue, review UI
  • SN
    Reviewer
    location verified · Queue
  • DO
    Reviewer · nights
    location verified · Queue
Access log
03:12 rv_014 reviewed inv_48213 (review UI)
03:09 rv_022 opened load #2291 (queue)

Try it: flip a switch to revoke access.

Use cases

Instead of hiring for these roles, run them with us.

If you're building
Finance and accounting ops AI
Roles you'd otherwise hire
Staff accountantsFiling specialistsOps analysts
What we run
  • Reconciliation and categorization QA
  • Sales-tax filing prep
  • AR and billing exceptions
  • KYB document checks
  • Fund-admin data QA
Compare

Three ways to staff the hard 1%.

Hire in-houseGeneric outsourcing Contra Machina
Time to startWeeks per hire, then trainingDays, but untrained on your work14 days, certified on a held-out test
Who grades qualityYou, by spot checksThe vendor grades itselfYou: your seeded errors, your re-review, raw logs
Ramp riskNew hires learn on live workLearns on live workShadow mode: nothing ships without your approval
SurgesAnother hiring roundContract changeMore certified reviewers in days
ToolingYou build and maintain itTheir platform, or you adaptYour tools, or custom tooling we build for you
Who can access your dataYour employeesOften a rotating poolOnly people you approve by name
Pricing$50k–$90k base per US hire*, plus benefits and management timeHourly, per seatPer item, month-to-month

*Range from AI-startup job postings for ops and review roles, September 2026.

The offer

Grade us before you pay.

A free blind test on your own work. You hold the answer key.

01

Start with a 20-minute call.

We answer the hard questions (who, where, security, price) and scope the test together.

02

Use our starter set, or your own items.

We bring 50 synthetic items for your industry (for freight: missing PODs, rate-con mismatches, TONU disputes). Tweak them and write your answers in about an hour. Or use your own golden set, inside your tools if data can't leave.

03

Plant your own errors.

We work every item blind, in your format.

04

Score us.

You get our corrections, proposed actions and escalation calls, plus a per-item quote: a head-to-head comparison with your current team.

Book a 20-min call

Fifty items is a smoke test, not a benchmark. It checks judgment. The shadow pilot checks follow-through on live volume.

Rather start by email?

Request your blind test. Takes about 30 seconds.

What's in your queue?
Coverage you need
Team working the queue today (optional)
Requirements (optional)

One email from Flo or Honore within one business day. No sales sequence.

Pilot and founding clients

We're new. So we carry the risk.

You'd be one of our first 5 clients. Here's what that gets you.

FTHT
Flo Turati & Honore Twagirayezu
Founders · ex-Meta (Wearable AI data team), ex-Amazon

A 2-week pilot in shadow mode.

Your live queue, your approval on everything, and daily quality reports.

$2,500 flat, and only if we hit the targets.

We agree on catch rate, false corrections, residual errors and SLA upfront. If we miss, the pilot is free. If we hit them, the fee is credited to your first month.

A founder in the loop.

Flo or Honore run your weekly review and are your escalation point for the first 90 days.

Pricing locked for 12 months.

Founding-client pricing doesn't move while you ramp.

No lock-in.

Month-to-month. You keep the SOP, golden set and every log.

Fit

Is this for you?

A great fit if
  • Your AI flags exceptions, or produces outputs a person must check or finish
  • You're hiring for ops, QA or review roles, or your founders are covering the queue
  • You face surges, or need coverage beyond US business hours
  • You want quality numbers you can recompute yourself
Not a fit yet if
  • Your work involves PHI, or your customer contracts bar offshore access
  • You need SOC 2 Type II today
  • Your contracts or marketing promise US-based staff
  • The work needs a licensed professional's judgment (we do the prep and the checks, not the sign-off)
FAQ

Questions buyers ask us.

One exception worked until it's closed (a two-week dispute is still one item), one document, or page-based bands for long documents. We agree on the definition during your blind test.

Give us your hardest 50 items. Then decide.

Start with a 20-minute call. We'll answer the hard questions and set up your free blind test.