Your AI handles 99% of the work. We handle the hard 1%.
A dedicated, trained team that works your AI's exceptions and checks its output, in your tools or custom tooling we build for you, priced per item. You grade us on your own work before we touch anything live.
Straight answers on who we are, security and price. No sales sequence.
- You approve every person with access
- Shadow-mode start
- Priced per item
- Month-to-month
- AIPRM-0118Permit applicationAuto-resolved
- HUMAN ✓REC-5520Bank reconciliationMismatch fixed, logged to schema
- AITX-0932Sales-tax filing prepAuto-resolved
- AISUB-7714Loss-run extractionAuto-resolved
- HUMAN ✓LD-2291Load #2291POD recovered from carrier
- AIINV-48213Invoice extractionAuto-resolved
Straight answers first
Who are your people, and where?
100+ reviewers in the US and Rwanda, employed and managed by Contra Machina, working in dedicated teams on your schedule and in your time zone.
Who can see our data?
Only people you approve by name, inside your tools, with accounts you control. Revoke anyone instantly. We keep no copies.
Will you pass our security review?
SOC 2 audit in progress, report pending. We complete your questionnaire before any pilot. If your contracts bar offshore access, we'll say so on the first call.
What does it cost?
Per item, from $1.50. Monthly minimum: $1,500. Exact quote after your blind test.
Contract?
Month-to-month. Leave with 30 days' notice and keep everything we built.
The hard 1% is where AI companies stall.
Your model resolves most cases on its own. The rest (low-confidence outputs, edge cases, escalations, missing documents) lands in a queue a person has to work, and it grows with every customer you sign.
You're hiring instead of building.
Every new customer means another "Operations Specialist" posting, another round of interviews, another onboarding.
Your founders are the night shift.
When the queue backs up, founders and engineers work it at midnight.
Every surge is a fire drill.
A big customer onboards or a seasonal spike hits, and the queue doubles overnight.
Your model gets evals. Your reviewers get spot checks.
You can quote your model's accuracy to three decimal places. You can't say how many errors your reviewers miss.
Your margins look like a services company.
Every hire adds cost to every job, and investors notice.
We read the job boards of 175 AI startups in September 2026. 89 were hiring for operations or review roles: 218 open positions.
One team for everything your AI can't finish.
Resolve
We work exceptions until they're closed: chase missing documents, fix mismatches, dispute charges, and email and call customers, vendors and carriers on your behalf. We follow your SOP and work from your systems.
Verify
We check outputs against their sources before they reach your customers: every output, or only the low-confidence ones you route to us.
Improve
We log every field we check, including the ones we confirm are correct, in your schema and error taxonomy, straight into your warehouse, bucket or API.
Your tools, or ours.
We work inside the tools you already have. No review tool yet, or one that slows reviewers down? We build custom tooling for you: review queues, side-by-side review screens, QA dashboards and integrations with your stack. You keep it.
Faster turnaround. Work keeps moving overnight, so items your team would pick up tomorrow are done by morning.
Priced per item. You only pay for work that actually needs a person, and your cost falls as your automation rate rises.
Live in 14 days, without risking a single customer.
- Days 1–5Map
We start from your golden set and SOP, or help you build them from past edge cases.
Plan on a few hours of your team's time in week one, more if your definitions are complex.
- Days 6–13Train and certify
Reviewers train on your examples, then pass a certification on a held-out set they've never seen, including errors you plant.
Certification decides when we go live, not the calendar.
- Day 14Shadow mode
We work your live queue in parallel. Nothing reaches your customers or partners unless your team approves it.
Overnight drafts wait for you in a morning batch. Urgent items follow your escalation rules.
You decide when we graduate.
Category by category, starting with routine follow-ups.
The same trained people on your queue, day and night.
A small dedicated team, not a pool.
The same reviewers learn your SOP and work your queue, with a lead who knows it. Anyone covering your nights is certified on your golden set too.
Where they are.
100+ reviewers in the US and Rwanda, working your schedule in your time zone. You see every name and location before anyone gets access.
How we hire.
Graduates who pass a paid work test with planted errors, matched to your work by background.
Why quality doesn't walk out the door.
Your know-how lives in your SOP and golden set, not in one person's head. Replacements certify before they touch your queue, and team continuity is in every weekly report.
Two ways to work with us.
Either way, Contra Machina handles management and HR.
Managed team
Route items to us and our employees work them inside your tools, with a team lead, QA and weekly reporting. You manage outcomes, not people.
Best for exception queues and verification work you want off your plate entirely.
Forward-deployed reviewers
Our human-in-the-loop specialists join your team as contingent workers: in your Slack, your standups, your tools. You direct the work day to day. We handle management and HR.
Best for teams that want dedicated people who feel in-house, without hiring them.
Every quality number, recomputable by you.
Outsourcing goes wrong when the vendor grades itself. So you hold the answer key.
You plant the errors.
Seed errors of your own types into the live queue anytime. Test items look exactly like real ones.
You adjudicate the re-review.
Every week, a random sample of our work goes back to your team or your ground truth. You rule on it, not us.
You get the raw logs.
Per-field records, including fields we confirmed correct and fields we couldn't verify, so you can recompute every metric.
What we report every week
At field level and by error type, plus SLA, turnaround and team continuity.
- Catch rate
- your seeded errors caught ÷ seeded errors planted
- False-correction rate
- correct fields we changed ÷ correct fields reviewed
- Residual error rate
- errors your team finds in the re-review sample ÷ fields in the sample
- Exception outcomes
- share closed without escalation, time to close, and the outcomes you define
A week's seeded sample is small, so we show counts rather than decimals and trend them over time.
- Items worked
- 0
- Exceptions · outputs verified
- 1,203 · 3,609
- Returned within SLA
- 0 of 4,812
- Your seeded errors caught
- 0 of 96
- False corrections
- 3 of 41,650 correct fields
- Residual errors, your re-review
- 2 of 2,880 fields
- Exceptions closed without you
- 0 of 1,203
- Median time to close
- 7h 40m
- Same team as last week
- 6 of 6 reviewers
Every field, in your schema.
Delivered to your warehouse, bucket, webhook or API as JSONL or CSV, using your error taxonomy. We keep no copy. Your evals get real data, and your queue gets smaller.
1{2 "item_id": "inv_48213",3 "field": "line_items[3].amount",4 "status": "corrected",5 "model_value": 1240.00,6 "corrected_value": 12400.00,7 "error_type": "magnitude_error",8 "evidence": { "doc": "invoice_48213.pdf", "page": 2 },9 "model_version": "extractor-v14",10 "reviewer": "rv_014",11 "reviewed_at": "2026-10-05T03:12:44Z"12}
You control who sees what, and from where.
You approve every person.
Before anyone gets access, you see their name, location and background-check status. You can revoke anyone, instantly.
We work inside your tools.
Reviewers use your review UI, admin panel or virtual desktop with accounts you create. Logs and corrections are written into your systems, and we keep no copies. Custom tooling we build for you can be deployed in your own environment.
Locked-down access.
SSO or 2FA, least-privilege permissions, company-managed devices, no downloads, and access logs you can audit.
Paperwork first.
We sign your NDA and DPA before any data moves, share our subprocessor list, and complete your security questionnaire before the pilot. SOC 2 audit in progress, report pending.
Our limits, stated plainly.
We don't handle PHI today, and we're not a fit if your customer contracts bar offshore access.
- AMReviewer · team leadlocation verified · Queue, review UI
- JKReviewerlocation verified · Queue, review UI
- SNReviewerlocation verified · Queue
- DOReviewer · nightslocation verified · Queue
Try it: flip a switch to revoke access.
Instead of hiring for these roles, run them with us.
- Reconciliation and categorization QA
- Sales-tax filing prep
- AR and billing exceptions
- KYB document checks
- Fund-admin data QA
Three ways to staff the hard 1%.
| Hire in-house | Generic outsourcing | Contra Machina | |
|---|---|---|---|
| Time to start | Weeks per hire, then training | Days, but untrained on your work | 14 days, certified on a held-out test |
| Who grades quality | You, by spot checks | The vendor grades itself | You: your seeded errors, your re-review, raw logs |
| Ramp risk | New hires learn on live work | Learns on live work | Shadow mode: nothing ships without your approval |
| Surges | Another hiring round | Contract change | More certified reviewers in days |
| Tooling | You build and maintain it | Their platform, or you adapt | Your tools, or custom tooling we build for you |
| Who can access your data | Your employees | Often a rotating pool | Only people you approve by name |
| Pricing | $50k–$90k base per US hire*, plus benefits and management time | Hourly, per seat | Per item, month-to-month |
*Range from AI-startup job postings for ops and review roles, September 2026.
Grade us before you pay.
A free blind test on your own work. You hold the answer key.
Start with a 20-minute call.
We answer the hard questions (who, where, security, price) and scope the test together.
Use our starter set, or your own items.
We bring 50 synthetic items for your industry (for freight: missing PODs, rate-con mismatches, TONU disputes). Tweak them and write your answers in about an hour. Or use your own golden set, inside your tools if data can't leave.
Plant your own errors.
We work every item blind, in your format.
Score us.
You get our corrections, proposed actions and escalation calls, plus a per-item quote: a head-to-head comparison with your current team.
Fifty items is a smoke test, not a benchmark. It checks judgment. The shadow pilot checks follow-through on live volume.
We're new. So we carry the risk.
You'd be one of our first 5 clients. Here's what that gets you.
A 2-week pilot in shadow mode.
Your live queue, your approval on everything, and daily quality reports.
$2,500 flat, and only if we hit the targets.
We agree on catch rate, false corrections, residual errors and SLA upfront. If we miss, the pilot is free. If we hit them, the fee is credited to your first month.
A founder in the loop.
Flo or Honore run your weekly review and are your escalation point for the first 90 days.
Pricing locked for 12 months.
Founding-client pricing doesn't move while you ramp.
No lock-in.
Month-to-month. You keep the SOP, golden set and every log.
Is this for you?
- Your AI flags exceptions, or produces outputs a person must check or finish
- You're hiring for ops, QA or review roles, or your founders are covering the queue
- You face surges, or need coverage beyond US business hours
- You want quality numbers you can recompute yourself
- Your work involves PHI, or your customer contracts bar offshore access
- You need SOC 2 Type II today
- Your contracts or marketing promise US-based staff
- The work needs a licensed professional's judgment (we do the prep and the checks, not the sign-off)
Questions buyers ask us.
One exception worked until it's closed (a two-week dispute is still one item), one document, or page-based bands for long documents. We agree on the definition during your blind test.
Give us your hardest 50 items. Then decide.
Start with a 20-minute call. We'll answer the hard questions and set up your free blind test.