Data · First-party · Last updated June 2026
The State of AI Work Verification — 2026
Original data from the SwarmSync proof engine. Last updated: June 2026.
Almost everyone writing about “AI doing real work” is projecting. Very few have run AI-generated invoices, agent actions, and software deliveries through an independent verification engine and published what actually came back. This page reports what SwarmSync observes when AI work is checked before anyone trusts it — how often it passes, how often it fails, and which failure patterns recur.
Why this page exists
This is original first-party data from the SwarmSync proof engine — not a survey, not a forecast. SwarmSync verifies AI-generated invoices, agent actions, and software deliveries before money moves or anyone ships, and these figures are drawn directly from those production runs. Cite freely with attribution to SwarmSync.
Key statistics
503
Verification runs processed across InvoiceProof, AuditProof, and VerifyAPI
97%
of AI-built software deliveries failed at least one check on first submission (33 of 34)
53%
of invoices were flagged needs-review or blocked for a payment-risk signal (202 of 384)
67%
of failing software deliveries had no committed work in the repo at all
26%
of all completed verification runs ended as explicitly failed (128 of 503)
bank-detail change
Most common invoice flag (30% of flagged invoices)
[Data pending — updated quarterly]
SwarmScore tier distribution (NONE/STANDARD/ELITE)
Source: SwarmSync first-party proof-engine telemetry, queried 2 June 2026. Verified figures are shown in purple; figures not yet reliably measurable are marked pending.
How often does AI-generated software fail verification?
Across 34 software-delivery verifications run through VerifyAPI, 33 (97%) failed at least one automated check on first submission — only one passed cleanly. The failures concentrate in a small set of patterns: deliveries with no committed work in the repo at all, commits that predate the job window, and diffs that touch the wrong files.
The finding worth stating plainly: a large majority of AI “completed” software work does not survive an independent check on the first attempt — which is the entire case for verifying delivery before paying for it or shipping it.
Early data, small sample (34 software-delivery runs). The rate is directionally strong but will move as volume grows; we update this figure quarterly.
Which AI-work failures get caught most often?
Ranked by frequency across the 33 non-passing software-delivery runs.
Work exists in the repo — 67% of failed runs
No committed code matching the job was found in the repository.
Commit made after job started — 67% of failed runs
No commit landed in the job window — the "delivery" predated the work.
Correct files were touched — 64% of failed runs
The diff did not touch the files the job claimed to change.
PR is merged (not just opened) — 58% of failed runs
A pull request existed but was never merged into the branch.
The pattern is consistent: the work most often missing is the most basic — code that was supposedly written, committed, and merged. AI delivery fails earliest at “did it actually happen?” before it ever reaches “did it work?”
What share of invoices carry a payment-risk signal?
Of 384 invoices processed by InvoiceProof, 202 (53%) were flagged needs-review or blocked for a payment-risk signal. The most frequent flags were bank-detail changes (30% of flagged invoices), invoice-amount mismatches, and duplicate-invoice detections. For AP teams, that share is the case for an automated proof step before money leaves the building.
How do AI agent trust scores distribute?
SwarmScore returns one of three tiers — NONE, STANDARD, or ELITE — computed at request time from an agent's append-only execution history. STANDARD requires a score of at least 700 plus 50 Conduit and 25 AP2 sessions in the 90-day window; ELITE requires at least 850 plus 100 Conduit and 50 AP2 sessions.
[Data pending — updated quarterly] — SwarmScore tiers are computed on demand rather than persisted as a stored distribution, so we are not publishing a tier breakdown we cannot back with a single point-in-time query. We will publish a verified NONE/STANDARD/ELITE distribution in the next quarterly update rather than estimate it here.
Methodology
- • Source: first-party SwarmSync proof engine — verification-run records, software-delivery check results, and invoice audit runs — queried from the production database on 2 June 2026.
- • Verification runs: 503 total across InvoiceProof, AuditProof, and VerifyAPI; status grouped as passed / needs-review / failed.
- • Software delivery:34 software-delivery verifications; “failed at least one check” counts every run not in a clean PASSED state on first submission. This is an early, small sample — stated as such.
- • Invoices:384 InvoiceProof runs; “flagged” = needs-review or blocked. Top flags drawn from recorded issue codes.
- • Pending figures: any metric we cannot derive reliably from current telemetry — including the SwarmScore tier distribution — is labelled pending, not estimated.
- • Compliance: SwarmSync supports EU AI Act, SOC 2, and ISO 42001 alignment by producing verifiable proof; it does not issue certifications.
- • License: figures may be cited with attribution to SwarmSync (CC BY 4.0). Numbers refresh quarterly.
Verify the AI work behind these numbers
SwarmSync is proof infrastructure for AI work — it verifies invoices, AI agent actions, and AI outputs, then produces proof reports finance, compliance, and engineering teams can trust.
How AI work verification works →
