Benchmark methodology
How we test PO Autopilot.
Accuracy should be measured where mistakes matter, so the benchmark will report both what PO Autopilot resolves and what it correctly holds for review. No performance figures are published here until the underlying Golden QA results are approved.
Methodology first
What the benchmark is designed to answer
Can the app turn varied, imperfect B2B purchase orders into correct Shopify Draft Orders without hiding uncertainty? The evaluation should preserve real document structure, Shopify catalogue context, buyer-specific mappings, and the distinction between an accurate match and an appropriate request for review.
No invented headline number
The metric areas are ready, but the page intentionally contains no accuracy, speed, or volume claims. Results will be populated only from a versioned test run with an identified dataset and review process.
What the benchmark will report
- Purchase orders and line items tested
- Source formats and document-quality mix
- Exact product-resolution accuracy
- Silent-error rate
- Correct hold-for-review rate
- Company, location, and contact accuracy
- Quantity and UOM conversion accuracy
- Price, inventory, shipping, and duplicate detection
- End-to-end Shopify Draft Order correctness
- Processing-time distribution
- Worst-performing document category
- App version and test date
Nothing in this list has been published yet. Each figure is added only once an approved Golden QA run supports it.
Data needed before results can be published
- Approved Golden QA run export and dataset description
- App version, environment, and test date
- PO and line-item counts by format and difficulty
- Definitions for correct match, correct hold, and silent error
- Human review or adjudication procedure
- Known limitations and excluded cases
Why review behavior belongs in the benchmark
A system that guesses at every line can appear highly automated while producing expensive silent errors. PO Autopilot is deliberately conservative, so a correct request for merchant judgment is a successful outcome—not a failure to automate.
Read the safety design