Grader Labs / shippingbench
Benchmark · customs document review

Can a model tell a filing blocker from something that only looks wrong?

Each task contains the nine documents for one export shipment. We plant one issue that would stop the customs filing and include six harmless inconsistencies. The model must find the real blocker without turning the harmless ones into false alarms.

Tasks available40
Tasks in this study5
Recorded attempts125
Best result11 / 25
August 2026 · five attempts per model on each study task · no network

What the score means

One task
Nine related documents describe one shipment. They include one real blocker and six plausible distractions.
One attempt
The model reviews the folder once, names anything that blocks filing, and explains what it cleared.
A successful attempt
The model finds the planted blocker and reports no harmless issue as blocking.
Find rate
Successful attempts divided by all attempts. For example, 11/25 is 44%.

Model comparison

five tasks · 25 attempts for each measured model
#ModelFind rate Successful attempts
filled width equals the percentage printed beside it; every bar uses a 0–100% scale

Which tasks were solved

successful attempts out of five; an empty bar means 0/5
TaskFiling blocker Fable 5GPT-5.6Opus 5 Gem FlashGem Pro

This view matters because one average can hide very different behavior. Task 006 was solved in 16 of 25 attempts across the five models; task 005 was never solved.

Try one · task 008

A cumin-seed shipment from Jaipur, moving by sea out of Mundra to Hamburg. Seven observations from its documents; six are normal trade practice. Pick the one that stops the filing.

task-008-documents.pdf · the full set models receive Open in new tab ↗
The full nine-document set, exactly as the models receive it. Open the PDF ↗
clicktap a row to reveal the verdict · 15 of 25 frontier model attempts missed this set
the exact prompt models get (with the full PDF)

Why models miss

We grouped 20 earlier GPT-5.5 attempts by the kind of check required. The model usually compared facts already printed in the folder, but struggled when it had to notice a missing item, apply an outside rule, or validate a number by itself.

GPT-5.5 · one attempt per task · 20 graded tasks · each bar uses its own sample count
"Chamber fumigation of goods already sealed inside the container two days earlier is physically impossible without breaking the seal."
Fable 5 on task 004, once in five attempts.

How we check task quality

01
Generate one shipment
one fact set; every field derived, none sampled independently
02
Plant one blocker
missing requirement, outside rule, bad self-check, or conflict between documents
03
Render nine documents
layouts calibrated on real exporters' filings
04
Check the answer
a build-time checker confirms that the rendered pages support the expected answer
05
Hide the answer
the model receives only the documents and prompt, with no network access
06
Grade afterward
the runner receives the fingerprinted scoring criteria only after the model exits
40 / 40
Current tasks build and pass the corpus checks before they enter the dataset.
7 answers refused
The checker has stopped seven expected answers because the documents did not support them.
1.0 vs 0.0
In Harbor, the reference answer scores 1.0. An answer that misses the blocker and flags a distraction scores 0.0.
0/5 → 3/5
A model found an unintended second defect in task 008. After we fixed the document and reran it, GPT-5.6-sol moved from 0/5 to 3/5.

Reasoning effort

Fable 5 · same five-task panel
EffortFind rateRateAttempts

Reasoning effort is the test-time compute budget: higher settings let the same model spend more tokens checking the folder. The default point comes from 11/25 recorded attempts; the rest of the curve is modeled from the observed task spread.

Inspect the tasks or run your model

Forty tasks are packaged in Harbor with the scoring criteria hidden until the model exits. Ask for the dataset, a private split, the QA records, or a measured column on this page.

Request full dataset