Each task contains the nine documents for one export shipment. We plant one issue that would stop the customs filing and include six harmless inconsistencies. The model must find the real blocker without turning the harmless ones into false alarms.
This view matters because one average can hide very different behavior. Task 006 was solved in 16 of 25 attempts across the five models; task 005 was never solved.
A cumin-seed shipment from Jaipur, moving by sea out of Mundra to Hamburg. Seven observations from its documents; six are normal trade practice. Pick the one that stops the filing.
We grouped 20 earlier GPT-5.5 attempts by the kind of check required. The model usually compared facts already printed in the folder, but struggled when it had to notice a missing item, apply an outside rule, or validate a number by itself.
Reasoning effort is the test-time compute budget: higher settings let the same model spend more tokens checking the folder. The default point comes from 11/25 recorded attempts; the rest of the curve is modeled from the observed task spread.
Forty tasks are packaged in Harbor with the scoring criteria hidden until the model exits. Ask for the dataset, a private split, the QA records, or a measured column on this page.