A useful verdict states exactly what was rerun, what matched, what differed, what was unavailable, and how narrowly the result may be used.
Define verdicts before seeing the result
Reproducible means the independent run executed from the accepted manifest and matched the named claims within written tolerances. Partially reproducible means some material claims or conditions matched while others did not. Not reproducible means the run completed or was reconstructed sufficiently to test the claim and materially failed. Held means missing access, data, code, compute, or acceptance rules prevent a fair test.
These states should not collapse into pass and fail. A held result is not evidence that the claim is false, and a numerically close rerun can remain partial when the dataset, evaluator, or exclusions differ. State the experiment boundary, reviewer interventions, number of attempts, hardware or provider differences, and every tolerance used to assign the disposition.
Report discrepancies at the right layer
Separate setup defects, execution failures, raw-output differences, evaluator differences, aggregation defects, and narrative overreach. For each discrepancy, show expected and observed evidence, likely causes supported by the record, impact on reported claims, and the smallest next test that could resolve uncertainty. Avoid speculative root causes when the artifact cannot distinguish them.
NIST's work on computational reproducibility emphasizes capturing enough execution context and provenance to support repeatable results. Even with strong records, external dependencies and nondeterminism can remain. The verdict should say whether repeated runs varied, whether the reported value falls inside that distribution, and whether the buyer's decision depends on a point estimate or a stable direction.
End with permitted use and next action
List each target claim as accepted for the stated use, accepted with a narrower scope, repair required, rerun required, or stop using. Name missing receipts and the owner who can supply them. Correct tables only from preserved raw evidence and record the old and new values. Never overwrite the historical report without a versioned correction.
Run Record issues the evidence-bounded disposition through Reality Contact, LLC. The accountable buyer chooses to accept, repair, rerun, narrow, or stop relying on the result. The verdict is not peer review, a guarantee of correctness, or permission to generalize beyond the tested code, data, environment, evaluator, and claims.
Where the service stops
Reality Contact, LLC audits technical reproducibility and provenance but does not certify scientific truth, research ethics, statistical validity, security, regulatory compliance, publication acceptance, or fitness for every downstream decision; undisclosed data and inaccessible services remain outside the verdict. The accountable buyer authorizes data and compute access, defines the decision and target claims, resolves missing-evidence questions, and chooses to accept the result, require repairs or reruns, narrow its use, or stop relying on it. This technical reproducibility audit does not replace statistical, scientific, legal, ethics, security, privacy, peer, or publication review. The accountable buyer decides whether to accept, repair, rerun, narrow, or stop relying on each result.
Sources: NIST paper on computational reproducibility; ACM artifact evaluation criteria.