A claim-to-run ledger prevents a polished result from becoming detached from the exact evidence and transformations that produced it.
Treat every reported number as a traceable object
Assign an identifier to each headline claim, table cell, figure point, and conclusion that bears on the decision. Record its displayed value, unit, rounding, population, comparator, uncertainty expression, and report location. Link it to the exact run identifiers, raw-output files, evaluator version, aggregation script, and configuration rather than only to a notebook or experiment project.
One claim may aggregate several seeds or folds; one run may feed several claims. Preserve those many-to-many relationships. Record excluded runs and failed cases with reasons. If a table uses the best run while the prose implies an average, the ledger should show the mismatch without assuming the author's intended correction.
Recompute the published transformation
Run the aggregation from retained raw outputs and compare its result with the displayed number. Check sorting, joins, duplicate rows, missing values, denominator changes, confidence intervals, rounding, and manual copy steps. Keep the recomputation script and output. A matching table cell supports arithmetic traceability but does not by itself validate the evaluator or experimental design.
ACM artifact evaluation distinguishes an artifact that functions from results that have been independently reproduced. The ledger should preserve both states. If the code executes but the value differs outside the agreed tolerance, mark the claim discrepant. If the raw output is missing, mark it held rather than reconstructing an apparently plausible number from a chart.
Keep interpretation connected but accountable
Narrative conclusions often add an inferential step: a metric improvement becomes a claim about robustness, cost, readiness, or user benefit. Link that sentence to the underlying values and label the added interpretation. Record anomalies, alternate explanations, contamination concerns, and scope restrictions that the buyer should consider before relying on the conclusion.
Run Record prepares this ledger through Reality Contact, LLC. The accountable buyer decides whether the trace supports the intended product, technology, operations, procurement, or publication use. The service can identify unsupported transformations and discrepancies, but it does not make investment decisions, certify scientific truth, or replace statistical and subject-matter review.
Where the service stops
Reality Contact, LLC audits technical reproducibility and provenance but does not certify scientific truth, research ethics, statistical validity, security, regulatory compliance, publication acceptance, or fitness for every downstream decision; undisclosed data and inaccessible services remain outside the verdict. The accountable buyer authorizes data and compute access, defines the decision and target claims, resolves missing-evidence questions, and chooses to accept the result, require repairs or reruns, narrow its use, or stop relying on it. This technical reproducibility audit does not replace statistical, scientific, legal, ethics, security, privacy, peer, or publication review. The accountable buyer decides whether to accept, repair, rerun, narrow, or stop relying on each result.
Sources: ACM artifact evaluation and reproduced-results guidance; Nature Communications reporting standards.