Proof treats evaluation as an engineering discipline. Fixtures define inputs and expected behaviour. Model responses are saved and can be replayed, so scoring is deterministic and repeatable. Claims must be backed by evidence, and catastrophic failures are detected on their own rather than averaged away.
Proof also checks itself. When a result looks wrong, it distinguishes a genuine model failure from an evaluator false positive or a problem in the gold corpus, and the scorers and corpus are verified for correctness in their own right.
Relay sits around AI actions at execution time. Tool use is bounded by explicit policy, every action is recorded with its provenance, and the evidence needed to verify the work later is captured as it happens.