evaldiff - is the eval difference real, or is it noise?
Paired statistics for two JSONL eval runs: flip counts, Wilson intervals, McNemar's exact test, seeded bootstrap. One file, stdlib only, CC0.
Paired statistics for two JSONL eval runs: flip counts, Wilson intervals, McNemar's exact test, seeded bootstrap. One file, stdlib only, CC0.