# Checks added after the run **None of this was decided before the study ran.** These are the checks a sceptical reader would ask for, run on 23 September 2026 after the main results. Every number here comes from `checks.cjs` in this folder; its output is `checks.json`, which also records the sha256 of every file it loaded. It takes about nine minutes and gives the same numbers on every run. ```sh node checks.cjs ``` Section numbers match the script. Counts are diaries out of 100 unless marked. A is the plain search, B the strictly corrected search, C the method I built. ## 0. The checks use the study's own diaries The script rebuilds the study's diaries with a copy of the generator that has switches (weekly rhythm, day to day persistence, slow wander, whole number outcome). With every switch at the published setting it produced the same diary files, byte for byte, as the study script for 48 diaries across all four lengths. It also reproduces every number in `results.json`. ## 1. How strong the planted pattern is The method file says a rank correlation of about 0.30. Measured on the 100 planted diaries per length, steps two days before against the outcome: | | 60 days | 90 days | 180 days | 292 days | |---|---|---|---|---| | Raw | 0.405 | 0.370 | 0.391 | 0.369 | | After weekday removal | 0.407 | 0.382 | 0.398 | 0.384 | On one planted diary 400,000 days long: 0.3825 raw, 0.3955 after weekday removal. **The planted pattern is about 0.38, not 0.30.** ## 2. The calendar correlation in the empty diaries At 292 days, the average same day rank correlation between the outcome and the three measures that share its weekend bump is 0.063; for the other measures it is about zero. After weekday removal it is about zero for all of them. On one very long diary the weekend correlation is 0.066. It is real, small, and about the calendar, not the person. ## 3. Where the plain search's 97 comes from (292 days) | Diaries built with | Flagged | Each single check wrong | |---|---|---| | Pure random numbers | 79 | 4.6 in 100 | | Day to day persistence only | 93 | 10.9 in 100 | | Weekly rhythm only | 86 | 7.3 in 100 | | Both, as published | 97 | 12.2 in 100 | 32 checks at 5 in 100 each would flag about 81 diaries by arithmetic alone. Persistence makes each check wrong more often than it promises. ## 4. Other corrections, on the same empty diaries | | 60 days | 90 days | 180 days | 292 days | |---|---|---|---|---| | A. No correction | 90 | 88 | 95 | 97 | | B. Benjamini and Yekutieli, q below 0.06 (the study's B) | 6 | 11 | 12 | 10 | | Benjamini and Yekutieli, q below 0.05 | 6 | 8 | 11 | 9 | | Benjamini and Hochberg, q below 0.05 | 11 | 24 | 25 | 32 | | Bonferroni, 0.05 | 11 | 23 | 25 | 31 | | B after weekday removal | 12 | 14 | 6 | 10 | Bonferroni promises at most 5 in 100 whatever the checks' relationship to each other, but only when nothing is there to find. The weekend bump (section 2) puts a small real link into some checks, so the 31 is not a clean test of that promise. The clean test is the same diaries, on the same random stream, with the bump switched off, so that every one of the 32 checks has nothing to find: | No weekend bump | 60 days | 90 days | 180 days | 292 days | |---|---|---|---|---| | Bonferroni, 0.05 | 15 | 22 | 14 | 19 | | Benjamini and Hochberg, q below 0.05 | 16 | 22 | 15 | 19 | | Benjamini and Yekutieli, q below 0.06 | 8 | 10 | 6 | 10 | Still three to four times Bonferroni's promise: the p values going in are too small. Removing the weekday from the published diaries instead gives Bonferroni 25, 24, 18 and 19 and does not rescue the corrected search. ## 5. Everyone held to method C's 21 comparisons Method C leaves out pain and counts one to three days before only: 21 comparisons, against 32 for A and B. | | 60 days | 90 days | 180 days | 292 days | |---|---|---|---|---| | A on the same 21 | 77 | 78 | 78 | 87 | | B on the same 21 | 4 | 9 | 9 | 9 | | A on the same 21, after weekday removal | 84 | 82 | 83 | 88 | | B on the same 21, after weekday removal | 12 | 12 | 7 | 9 | | C | 3 | 4 | 2 | 0 | | C, also counting its separate same day list | 3 | 5 | 3 | 0 | ## 6. The engine without the weekday step This is what the served engine file does on its own. | | 60 days | 90 days | 180 days | 292 days | |---|---|---|---|---| | Empty diaries flagged | 2 | 2 | 4 | 1 | | Also counting same day links | 3 | 3 | 7 | 4 | | Planted pattern found | 17 | 27 | 75 | 93 | On these diaries the weekday step barely matters. The block resampling does the work. ## 7. Planted diaries: clean finds and false names | | 60 days | 90 days | 180 days | 292 days | |---|---|---|---|---| | Found it: A / B / C | 72 / 29 / 16 | 82 / 38 / 29 | 96 / 80 / 71 | 100 / 97 / 96 | | Found it and named nothing false: A / B / C | 5 / 24 / 16 | 9 / 36 / 28 | 13 / 64 / 68 | 6 / 80 / 95 | | Named a wrong measure: A / B / C | 94 / 10 / 3 | 88 / 6 / 2 | 86 / 18 / 3 | 94 / 17 / 1 | | B held to C's 21 checks: found / clean | 34 / 31 | 41 / 37 | 83 / 70 | 98 / 86 | | Same, after weekday removal: found / clean | 35 / 29 | 46 / 38 | 85 / 71 | 99 / 89 | A and B are scored on all 32 checks, which include the same day and pain; C never counts those against itself. Scored on the same 21 checks, B has more clean finds than C up to 180 days, and C has more at 292 days. C names a wrong measure least often. ## 8. Fresh seeds 200 new empty diaries per length, on seeds that share nothing with the study's (a hash of the study seed plus 500,000; the rule is in the script). | Of 200 | 60 days | 90 days | 180 days | 292 days | |---|---|---|---|---| | A | 191 | 187 | 184 | 189 | | B | 12 | 16 | 32 | 18 | | C | 6 | 6 | 10 | 2 | | C, also counting same day | 7 | 7 | 11 | 4 | Method C's rate on fresh diaries is 1 to 5 in 100, in line with its pooled rate of about 2 in 100. The study's 0 at 292 days was its best cell. ## 9. Stress tests at 292 days **Slow wander, where method C breaks.** Each column, outcome included, gets its own independent random walk added: every day it moves by a random step with the standard deviation shown. For scale, the day to day noise has a standard deviation of about 1.15. | Daily wander step | A | B | C | C with same day | |---|---|---|---|---| | 0.03 | 98 | 14 | 2 | 3 | | 0.05 | 99 | 28 | 5 | 7 | | 0.10 | 99 | 62 | 35 | 36 | | 0.15 | 100 | 89 | 49 | 50 | **Other stresses.** Two versions of each: the "review" rows reproduce the first review's script, which drew one extra random number per day and so lands on a different random stream; the "study" rows use the study's own stream exactly. The two agree within chance. | Variant | Stream | A | B | C | |---|---|---|---|---| | Stronger persistence, phi 0.8 | review | 98 | 58 | 2 | | Stronger persistence, phi 0.8 | study | 100 | 64 | 1 | | Stronger persistence, phi 0.9 | review | 100 | 85 | 1 | | Stronger persistence, phi 0.9 | study | 100 | 90 | 2 | | Outcome as whole numbers 0 to 10 | review | 88 | 13 | 2 | | Outcome as whole numbers 0 to 10 | study | 96 | 11 | 0 | | Published settings, hashed seeds | review | 96 | 9 | 5 | | Published settings, hashed seeds | study | 92 | 13 | 2 | ## 10. Exact intervals and tests Clopper and Pearson 95% intervals, in percent. | | 60 days | 90 days | 180 days | 292 days | |---|---|---|---|---| | Empty, A | 90 (82.4 to 95.1) | 88 (80.0 to 93.6) | 95 (88.7 to 98.4) | 97 (91.5 to 99.4) | | Empty, B | 6 (2.2 to 12.6) | 11 (5.6 to 18.8) | 12 (6.4 to 20.0) | 10 (4.9 to 17.6) | | Empty, C | 3 (0.6 to 8.5) | 4 (1.1 to 9.9) | 2 (0.2 to 7.0) | 0 (0.0 to 3.6) | | Planted, A | 72 (62.1 to 80.5) | 82 (73.1 to 89.0) | 96 (90.1 to 98.9) | 100 (96.4 to 100) | | Planted, B | 29 (20.4 to 38.9) | 38 (28.5 to 48.3) | 80 (70.8 to 87.3) | 97 (91.5 to 99.4) | | Planted, C | 16 (9.4 to 24.7) | 29 (20.4 to 38.9) | 71 (61.1 to 79.6) | 96 (90.1 to 98.9) | Pooled over all four lengths: C flagged 9 of 400 empty diaries (1.0 to 4.2%), B flagged 39 of 400 (7.0 to 13.1%). Fisher exact test, C against B: p = 0.0000085. On the same 100 diaries at 292 days, B flagged 10 that C did not and C flagged none that B did not: exact McNemar p = 0.002. Pooled over all 400 paired diaries, B alone flagged 34 and C alone 4 (both 5, neither 357): exact McNemar p = 0.0000006. McNemar is the right test here, because both methods ran on the same diaries; the Fisher test treats them as separate samples. ## 11. The older fix: prewhitening The standard time series answer to "both columns lean on yesterday" is to take out each column's own leaning first and correlate what is left. Per column: the mean and the day to day coefficient are estimated from the logged pairs, and each day becomes its value minus the coefficient times the day before. Then the plain grid, textbook p values, and Benjamini and Yekutieli at q below 0.06. No resampling anywhere. | Without the weekday step | 60 days | 90 days | 180 days | 292 days | |---|---|---|---|---| | Empty diaries flagged, 32 checks | 1 | 6 | 4 | 2 | | Empty diaries flagged, C's 21 checks | 1 | 3 | 2 | 2 | | Planted found, 21 checks | 9 | 21 | 57 | 91 | | Planted found and nothing false, 21 checks | 7 | 20 | 56 | 89 | | With the weekday step first | 60 days | 90 days | 180 days | 292 days | |---|---|---|---|---| | Empty diaries flagged, 21 checks | 3 | 8 | 3 | 3 | | Planted found, 21 checks | 14 | 29 | 61 | 91 | On C's 21 checks, prewhitening flagged 8 of 400 empty diaries, against C's 9: the same false alarm rate. C found the planted pattern more often at every length (16, 29, 71 and 96 against 9, 21, 57 and 91). So what C's resampling adds over the older fix, on these diaries, is finds, not fewer false alarms. Neither was tested on slow wander here.