LOPO split — evaluation rerun
- Reproduce last week’s numbersrerun the baseline before changing anythingDone
- Fix the loader’s subject shufflefind why fold scores were too goodDone
- Rerun folds 1–4rerun with the fixed loaderDone (unverified)
- Rerun folds 5–6fold 5 is trainingIn progress
- Explain the subject-7 gaplook at why subject 7 scores 12 pt below the restNot started
Reproduce last week’s numbersturns 0–38 ↗
Done
Doing: rerun the baseline before changing anything
The baseline reproduced to within 0.2 pt on every fold, so later differences come from the fix, not the environment.
Fix the loader’s subject shuffleturns 39–97 ↗
Done
Doing: find why fold scores were too good
Subject IDs were reshuffled every epoch, so a subject could land in both train and test. Three attempts are condensed here; the one that held seeds the split once, before the first epoch.
split_subjects(seed=…)is now called once per run- a test asserts no subject appears on both sides
Rerun folds 1–4turns 98–161 ↗
Done (unverified)
Doing: rerun with the fixed loader
| Fold | Before | After |
|---|---|---|
| 1 | 0.91 | 0.84 |
| 2 | 0.89 | 0.83 |
| 3 | 0.92 | 0.86 |
| 4 | 0.90 | 0.82 |
Lower, as expected once the leak is gone. Not yet checked against a second seed.
Rerun folds 5–6turns 162–201 ↗
In progress
Doing: fold 5 is training
Fold 6 waits for fold 5; both use the same config as folds 1–4.
Explain the subject-7 gapturns 202–213 ↗
Not started
Doing: look at why subject 7 scores 12 pt below the rest
Noted at the end of the session; no work on it yet.
Summary
Done: loader leak fixed and tested; folds 1–4 rerun.
Left: folds 5–6; a second seed for 1–4; the subject-7 gap.
Written by your own agent from the redacted record of this conversation. Edit it freely: a later analysis will not overwrite your edits.