Follow-up study · Issue #1
Personal calibration and signal alignment for motor-imagery EEG
Comparing four approaches after ten initial recordings, with evaluation on later runs
Exploratory reanalysis of an existing benchmark, not an independent replication or a peer-reviewed publication
Research question and findings
This follow-up investigates whether initial recordings from a new participant can improve later motor-imagery decoding. Confidence calibration adjusts the probabilities attached to a prediction, while signal alignment transforms the inputs before classification. Four conditions were compared across three training seeds, using the same 630 later-run trials from 21 test participants for each comparison. The original production decoder and its reported results remain unchanged.
Mean participant-macro balanced accuracy was 61.28% for the baseline and 64.36% for Euclidean alignment, a paired change of +3.08 percentage points (95% participant-bootstrap interval -1.24 to +7.63). The interval includes no change, conditional on this split and these three fitted model pairs.
None of the twelve condition–seed combinations met the predefined validation reliability target, so every operating point uses the declared fallback; this comparison does not establish reliable control.
1 Study design
The question arose in a discussion with neowalter, who asked whether confidence calibration or participant-specific alignment would be the more useful next step. The follow-up protocol was committed before fitting these models, but the dataset and its original test outcomes had already been inspected. This study therefore provides exploratory evidence on the existing benchmark, with no claim of a newly untouched test population.
The original train, validation and test participant assignments and quality exclusions were preserved. For every participant, the first ten eligible trials from run 04 were selected chronologically without regard to their class labels or prediction correctness. Adaptation parameters use only this initial block; all eligible trials from runs 08 and 12 form the later evaluation block. Remaining run-04 trials are unused for validation/test adaptation and scoring. The 22 validation participants supply 660 evaluation trials, while the 21 test participants supply 630. These are later runs within the dataset, not independent-day sessions or live recordings.
| Condition | Initial recordings | Labels required | Model input and confidence |
|---|---|---|---|
| Baseline | 0 | 0 | Unaligned CNN with global temperature |
| Personal temperature | 10 | 10 | Same CNN with personal temperature |
| Euclidean alignment | 10 | 0 | Aligned CNN with global temperature |
| Alignment + personal temperature | 10 | 10 | Same aligned CNN with personal temperature |
Euclidean alignment uses the inverse square root of each participant’s mean channel second-moment matrix, estimated from their ten initial preprocessed recordings. A fixed 1% isotropic shrinkage and relative eigenvalue floor of 10−6 stabilise the matrix. The resulting transform is applied without refitting on later trials or renormalising after transformation. This is a regularised adaptation of He and Wu’s approach [1]; it does not implement Riemannian alignment.
Paired unaligned and aligned CNNs were fitted for seeds 2027, 2028 and 2029, with the same architecture, training participants, minibatch order and training settings within each pair. The aligned models were trained on consistently aligned data, rather than applying a new transform only at inference. All 63 training participants contribute their eligible training epochs. Checkpoints are selected using participant-macro balanced accuracy on validation later runs, with all six models frozen before test scoring.
Global temperatures are fitted on validation later-run logits, while personal temperatures minimise binary log loss on each participant’s ten initial labelled trials, bounded to [0.5, 5.0]. Single-class initial blocks retain the global value. Temperature scaling follows the post-hoc calibration principle described by Guo and colleagues [2]; it leaves left/right decisions unchanged, so any benefit must be assessed through probability quality or command selection rather than a gain in ordinary classification accuracy.
2 Classification and probability quality
| Condition | Balanced accuracy | Change from baseline 95% paired interval | Brier score | ECE |
|---|---|---|---|---|
| Baseline | 61.28% | Reference | 0.2315 | 5.85% |
| Personal temperature | 61.28% | +0.00 pp (+0.00, +0.00) | 0.2215 | 4.14% |
| Euclidean alignment | 64.36% | +3.08 pp (-1.24, +7.63) | 0.2109 | 5.87% |
| Alignment + personal temperature | 64.36% | +3.08 pp (-1.24, +7.63) | 0.2063 | 3.75% |
Balanced accuracy is calculated within each participant and then averaged, giving people equal weight. The change and interval are in percentage points: for each participant, scores are first averaged over the three seeds, then paired differences are bootstrapped over participants 2,000 times. These intervals condition on the chosen split and fitted models; the three seeds are not treated as additional independent people. Brier score is the mean squared error of the predicted right-hand probability against the binary label. Expected calibration error (ECE) uses ten equal-width confidence bins over [0, 1], weighted by the number of trials in each occupied bin. Both are pooled over trials within each seed before averaging across seeds, with smaller values indicating better probability quality under these metrics; ECE is sensitive to binning and sample size.
Personal temperature scaling leaves each model’s classifications unchanged, as expected. Its mean test Brier score was 0.2215, compared with 0.2315 for the global-temperature baseline. With alignment, the corresponding values were 0.2063 for personal temperature and 0.2109 for global temperature. These comparisons describe probability quality rather than an improvement in discrimination. For personal temperature versus baseline, the paired Brier change was -0.0100 (95% participant-bootstrap interval -0.0239 to +0.0016); this interval includes no change under the same conditional interpretation as the accuracy comparison.
Among test-participant personal temperature fits, 31 of 63 unaligned fits and 35 of 63 aligned fits reached a parameter bound; 0 fits used the single-class fallback. This count includes a separate fit for each seed and should not be read as a count of distinct people. A ten-trial confidence fit can be unstable, even though the parameter itself is only one-dimensional.
Inspect each training seed
| Condition | Seed | Balanced accuracy | Brier score | ECE |
|---|---|---|---|---|
| Baseline | 2027 | 61.63% | 0.2285 | 5.47% |
| Personal temperature | 2027 | 61.63% | 0.2222 | 3.44% |
| Euclidean alignment | 2027 | 66.34% | 0.2086 | 5.35% |
| Alignment + personal temperature | 2027 | 66.34% | 0.2031 | 4.22% |
| Baseline | 2028 | 62.04% | 0.2333 | 7.53% |
| Personal temperature | 2028 | 62.04% | 0.2184 | 4.25% |
| Euclidean alignment | 2028 | 61.70% | 0.2135 | 8.42% |
| Alignment + personal temperature | 2028 | 61.70% | 0.2117 | 3.83% |
| Baseline | 2029 | 60.17% | 0.2326 | 4.54% |
| Personal temperature | 2029 | 60.17% | 0.2238 | 4.72% |
| Euclidean alignment | 2029 | 65.03% | 0.2105 | 3.84% |
| Alignment + personal temperature | 2029 | 65.03% | 0.2041 | 3.19% |
3 Reliability, thresholds and coverage
For each condition and seed, the threshold-selection rule uses validation later-run predictions only. It chooses the lowest threshold from 0.50 to 0.95 in steps of 0.05 that achieves at least 75% retained-trial accuracy, at least 20% coverage and at least 100 retained examples. If none qualifies, 0.75 is the declared fallback. Table 4 additionally applies the same fixed 0.75 confidence threshold to every condition, showing how probability rescaling affects the number of accepted commands.
| Condition | Mean coverage | Mean accepted per seed | Accepted accuracy pooled across seeds | Validation target met seeds |
|---|---|---|---|---|
| Baseline | 1.22% | 7.7 | 100.00% | 0/3 |
| Personal temperature | 22.43% | 141.3 | 86.56% | 0/3 |
| Euclidean alignment | 8.62% | 54.3 | 99.39% | 0/3 |
| Alignment + personal temperature | 28.52% | 179.7 | 86.83% | 0/3 |
Accepted accuracy pools correct and accepted predictions across the three seeds, so a trial can contribute up to three predictions. Those repeated predictions do not create 1890 independent examples: the evaluation still contains 630 unique trials from 21 people. High accuracy on a small selected subset cannot establish useful command throughput, and failure of the validation criterion is not reversed by a favourable test subset.
Inspect validation-selected thresholds and test counts
| Condition / seed | Threshold | Validation target met | Test accepted | Test correct | Test accepted accuracy |
|---|---|---|---|---|---|
| Baseline / 2027 | 0.75 | No — fallback | 9/630 | 9 | 100.00% |
| Personal temperature / 2027 | 0.75 | No — fallback | 144/630 | 124 | 86.11% |
| Euclidean alignment / 2027 | 0.75 | No — fallback | 60/630 | 60 | 100.00% |
| Alignment + personal temperature / 2027 | 0.75 | No — fallback | 182/630 | 158 | 86.81% |
| Baseline / 2028 | 0.75 | No — fallback | 3/630 | 3 | 100.00% |
| Personal temperature / 2028 | 0.75 | No — fallback | 135/630 | 120 | 88.89% |
| Euclidean alignment / 2028 | 0.75 | No — fallback | 47/630 | 47 | 100.00% |
| Alignment + personal temperature / 2028 | 0.75 | No — fallback | 154/630 | 134 | 87.01% |
| Baseline / 2029 | 0.75 | No — fallback | 11/630 | 11 | 100.00% |
| Personal temperature / 2029 | 0.75 | No — fallback | 145/630 | 123 | 84.83% |
| Euclidean alignment / 2029 | 0.75 | No — fallback | 56/630 | 55 | 98.21% |
| Alignment + personal temperature / 2029 | 0.75 | No — fallback | 203/630 | 176 | 86.70% |
4 Interpretation and next steps
The two interventions address different weaknesses and are not interchangeable. Alignment produced a higher mean classification score, but its paired interval includes no improvement, so this experiment does not establish a dependable accuracy gain. Personal temperature scaling changes confidence and the mix of accepted predictions without improving left/right discrimination. Combining the methods accepted substantially more later test trials than the baseline, yet no condition passed the predefined validation reliability criterion at any seed. That discrepancy between validation and test participants is a reason to seek confirmation on fresh recordings, rather than promote the favourable test operating point. Frequent temperature estimates at the bounds also suggest that ten labelled trials provide limited information for individual confidence fitting. The evidence supports further investigation of alignment alongside more stable calibration, while the original decoder and its unmet reliability claim remain unchanged.
This comparison changes the deployment assumption: personal adaptation has access to recordings from the new participant, whereas the baseline requires none. The same ten-trial budget is held fixed across adaptation conditions, but their need for labels differs. A later-run split also differs from the original all-run evaluation, so the scores above must not be presented as direct replacements for the original 61.1% result. The newly trained baseline provides the matched comparison.
The study uses one dataset and one participant partition, with repeated training seeds but no independent recording environment. Validation later-run data are reused for checkpoint selection, global calibration and threshold selection. Signal alignment cannot eliminate cue, eye-movement or muscle confounds, and confidence calibration on ten trials may overfit. Further work would need fresh evaluation data, repeated participant splits and, if studied, a separately specified regularised or hierarchical temperature estimator.
The original C3/C4 occlusion results still suggest that the existing decoder depends on those inputs, but do not identify an optimal reduced montage. All nine selected channels were originally referenced using 64 electrodes. A smaller headset would need an appropriate reference, new training and evaluation, causal filtering and a rest/no-intention condition before prospective testing. This follow-up introduces neither a live headset mode nor a claim of reliable assistive control.
The public discussion also highlighted neoxai’s interest in processing EEG on the user’s own device. This experiment publishes code and public-data outputs that can be reproduced locally, but the existing demonstration still serves inference from the server; a browser-only implementation remains future work.
5 Reproduction and audit trail
The protocol was frozen at commit e78292e and the training implementation at 3f0190f. The completed record includes file hashes, software versions, calibration/evaluation trial IDs, alignment matrices, bounded temperature fits, validation histories, six ONNX exports and trial-level probabilities. The maximum observed PyTorch-to-ONNX logit difference was 8.34e-07, below the predefined 10−4 limit.
Reproduction requires the original preprocessed epoch archive, generated by the public acquisition pipeline, and the pinned training dependencies. The runner refuses to overwrite a completed evaluation. An interrupted run can be resumed only when its recorded implementation, protocol and data hashes match, preventing a silent change of method during an experiment.
pip install -r requirements-training.txt
python scripts/prepare_data.py
python scripts/run_adaptation.py
python scripts/render_adaptation_report.py
pytest -q
The repository already includes the completed study outputs, so a reproduction should use a separate checkout and move its supplied artifacts/adaptation-v1 directory to an archive before running. Keep the archived published results for comparison. Source data and derivatives retain their PhysioNet attribution and licence requirements.
References
- He, H. and Wu, D. (2020). Transfer learning for brain–computer interfaces: a Euclidean space data alignment approach. IEEE Transactions on Biomedical Engineering, 67(2), 399–410. doi:10.1109/TBME.2019.2913914 · Open manuscript
- Guo, C., Pleiss, G., Sun, Y. and Weinberger, K. Q. (2017). On calibration of modern neural networks. Proceedings of the 34th International Conference on Machine Learning, PMLR 70, 1321–1330. Paper and proceedings
Dataset, acquisition-system and architectural references are provided in the original report.