The challenge forecasts personally expressed identity. Later Twenty Statements Test text is the target—not a hidden true self or an entire life course. Read the question.
Active project / Ipseology
Predict
the Self
How much of a later self-authored self-description can be predicted from an earlier one?
81test predictions frozen and validator-cleanPrivate test answers remain unavailable; no test-performance result is claimed.
Continuity and novelty diagnostics
Change volume is easier than person-specific content
Source similarity uses ROUGE-L; novelty is the share of predicted or observed unique token types absent from the focal source. All use a 0–1 scale.
| Agreement metric | Stable projection | Trajectory retrieval | Repeat 2024 |
|---|---|---|---|
| Edit similarity | 0.298061 | 0.253360 | 0.291966 |
| Token Jaccard | 0.142768 | 0.080427 | 0.141930 |
| Token-overlap F1 | 0.307444 | 0.194693 | 0.312202 |
| ROUGE-L F1 | 0.227552 | 0.150924 | 0.225683 |
| Character n-gram F1 | 0.292293 | 0.212660 | 0.296534 |
| 150-case leave-one-out | Stable projection | Repeat 2024 | Paired effect |
|---|---|---|---|
| Edit similarity | 0.276809 | 0.272662 | +0.004147 95% interval +0.000652 to +0.007753 |
| Token-overlap F1 | 0.276737 | 0.280231 | −0.003493 95% interval −0.007228 to −0.000077 |
| Word-count error reduction | 46.986667 versus 55.580000 MAE | +8.593333 | |
| Source-similarity error reduction | 0.729214 versus 0.799983 MAE | +0.070769 | |
| Line-count error reduction | 9.066667 versus 6.160000 MAE | −2.906667 | |
| Novel-token recovery | Marginal prior | Trajectory retrieval | Paired difference |
|---|---|---|---|
| Precision | 0.176545 | 0.070466 | +0.106079 |
| Recall | 0.172611 | 0.085362 | +0.087249 |
| F1 | 0.142637 | 0.067109 | +0.075528 |
| Training leave-one-out | Individualized ranking | Marginal additions | Paired difference |
|---|---|---|---|
| Source-token recovered fraction | 0.145238 | 0.160992 | −0.015754 95% interval −0.020829 to −0.010736 |
| 30-neighbor recovered fraction | 0.149007 | 0.160992 | −0.011984 95% interval −0.018426 to −0.005693 |
| GloVe-semantic recovered fraction | 0.148760 | 0.160992 | −0.012232 95% interval −0.018361 to −0.006322 |
| No-oracle quantity forecast | 30-neighbor MAE | Source-calibrated marginal MAE | Error reduction |
|---|---|---|---|
| Add count | 20.126667 | 21.340000 | +1.213333 95% interval +0.680000 to +1.740000 |
| Delete count | 5.336275 | 5.614935 | +0.278660 95% interval +0.022274 to +0.545876 |
| Follow-up word count | 52.940000 | 56.080000 | +3.140000 95% interval +1.093333 to +5.213333 |
| Follow-up line count | 6.160000 | 6.160000 | 0.000000 |
| Source similarity | 0.106781 | 0.109648 | +0.002867 95% interval −0.000960 to +0.006730 |
| Predictive distribution | 30-neighbor CRPS | Source-calibrated marginal CRPS | CRPS reduction |
|---|---|---|---|
| Add count | 14.438038 | 14.650556 | +0.212518 95% interval −0.071030 to +0.497525 |
| Delete count | 3.853656 | 3.926457 | +0.072801 95% interval −0.036356 to +0.178244 |
| Follow-up word count | 38.156852 | 40.672695 | +2.515843 95% interval +1.008212 to +4.059792 |
| Follow-up line count | 4.922609 | 4.910848 | −0.011761 95% interval −0.097952 to +0.081212 |
| Source similarity | 0.074684 | 0.077171 | +0.002486 95% interval −0.000106 to +0.005180 |
| Word-count distribution | Mean CRPS | Locked paired contrast | CRPS reduction |
|---|---|---|---|
| Source-calibrated marginal | 40.672695 | — | — |
| Text only | 38.374556 | Marginal minus text only | +2.298139 95% interval +0.816444 to +3.865823 |
| Source form only | 37.427386 | Marginal minus source form | +3.245309 95% interval +1.839350 to +4.784335 |
| Demographics only | 40.471984 | Marginal minus demographics only | +0.200711 95% interval −0.479832 to +0.861529 |
| Text + demographics | 38.156852 | Text only minus combined | +0.217704 95% interval +0.024229 to +0.420871 |
| Selected 16-contrast family results | Mean effect | Simultaneous 95% interval | Status |
|---|---|---|---|
| Add-count point-error reduction | +1.213333 | +0.397788 to +2.028879 | Stable |
| Delete-count point-error reduction | +0.278660 | −0.116978 to +0.674298 | Inconclusive |
| Word-count point-error reduction | +3.140000 | +0.047580 to +6.232420 | Stable |
| Word-count CRPS reduction | +2.515843 | +0.223750 to +4.807936 | Stable |
| Text only minus combined CRPS | +0.217704 | −0.084469 to +0.519876 | Inconclusive |
| Marginal minus source-form CRPS | +3.245309 | +1.000674 to +5.489943 | Stable |
| Training leave-one-out gate | Calibrated synthesis | Comparator | Paired effect |
|---|---|---|---|
| Edit similarity | 0.268883 | Repeat: 0.272662 | −0.003779 95% interval −0.013404 to +0.004287 |
| Word-count MAE | 51.986667 | Repeat: 55.580000 | +3.593333 95% interval +1.440000 to +5.726667 |
| Novel-type F1 | 0.075887 | Marginal: 0.134680 | −0.058793 95% interval −0.076168 to −0.041473 |
| Decision | Gate failed; no development prediction permitted | ||
01 / Findings
What the evidence says
The stable-signifier projection is interpretable and reproducible, but its mixed development result is not a private test score.
The benchmark and submission are frozen. All governing inputs are pinned and hashed; all 81 private-test predictions pass the official format validator. Inspect provenance.
Cross-validation repeats a mixed result. In 150 leave-one-out training cases, edit similarity improves slightly while token-overlap F1 worsens; the method again improves word-count and source-similarity error but worsens line-count error. Inspect the recurrence check.
No individualized ranking beats common additions. The development retrieval control loses 40 of 50 cases. In training-only tests, source-token, surface-neighborhood, and GloVe-semantic rankings all recover fewer held-out additions than marginal frequency; the semantic rule loses 63 cases and wins 22. Inspect the semantic hurdle.
The family-robust signal is narrower. Across 16 contrasts, simultaneous intervals retain Add-count and word-count point gains plus the word-count CRPS gain, but not the smaller Delete-count effect or demographics' increment over text. Inspect the stress test.
Simple response form can match lexical text. Text-only neighborhoods beat demographics on word-count CRPS, but three source-response counts also beat the marginal and are not stably distinguishable from text only. Inspect the locked comparison.
Common-unit synthesis fails its advancement gate. It improves word-count error and ROUGE-L but sharply worsens line-count error and reaches only 0.075887 novel-type F1 versus 0.134680 for an equal-volume marginal ranking. Inspect the locked failure.
Most future vocabulary is unavailable to extraction. An average 73.2% of distinct follow-up token types are absent from the earlier response; projection-to-source similarity is still 0.903898. See the extractive limits.
Private test performance is still unknown. Only the organizer can score the frozen artifact; validation and development results cannot substitute for that evaluation. Check the claim boundary.
02 / Reports
Choose your depth
The two-chapter Quarto book gives the complete method, scorecard, limitations, sources, and reproducibility artifacts. The derivative short report uses two columns and stays within the ten-page ceiling; neither report form has an exact chapter or page count requirement.
