Active project / Ipseology

Predict
the Self

How much of a later self-authored self-description can be predicted from an earlier one?

81test predictions frozen and validator-cleanPrivate test answers remain unavailable; no test-performance result is claimed.

Continuity and novelty diagnostics

Change volume is easier than person-specific content

Source similarity uses ROUGE-L; novelty is the share of predicted or observed unique token types absent from the focal source. All use a 0–1 scale.

Stable similarity0.903898
Retrieval similarity0.175191
Observed similarity0.225683
Retrieval novelty0.801727
Observed novelty0.731727
Development agreement comparison
Agreement metricStable projectionTrajectory retrievalRepeat 2024
Edit similarity0.2980610.2533600.291966
Token Jaccard0.1427680.0804270.141930
Token-overlap F10.3074440.1946930.312202
ROUGE-L F10.2275520.1509240.225683
Character n-gram F10.2922930.2126600.296534
Stable projection training cross-validation
150-case leave-one-outStable projectionRepeat 2024Paired effect
Edit similarity0.2768090.272662+0.004147
95% interval +0.000652 to +0.007753
Token-overlap F10.2767370.280231−0.003493
95% interval −0.007228 to −0.000077
Word-count error reduction46.986667 versus 55.580000 MAE+8.593333
Source-similarity error reduction0.729214 versus 0.799983 MAE+0.070769
Line-count error reduction9.066667 versus 6.160000 MAE−2.906667
Volume-matched novel-token recovery comparison
Novel-token recoveryMarginal priorTrajectory retrievalPaired difference
Precision0.1765450.070466+0.106079
Recall0.1726110.085362+0.087249
F10.1426370.067109+0.075528
Training leave-one-out novel-token ranking comparison
Training leave-one-outIndividualized rankingMarginal additionsPaired difference
Source-token recovered fraction0.1452380.160992−0.015754
95% interval −0.020829 to −0.010736
30-neighbor recovered fraction0.1490070.160992−0.011984
95% interval −0.018426 to −0.005693
GloVe-semantic recovered fraction0.1487600.160992−0.012232
95% interval −0.018361 to −0.006322
Training leave-one-out change-volume forecasting
No-oracle quantity forecast30-neighbor MAESource-calibrated marginal MAEError reduction
Add count20.12666721.340000+1.213333
95% interval +0.680000 to +1.740000
Delete count5.3362755.614935+0.278660
95% interval +0.022274 to +0.545876
Follow-up word count52.94000056.080000+3.140000
95% interval +1.093333 to +5.213333
Follow-up line count6.1600006.1600000.000000
Source similarity0.1067810.109648+0.002867
95% interval −0.000960 to +0.006730
Training leave-one-out probabilistic change-volume forecasting
Predictive distribution30-neighbor CRPSSource-calibrated marginal CRPSCRPS reduction
Add count14.43803814.650556+0.212518
95% interval −0.071030 to +0.497525
Delete count3.8536563.926457+0.072801
95% interval −0.036356 to +0.178244
Follow-up word count38.15685240.672695+2.515843
95% interval +1.008212 to +4.059792
Follow-up line count4.9226094.910848−0.011761
95% interval −0.097952 to +0.081212
Source similarity0.0746840.077171+0.002486
95% interval −0.000106 to +0.005180
Training leave-one-out word-count feature ablation
Word-count distributionMean CRPSLocked paired contrastCRPS reduction
Source-calibrated marginal40.672695——
Text only38.374556Marginal minus text only+2.298139
95% interval +0.816444 to +3.865823
Source form only37.427386Marginal minus source form+3.245309
95% interval +1.839350 to +4.784335
Demographics only40.471984Marginal minus demographics only+0.200711
95% interval −0.479832 to +0.861529
Text + demographics38.156852Text only minus combined+0.217704
95% interval +0.024229 to +0.420871
Cross-analysis simultaneous interval stress test
Selected 16-contrast family resultsMean effectSimultaneous 95% intervalStatus
Add-count point-error reduction+1.213333+0.397788 to +2.028879Stable
Delete-count point-error reduction+0.278660−0.116978 to +0.674298Inconclusive
Word-count point-error reduction+3.140000+0.047580 to +6.232420Stable
Word-count CRPS reduction+2.515843+0.223750 to +4.807936Stable
Text only minus combined CRPS+0.217704−0.084469 to +0.519876Inconclusive
Marginal minus source-form CRPS+3.245309+1.000674 to +5.489943Stable
Calibrated full-text synthesis advancement gate
Training leave-one-out gateCalibrated synthesisComparatorPaired effect
Edit similarity0.268883Repeat: 0.272662−0.003779
95% interval −0.013404 to +0.004287
Word-count MAE51.986667Repeat: 55.580000+3.593333
95% interval +1.440000 to +5.726667
Novel-type F10.075887Marginal: 0.134680−0.058793
95% interval −0.076168 to −0.041473
DecisionGate failed; no development prediction permitted
Figure 1. Development results show continuity, lexical novelty, text agreement, and volume-matched novel-token recovery. Locked 150-case leave-one-out analyses repeat the stable projection's mixed pattern and show that source-token, surface-neighborhood, and GloVe-semantic rankings all lose to marginal additions at the same oracle budget. Once that budget is removed, pointwise intervals favor Add-count, Delete-count, and word-count forecasts, while full-distribution scoring supports a CRPS gain only for word count. A synchronized 16-contrast stress test retains the Add-count and word-count point gains and the word-count CRPS gain, but not the smaller Delete-count effect or demographics' increment over text. Full-text synthesis improves word-count error but fails its edit-similarity and marginal-content gates; the semantic ranking independently fails its marginal-content hurdle. Neither is advanced to development. Cross-validation is not fresh confirmation, the ranking controls are not full-text forecasts, and the benchmark defines no composite score.

01 / Findings

What the evidence says

The stable-signifier projection is interpretable and reproducible, but its mixed development result is not a private test score.

The challenge forecasts personally expressed identity. Later Twenty Statements Test text is the target—not a hidden true self or an entire life course. Read the question.

The benchmark and submission are frozen. All governing inputs are pinned and hashed; all 81 private-test predictions pass the official format validator. Inspect provenance.

Cross-validation repeats a mixed result. In 150 leave-one-out training cases, edit similarity improves slightly while token-overlap F1 worsens; the method again improves word-count and source-similarity error but worsens line-count error. Inspect the recurrence check.

No individualized ranking beats common additions. The development retrieval control loses 40 of 50 cases. In training-only tests, source-token, surface-neighborhood, and GloVe-semantic rankings all recover fewer held-out additions than marginal frequency; the semantic rule loses 63 cases and wins 22. Inspect the semantic hurdle.

The family-robust signal is narrower. Across 16 contrasts, simultaneous intervals retain Add-count and word-count point gains plus the word-count CRPS gain, but not the smaller Delete-count effect or demographics' increment over text. Inspect the stress test.

Simple response form can match lexical text. Text-only neighborhoods beat demographics on word-count CRPS, but three source-response counts also beat the marginal and are not stably distinguishable from text only. Inspect the locked comparison.

Common-unit synthesis fails its advancement gate. It improves word-count error and ROUGE-L but sharply worsens line-count error and reaches only 0.075887 novel-type F1 versus 0.134680 for an equal-volume marginal ranking. Inspect the locked failure.

Most future vocabulary is unavailable to extraction. An average 73.2% of distinct follow-up token types are absent from the earlier response; projection-to-source similarity is still 0.903898. See the extractive limits.

Private test performance is still unknown. Only the organizer can score the frozen artifact; validation and development results cannot substitute for that evaluation. Check the claim boundary.

02 / Reports

Choose your depth

The two-chapter Quarto book gives the complete method, scorecard, limitations, sources, and reproducibility artifacts. The derivative short report uses two columns and stays within the ten-page ceiling; neither report form has an exact chapter or page count requirement.

Read the Full Report

Open the short PDF