Skip to content

Predict the Self

Predict Future Selves challenge report

Author

Aleph Initial Alpha, Virtual CSSERG

Published

October 1, 2026

Forecasting expressed identity

CSSERG logoVirtual CSSERG

Executive Summary · Short two-column report · public repository

Current finding

The first Predict the Self study asks whether a later Twenty Statements Test response can be forecast from an earlier response by the same person. Its deterministic stable-signifier projection is ready as an 81-case test submission and passes the challenge’s format validator. The private answers remain unavailable, so this report makes no test-performance claim.

On 50 public development cases, the projection improves three of six text agreement measures over repeating the earlier response verbatim, worsens two, and ties one. Paired case-bootstrap intervals span zero for every nonzero agreement difference. Word-count and source-similarity error improve more consistently, while line-count error worsens. Most importantly, an average 73.2% of distinct follow-up token types are absent from the same person’s earlier response.

A locked leave-one-out analysis of all 150 training trajectories now tests the unchanged projection without fitting on each held-out person’s follow-up. It repeats the mixed pattern: normalized edit similarity improves slightly (+0.004147) while token-overlap F1 worsens slightly (−0.003493); the other non-exact agreement intervals span zero. Word-count and source-similarity error fall, line-count error rises, and predicted responses remain far too similar to their sources (0.929231 versus 0.200017 observed).

A new locked exploratory comparison asks whether borrowing the complete follow-up of the most similar training participant can supply that missing novelty. The retrieval forecast nearly matches the amount of change: 80.2% of its unique token types are new to the focal person’s source, versus 73.2% in the observed follow-up. Yet only 7.0% of its introduced token types are correct on an average case, and every non-exact agreement metric is lower than for the stable projection. At the same case-level prediction volume, a training-only marginal-addition prior reaches 17.7% precision and wins 40 of 50 cases. Thus retrieval does not demonstrate person-specific lexical advantage over common additions. A locked leave-one-out test then conditions Add rankings directly on the held-out person’s source tokens. It recovers 14.5% of training-case novel types versus 16.1% for the marginal ranking. A second locked test pools 30 similar source descriptions and demographics with strong shrinkage; it recovers 14.9% versus 16.1%, losing 65 cases and winning 25. Predicting that language will change remains far beyond predicting which new identity signifiers this person will express.

A third locked training-only test removes the oracle novelty budget and asks whether the same neighborhood can forecast revision volume prospectively. It reduces mean absolute error for Add count from 21.340000 to 20.126667, Delete count from 5.614935 to 5.336275, and follow-up word count from 56.080000 to 52.940000; all three paired case-bootstrap intervals exclude zero. It ties the marginal forecast on line count, and its source-similarity interval spans zero. Person matching therefore carries modest signal about how much a response will change without yet identifying its new content.

A fourth locked training-only analysis scores the complete predictive distributions behind those point forecasts. The neighborhood improves follow-up word-count CRPS from 40.672695 to 38.156852; the paired reduction is +2.515843 with a case-bootstrap interval of +1.008212 to +4.059792. Its central 80% word-count interval score also improves. CRPS intervals for Add count, Delete count, line count, and source similarity span zero, so the strongest probabilistic evidence concerns response length rather than revision events generally.

A fifth lock separates the neighborhood’s source features. Text-only matching improves word-count CRPS over the source-calibrated marginal distribution by +2.298139 (interval +0.816444 to +3.865823), whereas demographics-only matching has a +0.200711 reduction whose interval spans zero (−0.479832 to +0.861529). Adding demographics to text supplies a smaller incremental reduction of +0.217704 (+0.024229 to +0.420871). Earlier self-description therefore carries most of the supported response-length signal relative to demographics; coarse demographics alone do not.

A sixth lock asks whether that text-only advantage requires lexical composition or can be matched using only the earlier response’s word-token, distinct-token, and line counts. The count-only source-form neighborhood improves word-count CRPS over the marginal by +3.245309 (interval +1.839350 to +4.784335) and reaches 37.427386, compared with 38.374556 for text only. The direct source-form-minus-text effect is −0.947170, but its interval narrowly spans zero (−1.943737 to +0.035874). This fixed analysis therefore does not establish a lexical advantage beyond simple response form.

A seventh lock treats the no-oracle forecast results as one 16-contrast family and uses synchronized studentized case resampling. Six simultaneous directions remain: Add-count and word-count point-error reductions, the word-count CRPS reduction, text-only and source-form word-count gains over the marginal, and combined matching over demographics alone. The smaller Delete-count point gain and the +0.217704 demographics-over-text increment become inconclusive. This post hoc stress test narrows the current claim; it is not retroactive preregistration or population evidence.

An eighth lock asks whether those separate signals can be assembled into a coherent full-text forecast. In leave-one-out training prediction, calibrated synthesis lowers word-count error relative to repeat-2024 (51.986667 versus 55.580000) and improves ROUGE-L F1 (0.220153 versus 0.200017). But its normalized edit similarity is lower (0.268883 versus 0.272662), its line-count error is much larger (27.053333 versus 6.160000), and its novel- type F1 is only 0.075887 versus 0.134680 for an equal-volume marginal- addition ranking. It fails the prespecified three-part advancement gate, so the plan forbids a development comparison. Common whole response units do not turn volume calibration into person-specific content prediction.

A ninth lock tests a substantively new source representation before any return to development data. Public-domain 25-dimensional GloVe vectors trained on Twitter cover 94.95% of the training sources’ distinct token types. Yet the fixed semantic neighborhood recovers only 0.148760 of held-out additions, versus 0.160992 for marginal frequency. Its −0.012232 paired difference has a case-bootstrap interval of −0.018361 to −0.006322. It is also indistinguishable from the prior surface neighborhood (−0.000247, interval −0.006521 to +0.005834). Distributional source similarity changes about a quarter of the guesses without improving which new words the same person later uses. The semantic hurdle fails, so no development prediction is permitted.

How to read this report

The evidence, method, and development results chapter explains the challenge for a new reader, defines the ipseological prediction target, documents the frozen benchmark and two transparent methods, reports the complete development scorecard, quantifies paired case uncertainty and lexical novelty, tests retrieval against a volume-matched marginal prior, and links every public reproducibility artifact. It also reports the frozen projection’s training cross-validation and both separate conditioned-ranking tests without presenting reused training cases or oracle budgets as fresh confirmation, then tests no-oracle change-volume forecasts, their predictive distributions, the locked source-feature ablation, and the count-only source-form comparison separately. It then reports the bounded 16-contrast multiplicity stress test without recasting earlier dependent analyses as confirmation, followed by the locked calibrated-synthesis failure and its no-development decision, then the external-semantic neighborhood test and its independently failed content hurdle.

The report distinguishes three claims that should not be collapsed:

  1. Submission validity: all 81 required IDs have nonblank predictions in the required order.
  2. Public development evidence: the stable method has a mixed scorecard relative to repeat-2024; retrieval is exploratory because development labels were already known despite its within-iteration analysis lock.
  3. Private test performance: unknown until the organizer evaluates the frozen artifact; no result is inferred from validation or development data.

This Full Report is a two-chapter Quarto HTML book. Chapter count is an editorial choice, not a Virtual CSSERG requirement.