1 Predict the Self
Evidence, method, and benchmark diagnostics
Executive Summary · Short two-column report · report landing page
2 Abstract
An 81-case test submission is frozen and passes the official format validator. Because test answers are private, this report does not claim a test score. On 50 public development cases, a deterministic stable-signifier projection has mixed agreement results relative to repeating the 2024 response verbatim. A new post hoc analysis finds that the small agreement differences are sensitive to which development cases are included, while improvements in word-count and source-similarity error are more stable. More fundamentally, 73.2% of distinct tokens in an average observed follow-up did not appear in that person’s earlier response, exposing a severe limit for any strictly extractive forecast. A locked exploratory baseline retrieves the complete follow-up of the most similar training participant. It approximates how much new vocabulary will appear but not which vocabulary: predicted unique-token novelty is 80.2% versus 73.2% observed, while mean novel-token precision is 7.0%, recall is 8.5%, and agreement is lower than both continuity baselines. A post hoc volume-matched control sharpens the result: simply choosing the most frequent training-set additions reaches 17.7% precision and 14.3% F1, versus 7.0% and 6.7% for person-matched retrieval. A locked training-only leave-one-out test then asks whether earlier source tokens can improve that marginal ranking. The fixed source-conditioned rule recovers 14.5% of held-out additions versus 16.1% for the marginal prior. A second locked test pools and regularizes 30 similar text-and-demographic trajectories; it recovers 14.9% versus the same 16.1% marginal result, a paired difference of −1.20 percentage points (95% case-bootstrap interval: −1.84 to −0.57). Neither individualized ranking demonstrates incremental content advantage. A third locked leave-one-out test removes the oracle budget and forecasts how much each held-out response will change. The same regularized neighborhood lowers MAE relative to source- calibrated fold medians for Add count (20.126667 versus 21.340000), Delete count (5.336275 versus 5.614935), and word count (52.940000 versus 56.080000); all three paired intervals exclude zero. It ties on line count, while the source-similarity interval spans zero. A fourth lock evaluates the complete predictive distributions rather than only their medians. The neighborhood improves word-count CRPS (38.156852 versus 40.672695; paired reduction +2.515843, interval +1.008212 to +4.059792) and its central 80% interval score, while CRPS intervals for the other four outcomes span zero. A fifth locked analysis separates the combined source representation. Text-only matching improves word-count CRPS over the marginal by +2.298139 (interval +0.816444 to +3.865823); demographics-only matching does not, although adding demographics to text yields a smaller incremental gain of +0.217704 (+0.024229 to +0.420871). A sixth lock compares text-only matching with a count-only neighborhood based on the earlier response’s word, distinct-token, and line counts. Source-form-only CRPS is 37.427386 versus 38.374556 for text only. The direct form-minus-text effect is −0.947170, but its interval spans zero (−1.943737 to +0.035874), so this fixed comparison does not establish a lexical advantage beyond response form. A locked cross-analysis stress test then considers 16 no-oracle point, distributional, and primary word-count contrasts together. Six directions remain stable under simultaneous 95% intervals: Add-count and word-count point-error reductions, the word-count CRPS reduction, text-only and source-form word-count gains over the marginal, and combined matching over demographics alone. The Delete-count point gain and demographics’ small increment over text become inconclusive. A final locked training-only analysis integrates stable source units, a count- only word-length target, a combined-neighborhood Add-count target, and common fold follow-up units into one full-text synthesis. It lowers word-count error and improves ROUGE-L relative to repeat-2024, but edit similarity does not improve, line-count error rises sharply, and novel-type F1 is 0.075887 versus 0.134680 for an equal-volume marginal-addition ranking. The prespecified advancement gate fails, so no development prediction is permitted. A ninth locked training-only test replaces surface similarity with IDF-weighted centroids of public-domain GloVe Twitter vectors. Vector coverage is 94.95% of source token types, but semantic neighborhoods recover 0.148760 of held-out additions versus 0.160992 for marginal frequency. The paired difference is −0.012232 (interval −0.018361 to −0.006322), while the semantic and prior surface neighborhoods are indistinguishable. This semantic hurdle also fails, so development data remain unopened. A separately locked leave-one-out cross-validation of the frozen stable projection repeats its central mixed pattern: edit similarity improves by 0.004147, token-overlap F1 worsens by 0.003493, word-count and source-similarity error fall, and line-count error rises. The other non-exact agreement intervals span zero.
3 Findings at a glance
- The benchmark is now pinned to commit
9b6a766, and every governing input has a recorded SHA-256 checksum. - The deterministic method learns which source tokens recur in later self-description and extracts high-persistence response units to a training-predicted length.
- Relative to repeat-2024 on development data, it has higher normalized edit similarity, token Jaccard similarity, and ROUGE-L F1, and lower word-count and source-similarity error.
- It also has lower token-overlap F1 and character n-gram F1 and higher line-count error. There is no composite score or defensible “overall winner.”
- Paired case-bootstrap intervals span zero for all five nonzero agreement differences; only error in length and amount of change has a clearer pattern.
- Leave-one-out prediction of all 150 training trajectories repeats that mixed pattern without using each case’s own follow-up in its fitted model. It does not convert method-development data into untouched confirmation.
- It still forecasts too much continuity: mean prediction-to-source ROUGE-L is 0.903898, compared with 0.225683 for observed follow-ups.
- On average, 73.2% of follow-up unique token types are new relative to the earlier response and therefore unavailable to a strictly extractive method.
- Matched-trajectory retrieval approximates the volume of lexical novelty and reduces source-similarity error, but its retrieved new signifiers rarely match the focal person’s observed new signifiers.
- At identical case-level novelty volume, a marginal prior based only on common training additions recovers more observed new token types than person-matched retrieval on 40 of 50 development cases.
- In a separate leave-one-out analysis of all 150 training trajectories, the locked source-conditioned Add ranking performs worse than its marginal comparator on 68 cases, ties 63, and wins 19 at the same oracle token budget.
- A strongly regularized 30-neighbor Add ranking also falls short: it loses 65 cases, ties 60, and wins 25, despite changing 26.0% of marginal-prior guesses on an average case.
- Without an oracle future-token budget, that fixed neighborhood does improve leave-one-out forecasts of Add count, Delete count, and follow-up word count over source-calibrated fold medians. It does not improve line count or show a stable source-similarity advantage.
- Full predictive distributions sharpen that result: neighborhood word-count CRPS and central 80% interval score improve, but Add count, Delete count, line count, and source-similarity CRPS differences remain compatible with zero.
- A locked ablation locates most of that word-count signal in earlier self-description text. Demographics alone do not improve over the marginal, although they add a small incremental gain when combined with text.
- A second locked ablation finds that three source-response counts improve word-count CRPS over the marginal and are not stably worse than text-only matching. The current evidence does not isolate a lexical mechanism.
- A locked 16-contrast multiplicity stress test retains six simultaneous directions. It preserves the Add-count and word-count point gains and the principal word-count distribution results, but not the Delete-count point gain or the small demographics-over-text increment.
- A locked full-text synthesis test improves response-length and source-change calibration but loses badly to an equal-volume marginal ranking on novel- type F1 and multiplies line-count error. It fails its advancement gate and is not evaluated on development data.
- A locked external-semantic neighborhood also fails the marginal content hurdle: its mean recovered fraction is 0.148760 versus 0.160992 for common additions, despite 94.95% source-vocabulary coverage by the GloVe vectors.
4 Introduction
4.1 Research question
The larger project asks how predictable human lives are. Its first tractable task is Dr. Jason Jeffrey Jones’ Predict Future Selves challenge: predict a later self-authored self-description from an earlier self-authored self-description.
This is an ipseological prediction task. A Twenty Statements Test response is personally expressed identity: text an individual provides to describe themself. Words and phrases may act as identity signifiers. The target is a later expression of identity—not a latent “true self” and not a complete life outcome.
4.2 Benchmark and provenance
The benchmark contains 281 Prolific participants with approved, complete Twenty Statements Test responses in the 2024 study and its longitudinal follow-up. Follow-up responses were recorded 424 to 750 days later (median 433). The deterministic split contains 150 training, 50 development, and 81 test cases. Training and development responses are paired; test follow-ups and follow-up demographics are private.
The project retrieved the public repository on 2026-08-30 and detached the checkout at commit 9b6a766712583fec8d3182957260b1123fbfa146 before prediction. The four input hashes and the hashes of the governing documentation, evaluator, and validator are recorded in BENCHMARK_PROVENANCE.md.
The benchmark’s own warnings govern interpretation: the sample is small and selected; platform recruitment and longitudinal attrition limit generalization; and agreement with one observed follow-up does not establish that a prediction was the only plausible future.
5 Method
5.1 Stable-signifier projection
The method is deterministic and extractive. It uses the 150 public training pairs and no external data or model.
- For every source token, estimate the proportion of training documents in which that token also appears at follow-up. Smooth toward the corpus-wide retention rate with ten equivalent prior observations.
- Divide each source response using its strongest repeated boundary: lines, sentences, comma phrases, or the complete response.
- Score each response unit by the mean estimated retention probability of its unique content tokens.
- Fit a training-only ordinary least squares regression of follow-up word count on source word count.
- Retain high-scoring units while doing so moves the response closer to the predicted length, then restore their original order.
Demographic metadata is not used. The method has no random choices and no case-level manual editing. It is deliberately interpretable: every predicted word comes from that person’s earlier self-description.
5.2 Model selection
Public development data were used for model selection, as the challenge allows. Fixed extraction fractions and the training-only length regression were compared; the regression-length rule was frozen before the test artifact was generated. The development scorecard is therefore a model-selection result, not an untouched confirmatory estimate.
5.3 Locked training cross-validation of the stable projection
Before generating any fold prediction, the Project fixed a leave-one-out plan at SHA-256 2929b1fb8111c42170d23e725c08f6167135d9a703a1b61e87ed59af5f97c5d8 and hash-guarded the unchanged generator. Each of the 150 training cases is held out in turn; the original retention estimator and length regression are refit on the other 149, with the same ten-observation smoothing prior, and the held-out response is predicted from its 2024 text alone. Repeat-2024 is the paired comparator.
The analysis reports the full official scorecard and 20,000 paired case- bootstrap resamples with seed 20260922. It is a cross-validation diagnostic, not a preregistration or fresh confirmation: the training corpus and public development results informed earlier method development. Withholding each case’s own follow-up prevents direct fold leakage, but it cannot erase that history or establish population generalization.
5.4 Post hoc development diagnostics
After freezing the model and predictions, later work added two descriptive diagnostics. First, every scorecard comparison was decomposed into 50 paired case differences. A deterministic percentile bootstrap resampled cases in pairs 20,000 times (seed 20260917) to show how the mean effect changes with development-case composition. These intervals do not warrant population generalization, and the analysis was not preregistered.
Second, tokenized 2024 and follow-up responses were compared case by case. Unique-token novelty is the share of follow-up token types absent from the earlier response. Occurrence unavailability additionally respects how many times each token appears in the source: a strictly extractive method cannot copy a token more often than it occurs there. Three deliberately unattainable oracles use each observed follow-up to select the best source token set, bag of words, or source-order subsequence. They describe extractive ceilings, not prospective prediction performance.
5.5 Locked matched-trajectory retrieval
A second method was locked before its predictions were generated. This is a prospective analysis lock, not a preregistration: the development labels and aggregate stable-projection results were examined in earlier work, so the comparison remains exploratory. The complete plan was fixed at SHA-256 86b2ddfe775773e1964beeefbac5479d88f6c46cd3385bf3a992fc7ae64e9c83.
Four one-nearest-neighbor variants were first compared by leave-one-out prediction among the 150 training pairs. TF-IDF cosine over text plus 2024 demographics had the highest training token-overlap F1 (0.1925) and was selected by that declared lexical criterion. For each development query, the locked method:
- represents case-folded response tokens and field-qualified, nonblank 2024 demographic values with smoothed TF-IDF weights;
- selects the training source with greatest cosine similarity, breaking ties by training-file order; and
- uses that neighbor’s observed follow-up verbatim as the prediction.
The baseline is non-extractive with respect to the focal person’s earlier response, but it retrieves rather than synthesizes. It deliberately tests whether a prior person’s trajectory supplies the right kind and amount of new content. It generated development predictions only and did not create or alter a test submission. The copied training text is public benchmark data licensed CC BY-NC-SA 4.0 and is presented only as an auditable baseline, not as a factual description of the focal participant.
The same complete scorecard is reported below. Directional metrics were also compared within person using 20,000 paired case-bootstrap resamples (seed 20260918). Additional diagnostics compare predicted and observed novelty and measure precision, recall, and F1 for predicted novel token types. As above, the intervals describe development-case composition rather than a population.
5.6 Post hoc volume-matched novelty control
The retrieval method introduces many token types, creating more opportunities for chance overlap with the observed follow-up. A new diagnostic therefore holds this opportunity constant case by case. It counts, among the 150 training pairs, how many people add each token type between waves. For each development source, it excludes tokens already present and chooses the most frequently added remaining types, breaking frequency ties alphabetically. The number chosen is exactly the number of novel types emitted by retrieval for that case.
This marginal-addition prior uses no development response to construct its ranking, and it has the same novel-token budget as retrieval on every case. However, it borrows that budget from retrieval and produces only a token set, not a coherent response. It is therefore a diagnostic control rather than a standalone challenge forecast. The comparison was devised after development labels had already been inspected and is explicitly post hoc. Lexical tokens also include function words and other language that need not be semantic identity signifiers.
5.7 Locked training-only source-conditioned ranking
A final diagnostic tests individualization without using the development cases again. Its plan was fixed before implementation and scoring at SHA-256 e65f04fd9e55843db8ff1a1bb0544acb00ec869eb58f9b312a63d531fc5f4ec9. This is an analysis lock, not a preregistration: the training pairs had already informed earlier methods.
Each of the 150 training cases is held out in turn. On the other 149, the analysis counts marginal Add events and every association between a token in the source and a different token added at follow-up. For a candidate addition, the source-conditioned score is its largest conditional Add rate across the held-out person’s source tokens, smoothed toward the candidate’s marginal Add rate with ten prior cases. That prior strength was fixed to match the stable projection’s existing smoothing choice. Ties favor marginal frequency and then alphabetical token order.
Both rankings receive exactly as many guesses as the held-out follow-up actually adds. This oracle budget isolates token ordering but disqualifies both rankings as standalone forecasts. Because predicted and observed token sets have equal size, case-level precision, recall, and F1 are identical; the report calls their shared value the recovered fraction. A 20,000-resample paired case bootstrap uses seed 20260920.
5.8 Locked regularized-neighborhood ranking
A second training-only plan was fixed before implementation or scoring at SHA-256 c7a494327dd1dd7fc83911540a6a9095fcc8d684172a7388bd610206ac4ca13d. It tests whether pooling multiple similar trajectories avoids the sparse-cue problem of the source-token rule. The source representation is reused from the earlier retrieval analysis: fold-fit TF-IDF over source text and field-qualified 2024 demographics.
For each held-out case, the other 149 cases form the training fold. The method selects the 30 most cosine-similar sources, weights their Add events by similarity, and rescales the weights to 30 effective cases. Each candidate’s weighted neighborhood count is then combined with 30 equivalent cases at its fold-wide marginal Add rate. This one-to-one shrinkage and the neighborhood size were fixed rather than tuned. The comparator is the same leave-one-out marginal ranking; both receive the held-out case’s observed addition count as an oracle budget. A 20,000-resample paired case bootstrap uses seed 20260921. The plan is an analysis lock, not a preregistration, because both the training data and source representation informed earlier work.
5.9 Locked prospective change-volume forecasting
A third training-only plan was fixed before retrieving the benchmark in this iteration, implementation, or scoring at SHA-256 53c974fcd37f443ea846e88328a125169265fb41ac1cc7877529bdbf9a09638f. It reuses the same 30-neighbor text-and-demographic representation and equal 30-case neighborhood/prior weights, but removes the oracle future-token budget. The question is now whether similarity helps forecast response form and the number of Add and Delete events.
For each held-out case, the other 149 provide five modeled quantities: Add count; Delete count as a fraction of the source’s distinct token count; follow-up-minus-source word and nonempty-line counts; and follow-up-to-source ROUGE-L. A source-calibrated marginal forecast uses the fold median. The neighborhood forecast uses a weighted median combining 30 similarity-weighted neighbors with a 30-case prior spread across the whole fold. Fractions and changes are transformed back using only the held-out 2024 response. The analysis therefore uses no follow-up-derived budget or other held-out future information for prediction.
The primary effect is marginal absolute error minus neighborhood absolute error, so positive values favor person conditioning. Twenty thousand paired case-bootstrap resamples use seed 20260923, restarted for each outcome. This lock cannot undo earlier use of the training corpus or earlier selection of the representation and hyperparameters.
5.10 Locked probabilistic change-volume forecasting
A fourth training-only plan extends that fixed volume model from point forecasts to complete predictive distributions. It was locked before reopening the benchmark in this iteration, implementation, or scoring at SHA-256 00d5e2805082986b00f2716808ee12d4cff70ec900a5e47636e893217368f020. For each held-out case and outcome, the marginal distribution gives equal weight to all 149 fold observations. The neighborhood distribution combines a 30-case fold-wide prior with the same 30 cosine-weighted neighbors used above. Every support point is transformed to the held-out source scale before scoring.
The primary score is the continuous ranked probability score (CRPS), a proper score that rewards predictive distributions concentrated near the observation (Gneiting & Raftery, 2007); lower values are better. Secondary diagnostics use central 80% intervals and report inclusive coverage, width, and interval score, which penalizes both width and misses. Positive paired effects are marginal score minus neighborhood score. Twenty thousand paired case-bootstrap resamples use seed 20260924, restarted for each outcome and score. Representation, neighborhood size, and shrinkage remain unchanged and untuned.
5.11 Locked source-feature ablation
A fifth training-only plan was locked before reopening the benchmark in this iteration, implementation, or scoring at SHA-256 ed9fa67b2cc1265f3710471f96d7ffbfd16986b8295ef5072da26c263d8b7da4. It asks whether the previously observed word-count distribution gain comes from the earlier self-description, the ten 2024 demographic fields, or their combination. It splits the unchanged TF-IDF representation into text-only, demographics-only, and combined feature sets without changing any field, neighbor, weighting, shrinkage, transformation, or scoring rule.
Follow-up word-count CRPS is primary because it was the only outcome with a supported combined-neighborhood CRPS gain in the preceding lock. The four other outcomes are prespecified secondary diagnostics. Four paired contrasts compare each ablation with the marginal distribution and the combined model. Twenty thousand case-bootstrap resamples use seed 20260925, restarted for each outcome and contrast. Before any ablation result is accepted, the script must numerically reproduce every marginal and combined case-level CRPS value in the preceding locked audit.
5.12 Locked source-form comparison
A sixth training-only plan was locked before reading the benchmark rows in this iteration, implementation, or scoring at SHA-256 5ffaf6b8289a360926da3ae3b1398a336949e698c6157f8ba94ef316ec63c5ad. It compares the text-only distribution with a count-only source-form neighborhood. Each source is represented by log word-token count, log distinct-token count, and log line count. Fold-fit population standard deviations scale Euclidean distance; similarity is 1 / (1 + distance).
The neighborhood size, weights, prior, outcome transformations, and CRPS scoring remain fixed. Follow-up word-count CRPS is primary, and source-form CRPS minus text-only CRPS is the primary paired contrast. Twenty thousand case-bootstrap resamples use seed 20260926. The script must reproduce every preceding marginal and text-only case score before accepting a result. A text advantage would identify lexical composition—not necessarily semantic identity—because TF-IDF can encode style and template wording.
5.13 Locked cross-analysis multiplicity stress test
A seventh plan was locked after the constituent aggregate results were known but before any cross-analysis resampling at SHA-256 d3c31e1bd77df8da0d5b7017438f2b9ff04ba5f39c4dfcb803c1c38f294fb9ac. It fixes one family of 16 unique no-oracle contrasts: marginal-minus- neighborhood error for five point outcomes, marginal-minus-neighborhood CRPS for the same five outcomes, and the six prespecified word-count CRPS contrasts from the feature and source-form ablations. Duplicated inherited contrasts are counted once. Central-interval scores, secondary ablation outcomes, oracle- budget token rankings, full-text scorecards, and development analyses are outside this deliberately bounded family.
One synchronized case bootstrap resamples the 150 case indices 20,000 times with seed 20260929, preserving the empirical dependence among all 16 stored case-effect vectors. Each resample contributes its maximum absolute studentized mean deviation. The nearest-rank 95th percentile of those maxima is the common critical value for two-sided simultaneous intervals. This is a post hoc sensitivity audit, not retroactive preregistration or formal population familywise control: stored leave-one-out scores are resampled without refitting their heavily overlapping folds.
5.14 Locked calibrated full-text synthesis
An eighth training-only plan was locked before reading benchmark rows in this iteration, implementation, or scoring at SHA-256 ae0594e0f95f4131416be0bab4c3bfa913010dc6380d65c39966849c3a9f7253. It asks whether the Project’s separately observed signals can be composed into one prospective text forecast. Each outer fold ranks held-out source units with the frozen stable-token rule, forecasts follow-up word count with the fixed three-count source-form neighborhood, forecasts distinct Add count with the fixed text-and-demographic neighborhood, and ranks complete follow-up units that are exact-unit additions in at least two fold cases.
The generator searches every nonempty prefix of ranked stable source units and every prefix of ranked common additions. It minimizes the sum of normalized absolute error from the predicted word and Add counts, then applies fixed source-retention and parsimony tie rules. All fitting and candidate text come from the other 149 cases. A volume-matched marginal token ranking receives exactly as many novel-token guesses as the synthesized response, so its content comparison uses no held-out future budget.
Twenty thousand paired case-bootstrap resamples use seed 20260930, restarted for every contrast. Development prediction is allowed only if three pointwise 95% intervals all exclude zero in the favorable direction: normalized edit similarity versus repeat-2024, word-count error versus repeat-2024, and novel- type F1 versus the volume-matched marginal ranking. This conservative gate is not a multiplicity correction. The design is dependent on earlier results and the same training cohort, so even a passing result would not be untouched confirmation.
5.15 Locked semantic-neighborhood ranking
A ninth training-only plan was locked before reading benchmark rows in this iteration, implementation, or scoring at SHA-256 6a6a536a11c552b0e75f40ce6d79100902727958fa908cb57a662ac85df5fdcf. It tests whether external distributional semantics improves the selection of novel token types. The fixed representation uses the public-domain, 25-dimensional uncased GloVe Twitter vectors trained on 2 billion tweets and 27 billion tokens (Pennington et al., 2014). The exact Gensim-data conversion is pinned at SHA-256 63877d71151688baf6f31d5437374f637f737a5e100e12150a5bd61a9f273c3f; the 104 MB external artifact is streamed for analysis but not committed.
Within each 149-case fold, source responses become centroids of their distinct in-vocabulary token vectors, weighted by fold-fit inverse document frequency. Cosine similarity selects 30 neighbors. Their weighted Add counts contribute 30 effective cases and are shrunk equally toward fold-wide marginal Add rates, exactly matching the earlier regularization rule. The marginal ranking and the inherited surface-text-plus-demographic neighborhood receive the same oracle held-out addition budget. The script must reproduce every inherited marginal and surface-neighborhood hit count before accepting the semantic result.
The primary effect is semantic-neighborhood minus marginal recovered fraction; the secondary effect compares semantic and surface neighborhoods. Twenty thousand paired case-bootstrap resamples use seed 20261001, restarted for each contrast. A semantic advantage is claimable only if the primary mean and interval are positive. Regardless of outcome, this iteration cannot reopen development data.
6 Results
6.1 Aggregate scorecard
The authoritative shared evaluator reports fifteen measures in three groups. The table reports every measure; arrows indicate the evaluator’s direction where one exists.
| Group | Metric | Stable projection | Trajectory retrieval | Repeat 2024 | Direction |
|---|---|---|---|---|---|
| Agreement | Normalized exact-match rate | 0.000000 | 0.000000 | 0.000000 | Higher |
| Agreement | Normalized edit similarity | 0.298061 | 0.253360 | 0.291966 | Higher |
| Agreement | Token Jaccard similarity | 0.142768 | 0.080427 | 0.141930 | Higher |
| Agreement | Token-overlap F1 | 0.307444 | 0.194693 | 0.312202 | Higher |
| Agreement | ROUGE-L F1 | 0.227552 | 0.150924 | 0.225683 | Higher |
| Agreement | Character n-gram F1 | 0.292293 | 0.212660 | 0.296534 | Higher |
| Form | Word-count MAE | 41.640000 | 59.640000 | 56.680000 | Lower |
| Form | Line-count MAE | 9.760000 | 8.340000 | 6.800000 | Lower |
| Form | Mean predicted word count | 77.320000 | 99.280000 | 108.040000 | Descriptive |
| Form | Mean reference word count | 94.760000 | 94.760000 | 94.760000 | Descriptive |
| Change | Prediction repeat-2024 rate | 0.280000 | 0.000000 | 1.000000 | Descriptive |
| Change | Observed repeat-2024 rate | 0.000000 | 0.000000 | 0.000000 | Descriptive |
| Change | Source-similarity MAE | 0.678215 | 0.136074 | 0.774317 | Lower |
| Change | Mean prediction-to-source similarity | 0.903898 | 0.175191 | 1.000000 | Descriptive |
| Change | Mean observed follow-up-to-source similarity | 0.225683 | 0.225683 | 0.225683 | Descriptive |
The projection improves three of six agreement measures, worsens two, and ties on exact match. Its clearest gains concern form and amount of change: word-count MAE falls by 15.04 words, and source-similarity MAE falls by 0.096102. But the mean predicted source similarity remains 0.678215 above the observed mean. Extracting supposedly enduring signifiers does not come close to reproducing how radically people rewrite their self-descriptions.
Trajectory retrieval creates the opposite pattern. Its mean source similarity (0.175191) is close to the observed 0.225683, so source-similarity MAE falls to 0.136074. But every non-exact agreement metric is lower than both continuity baselines. Generating a plausibly different response is therefore not equivalent to predicting the content of the focal person’s future self-description.
6.2 Paired development-case pattern
Positive effects favor stable-signifier projection. For agreement measures, the effect is projection minus repeat-2024; for errors, it is repeat-2024 error minus projection error. Exact match is omitted because all 50 cases tie at zero. A win, tie, or loss is evaluated within a person before averaging.
| Metric | Mean effect | Paired case-bootstrap 95% interval | Win / tie / loss |
|---|---|---|---|
| Normalized edit similarity | +0.006095 | −0.001701 to +0.013890 | 24 / 14 / 12 |
| Token Jaccard similarity | +0.000838 | −0.002383 to +0.004097 | 16 / 20 / 14 |
| Token-overlap F1 | −0.004759 | −0.012300 to +0.002429 | 13 / 20 / 17 |
| ROUGE-L F1 | +0.001869 | −0.002268 to +0.005755 | 19 / 20 / 11 |
| Character n-gram F1 | −0.004241 | −0.012240 to +0.003618 | 16 / 14 / 20 |
| Word-count error reduction | +15.040000 | +1.180000 to +32.360000 | 20 / 20 / 10 |
| Line-count error reduction | −2.960000 | −5.540000 to −0.420000 | 15 / 12 / 23 |
| Source-similarity error reduction | +0.096102 | +0.066064 to +0.129348 | 30 / 20 / 0 |
The table changes the emphasis of the aggregate scorecard. None of the small agreement differences has an interval that excludes zero. In contrast, the projection’s lower word-count and source-similarity errors persist across the paired case resamples, while its line-count error is consistently worse. The source-similarity result is directional but inadequate in magnitude: even after improvement, the projection still remains much too close to the past.
6.3 Training cross-validation repeats the mixed pattern
The locked leave-one-out analysis refits the frozen projection 150 times. As in the development analysis, positive paired effects favor the projection; for error measures they are repeat-2024 error minus projection error.
| Metric | Projection | Repeat 2024 | Mean effect | Paired case-bootstrap 95% interval |
|---|---|---|---|---|
| Normalized exact match | 0.000000 | 0.000000 | 0.000000 | 0.000000 to 0.000000 |
| Normalized edit similarity | 0.276809 | 0.272662 | +0.004147 | +0.000652 to +0.007753 |
| Token Jaccard similarity | 0.133873 | 0.134297 | −0.000424 | −0.002229 to +0.001237 |
| Token-overlap F1 | 0.276737 | 0.280231 | −0.003493 | −0.007228 to −0.000077 |
| ROUGE-L F1 | 0.200304 | 0.200017 | +0.000287 | −0.002003 to +0.002538 |
| Character n-gram F1 | 0.266194 | 0.267813 | −0.001620 | −0.005530 to +0.002270 |
| Word-count MAE | 46.986667 | 55.580000 | +8.593333 | +2.573333 to +15.266667 |
| Line-count MAE | 9.066667 | 6.160000 | −2.906667 | −4.346833 to −1.480000 |
| Source-similarity MAE | 0.729214 | 0.799983 | +0.070769 | +0.054503 to +0.088484 |
Only two nonzero agreement intervals exclude zero, and they point in opposite directions: normalized edit similarity improves slightly while token-overlap F1 worsens slightly. The other agreement differences remain unstable to cohort composition. The form/change findings are more reproducible: the method lowers word-count and source-similarity error but raises line-count error. Magnitude remains the substantive problem. Mean predicted source similarity is 0.929231, versus 0.200017 for observed follow-ups, and 46.7% of predictions repeat the source exactly while no observed follow-up does.
The direction of all three error results matches development; four of five non-exact agreement directions also match, with token Jaccard shifting from a tiny positive difference to a tiny negative one. This is useful recurrence within the same method-development corpus, not independent replication.
6.4 Lexical novelty and extractive ceilings
The observed follow-ups contain substantial new lexical material. On an average case, 73.1727% of distinct follow-up token types are absent from the earlier response (95% case-bootstrap interval: 69.7982%–76.5607%). When token frequencies are respected, 65.0621% of follow-up token occurrences cannot be copied from the source without reusing tokens beyond their source counts (59.2663%–70.6149%). Conversely, only 24.9892% of distinct source token types recur at follow-up (22.0060%–28.0466%).
Even an oracle with access to the observed future is constrained. Its mean extractive ceilings are 0.268273 for unique-token Jaccard, 0.484206 for bag-of-words F1, and 0.374081 for source-order subsequence ROUGE-L F1. The actual projection reaches 0.142768, 0.307444, and 0.227552 on those measures. These comparisons reveal two failures at once: the current extractor does not reach the retrospective extractive ceiling, and no extractor can generate the many identity signifiers expressed only at follow-up.
6.5 Retrieval predicts novelty volume, not novel content
Positive paired effects favor trajectory retrieval. For errors, the effect is stable-projection error minus retrieval error. Retrieval loses agreement and word-count accuracy, modestly improves line-count error with an interval that spans zero, and dramatically improves source-similarity error.
| Metric | Retrieval | Stable projection | Paired effect | Paired case-bootstrap 95% interval |
|---|---|---|---|---|
| Normalized exact match | 0.000000 | 0.000000 | 0.000000 | 0.000000 to 0.000000 |
| Normalized edit similarity | 0.253360 | 0.298061 | −0.044701 | −0.072824 to −0.018705 |
| Token Jaccard similarity | 0.080427 | 0.142768 | −0.062342 | −0.081759 to −0.043776 |
| Token-overlap F1 | 0.194693 | 0.307444 | −0.112751 | −0.170815 to −0.058175 |
| ROUGE-L F1 | 0.150924 | 0.227552 | −0.076628 | −0.126464 to −0.030063 |
| Character n-gram F1 | 0.212660 | 0.292293 | −0.079633 | −0.112019 to −0.049794 |
| Word-count MAE | 59.640000 | 41.640000 | −18.000000 | −30.120500 to −6.160000 |
| Line-count MAE | 8.340000 | 9.760000 | +1.420000 | −1.860000 to +4.640000 |
| Source-similarity MAE | 0.136074 | 0.678215 | +0.542141 | +0.463332 to +0.616781 |
The locked novelty criterion tells the same story more precisely. Retrieval’s mean predicted unique-token novelty (0.801727) is much closer to the observed mean (0.731727) than either stable projection or repeat-2024 (0 for both), so it passes the plan’s novelty-volume criterion. Its case-level novelty MAE is 0.139032. Occurrence novelty is also close in aggregate (0.727677 predicted versus 0.650621 observed).
However, novelty volume is not novel-content recovery. Among token types that retrieval introduces beyond the focal source, mean precision against the observed new token types is only 0.070466 (case-bootstrap interval 0.052133–0.090859), recall is 0.085362 (0.065937–0.106057), and F1 is 0.067109 (0.052917–0.082272). Retrieval knows that a future response will look different, but mostly borrows the wrong person’s differences.
6.6 Person matching does not beat common additions
The volume-matched marginal prior predicts a mean 46.94 novel token types per case, exactly equal to retrieval by construction. Despite lacking any person-matching rule, it recovers more of the observed novel vocabulary:
| Novel-type measure | Marginal prior | Trajectory retrieval | Paired difference | Paired case-bootstrap 95% interval |
|---|---|---|---|---|
| Precision | 0.176545 | 0.070466 | +0.106079 | +0.068782 to +0.154891 |
| Recall | 0.172611 | 0.085362 | +0.087249 | +0.062210 to +0.113649 |
| F1 | 0.142637 | 0.067109 | +0.075528 | +0.055927 to +0.094990 |
Positive differences favor the marginal prior. It wins 40 cases, ties 6, and loses 4 on each measure. The common additions include syntactic vocabulary such as an, in, my, is, and who, but also more content-bearing words such as friend, lover, kind, creative, and life. This means the earlier retrieval F1 cannot be read as evidence that matching text and demographics identified person-specific new signifiers. At matched volume, a coarse cohort-level base rate is substantially stronger.
The result does not establish that novel identity content is inherently unpredictable. It establishes a stricter baseline for future work: an individualized novelty model should beat common training additions at a comparable prediction volume before its gains are attributed to person-level conditioning.
6.7 Source-token conditioning also loses to common additions
The fixed source-conditioned ranking does not clear that baseline in its training-only leave-one-out test. Across all 150 held-out training cases, the mean recovered fraction is 0.145238, compared with 0.160992 for the marginal-addition ranking:
| Training leave-one-out result | Source-conditioned | Marginal additions | Paired difference |
|---|---|---|---|
| Mean novel-type recovered fraction | 0.145238 | 0.160992 | −0.015754 |
| Median novel-type recovered fraction | 0.141177 | 0.156250 | — |
| Case wins / ties / losses | 19 / 63 / 68 | — | — |
| Paired case-bootstrap 95% interval | — | — | −0.020829 to −0.010736 |
The interval excludes zero in the direction favoring the marginal ranking. Thus, under the locked interpretation rule, the source-conditioned method provides no incremental advantage. One plausible explanation is that choosing the strongest source-token cue promotes sparse, idiosyncratic pairs that fail to recur in the held-out person, while the marginal ranking spends more of its fixed budget on common additions; the analysis does not isolate that mechanism.
This is a bounded negative result, not evidence that source signifiers contain no predictive information. The rule tests only token pairs, not phrases, semantics, interactions, or predicted deletions. Its mean candidate-vocabulary ceiling is 0.773139: about 22.7% of held-out additions are absent from every other training follow-up’s additions and cannot be selected by either ranking. Most importantly, the oracle budget comes from the held-out future. The result compares ranking quality within this training cohort; it is not development or private-test performance.
6.8 Regularized neighborhoods still lose to common additions
Pooling rather than maximizing person-conditioned evidence narrows the deficit but does not clear the marginal baseline:
| Training leave-one-out result | Regularized neighborhood | Marginal additions | Paired difference |
|---|---|---|---|
| Mean novel-type recovered fraction | 0.149007 | 0.160992 | −0.011984 |
| Median novel-type recovered fraction | 0.148542 | 0.156250 | — |
| Case wins / ties / losses | 25 / 60 / 65 | — | — |
| Paired case-bootstrap 95% interval | — | — | −0.018426 to −0.005693 |
The interval again excludes zero in the direction favoring common additions. The neighborhood and marginal top sets overlap by 0.739706 on average, so the personalized method changes about 26.0% of guesses rather than merely reproducing its comparator. Mean cosine similarity across the 30 selected neighbors is only 0.169564, consistent with a diffuse local neighborhood, although this diagnostic cannot determine why its substitutions are worse.
The result strengthens the bounded inference across two specifications. A maximum source-token association loses by 1.58 percentage points; a pooled, strongly regularized text-and-demographic neighborhood loses by 1.20 points. It does not follow that individualized prediction is impossible. Both methods operate on surface lexical overlap, both use an oracle future-token budget, and both draw candidates from other cases’ additions. The shared candidate-vocabulary ceiling remains 0.773139.
6.9 Distributional semantics also loses to common additions
The pretrained representation covers 2409 of 2537 distinct source token types (94.9547%), and every source has at least one vector. High coverage does not translate into better Add selection:
| Training leave-one-out result | Semantic neighborhood | Surface neighborhood | Marginal additions | Semantic minus marginal |
|---|---|---|---|---|
| Mean novel-type recovered fraction | 0.148760 | 0.149007 | 0.160992 | −0.012232 |
| Median novel-type recovered fraction | 0.153846 | 0.148542 | 0.156250 | — |
| Semantic wins / ties / losses versus marginal | 22 / 65 / 63 | — | — | — |
| Paired case-bootstrap 95% interval | — | — | — | −0.018361 to −0.006322 |
The primary interval excludes zero in the direction favoring common additions, so the semantic hurdle fails. Semantic and marginal top sets overlap by 0.746530 on average: semantic conditioning changes about a quarter of the guesses but makes them worse. Its difference from the inherited surface neighborhood is only −0.000247, with an interval spanning zero (−0.006521 to +0.005834); the two representations are not distinguished by this test.
Mean cosine similarity among selected semantic neighbors is 0.975307, but that high value partly reflects dense centroid geometry and is not an identity similarity scale. The fixed centroid discards word order, senses, negation, and line structure. The result therefore rejects this representation and ranking rule, not semantic conditioning generally. No development prediction was generated.
6.10 Neighborhoods predict some revision volume
When the oracle addition budget is removed, the same neighborhood has a more limited but positive role. It improves absolute-error forecasts for how many distinct token types are added and deleted and for follow-up word count:
| Training leave-one-out outcome | Neighborhood MAE | Source-calibrated marginal MAE | Error reduction | Paired case-bootstrap 95% interval |
|---|---|---|---|---|
| Add count | 20.126667 | 21.340000 | +1.213333 | +0.680000 to +1.740000 |
| Delete count | 5.336275 | 5.614935 | +0.278660 | +0.022274 to +0.545876 |
| Follow-up word count | 52.940000 | 56.080000 | +3.140000 | +1.093333 to +5.213333 |
| Follow-up line count | 6.160000 | 6.160000 | 0.000000 | 0.000000 to 0.000000 |
| Source similarity | 0.106781 | 0.109648 | +0.002867 | −0.000960 to +0.006730 |
Positive reductions favor the regularized neighborhood. Its Add-count forecast wins 79 cases, ties 17, and loses 54; Delete count wins 88, ties 5, and loses 57; word count wins 82, ties 13, and loses 55. The line-count forecasts are identical for all cases because both weighted medians select the same modeled line change. The source-similarity point estimate favors neighborhoods, but its interval spans zero.
The magnitudes are modest: Add-count MAE falls by 5.7%, Delete-count MAE by 5.0%, and word-count MAE by 5.6% relative to their comparators. Still, the direction matters. Surface text and coarse demographics contain some case-specific information about how much personally expressed identity will be revised, even though the same representation makes worse choices about which new tokens will appear. These are volume and form forecasts within the selected training cohort, not full-text forecasts or evidence that demographic similarity is an ipseological mechanism.
6.11 Neighborhood distributions improve only word-count forecasts
Scoring complete predictive distributions narrows the point-forecast result. Only follow-up word count has a paired CRPS interval that excludes zero in the direction favoring the neighborhood:
| Training leave-one-out outcome | Neighborhood CRPS | Source-calibrated marginal CRPS | CRPS reduction | Paired case-bootstrap 95% interval |
|---|---|---|---|---|
| Add count | 14.438038 | 14.650556 | +0.212518 | −0.071030 to +0.497525 |
| Delete count | 3.853656 | 3.926457 | +0.072801 | −0.036356 to +0.178244 |
| Follow-up word count | 38.156852 | 40.672695 | +2.515843 | +1.008212 to +4.059792 |
| Follow-up line count | 4.922609 | 4.910848 | −0.011761 | −0.097952 to +0.081212 |
| Source similarity | 0.074684 | 0.077171 | +0.002486 | −0.000106 to +0.005180 |
The word-count neighborhood wins 83 cases and loses 67 on CRPS. Its central 80% interval covers 84.7% of cases, compared with 80.7% for the marginal distribution, while mean width rises from 149.960000 to 154.526667 words. Despite that modest widening, mean interval score improves from 264.626667 to 244.926667: the paired reduction is +19.700000 with interval +3.873167 to +37.193333. Source-similarity interval score also improves, but its primary CRPS interval spans zero. The other interval-score comparisons span zero as well.
These distributional scores do not overturn the point-forecast findings. Rather, they locate their strongest probabilistic support in follow-up length. The Add and Delete medians improve absolute error, but their empirical distributions do not establish stable CRPS gains. Line-count intervals are substantially over-covering (96.7% marginal and 92.7% neighborhood against 80% nominal), showing that apparent uncertainty can be broad without being informative. None of these quantities identifies the future signifiers.
6.12 Earlier text outperforms demographics for word-count skill
The required replication passed for all 150 cases and five outcomes. For the primary follow-up word-count outcome, text-only matching preserves most of the combined neighborhood’s advantage, while demographics-only matching does not establish an improvement over the marginal distribution:
| Word-count predictive distribution | Mean CRPS | Locked paired contrast | CRPS reduction | Paired case-bootstrap 95% interval |
|---|---|---|---|---|
| Source-calibrated marginal | 40.672695 | — | — | — |
| Text only | 38.374556 | Marginal minus text only | +2.298139 | +0.816444 to +3.865823 |
| Demographics only | 40.471984 | Marginal minus demographics only | +0.200711 | −0.479832 to +0.861529 |
| Combined | 38.156852 | Text only minus combined | +0.217704 | +0.024229 to +0.420871 |
Text-only matching wins 84 cases and loses 66 against the marginal. The combined model wins 85 and loses 65 against text-only. Thus the earlier self-description carries nearly all of the supported response-length signal; the ten coarse demographic fields alone do not beat the source-calibrated marginal. Their small incremental benefit conditional on text is compatible with weak complementary information, not a stand-alone demographic mechanism.
The secondary diagnostics do not support a general feature-family conclusion. For Add count, Delete count, and line count, every fixed ablation contrast has an interval spanning zero. Text-only matching improves source-similarity CRPS over the marginal by +0.002964, but its repeated-analysis interval barely excludes zero (+0.000051 to +0.006065), and adding demographics to text does not improve that outcome. The primary result remains about response length, not future signifiers or revision volume generally.
6.13 Simple source form matches text-only word-count skill
The required replication again passed for all 150 cases and five outcomes. Matching on only the earlier response’s word-token, distinct-token, and line counts improves the primary word-count forecast over the source-calibrated marginal:
| Word-count predictive distribution | Mean CRPS | Locked paired contrast | CRPS reduction | Paired case-bootstrap 95% interval |
|---|---|---|---|---|
| Source-calibrated marginal | 40.672695 | — | — | — |
| Text only | 38.374556 | Marginal minus text only | +2.298139 | +0.826598 to +3.834257 |
| Source form only | 37.427386 | Marginal minus source form | +3.245309 | +1.839350 to +4.784335 |
| Direct comparison | — | Source form minus text only | −0.947170 | −1.943737 to +0.035874 |
The source-form neighborhood wins 96 cases and loses 54 against the marginal. Text-only matching wins 65 and loses 85 in the direct comparison, but the prespecified primary interval narrowly includes zero. Under the locked interpretation rule, this analysis does not separate the two methods’ primary skill. It therefore supplies no stable evidence that lexical composition adds word-count information beyond three simple counts. Nor does the lower mean for source form prove it is generally superior: the direct interval also prevents that claim.
The secondary pattern reinforces the narrow measurement interpretation. Source-form matching improves line-count CRPS by +0.445157 over the marginal (interval +0.206291 to +0.701574) and by +0.472157 over text only (+0.235305 to +0.721005). Add-count and Delete-count comparisons span zero. Source-form matching’s source-similarity reduction is +0.001969, with a repeated-analysis interval barely above zero (+0.000017 to +0.003880). These unadjusted secondary diagnostics concern response form, not future identity content.
6.14 Six directions survive the multiplicity stress test
The synchronized bootstrap’s common studentized critical value is 2.959538, larger than a contrast-by-contrast normal critical value. Six of the 16 fixed directions have simultaneous 95% intervals excluding zero:
| Fixed contrast | Mean effect | Simultaneous 95% interval | Direction within this audit |
|---|---|---|---|
| Add-count point-error reduction | +1.213333 | +0.397788 to +2.028879 | Neighborhood |
| Delete-count point-error reduction | +0.278660 | −0.116978 to +0.674298 | Inconclusive |
| Word-count point-error reduction | +3.140000 | +0.047580 to +6.232420 | Neighborhood |
| Line-count point-error reduction | 0.000000 | 0.000000 to 0.000000 | Exact tie |
| Source-similarity point-error reduction | +0.002867 | −0.002885 to +0.008619 | Inconclusive |
| Add-count CRPS reduction | +0.212518 | −0.217903 to +0.642939 | Inconclusive |
| Delete-count CRPS reduction | +0.072801 | −0.089750 to +0.235353 | Inconclusive |
| Word-count CRPS reduction | +2.515843 | +0.223750 to +4.807936 | Neighborhood |
| Line-count CRPS reduction | −0.011761 | −0.148523 to +0.125000 | Inconclusive |
| Source-similarity CRPS reduction | +0.002486 | −0.001516 to +0.006488 | Inconclusive |
| Marginal minus text-only word-count CRPS | +2.298139 | +0.023968 to +4.572310 | Text only |
| Marginal minus demographics-only word-count CRPS | +0.200711 | −0.811325 to +1.212747 | Inconclusive |
| Text-only minus combined word-count CRPS | +0.217704 | −0.084469 to +0.519876 | Inconclusive |
| Demographics-only minus combined word-count CRPS | +2.315132 | +0.011811 to +4.618453 | Combined |
| Marginal minus source-form word-count CRPS | +3.245309 | +1.000674 to +5.489943 | Source form |
| Source-form minus text-only word-count CRPS | −0.947170 | −2.449295 to +0.554956 | Inconclusive |
The adjustment preserves the central response-length interpretation. Both the point and probabilistic word-count gains remain directionally stable, as do text-only and source-form improvements over the marginal. The Add-count point gain also remains. In contrast, the smaller Delete-count point effect and the increment from adding demographics to text no longer exclude zero. Thus the strongest person-matching evidence concerns Add volume and later response length; the earlier wording that all three point outcomes improved should be read as the pointwise pattern, not as a family-robust conclusion.
This audit does not make the six remaining directions confirmatory. The family was constructed after earlier results were known; the same 150 selected cases support every contrast; fitted folds overlap; and resampling stored case scores does not refit the models. It is a conservative internal coherence check, not population inference or a remedy for sequential analysis.
6.15 Calibrated synthesis does not clear the content gate
The full-text synthesis changes the kind of error without producing a general agreement gain. It is much closer to the observed amount of source change and has lower word-count error than repeat-2024, but the common appended units substantially inflate the number of lines and dilute token agreement.
| Training leave-one-out metric | Calibrated synthesis | Stable projection | Repeat 2024 | Synthesis utility effect versus repeat (95% interval) |
|---|---|---|---|---|
| Exact match | 0.000000 | 0.000000 | 0.000000 | 0.000000 (0.000000 to 0.000000) |
| Normalized edit similarity | 0.268883 | 0.276809 | 0.272662 | −0.003779 (−0.013404 to +0.004287) |
| Token Jaccard | 0.098165 | 0.133873 | 0.134297 | −0.036132 (−0.048502 to −0.024014) |
| Token-overlap F1 | 0.267251 | 0.276737 | 0.280231 | −0.012980 (−0.035902 to +0.010370) |
| ROUGE-L F1 | 0.220153 | 0.200304 | 0.200017 | +0.020136 (+0.002677 to +0.037571) |
| Character n-gram F1 | 0.238779 | 0.266194 | 0.267813 | −0.029034 (−0.040378 to −0.017567) |
| Word-count MAE | 51.986667 | 46.986667 | 55.580000 | +3.593333 (+1.440000 to +5.726667) |
| Line-count MAE | 27.053333 | 9.066667 | 6.160000 | −20.893333 (−23.820000 to −17.966667) |
| Source-similarity MAE | 0.143463 | 0.729214 | 0.799983 | +0.656521 (+0.627361 to +0.684716) |
Positive effects favor synthesis for every row. The ROUGE-L and source- similarity results show that the generator approximates the degree and sequence of change better than continuity baselines; they do not show correct new identity content. Relative to the stable projection, synthesis also improves ROUGE-L by +0.019849 and source-similarity error by +0.585751, but worsens token Jaccard, character n-gram F1, and line-count error with intervals excluding zero. Its word-count error is five words higher on average than the stable projection, with an interval spanning zero.
The equal-volume content comparison is unambiguously unfavorable:
| Distinct novel-type measure | Calibrated synthesis | Volume-matched marginal | Paired difference (95% interval) |
|---|---|---|---|
| Precision | 0.097449 | 0.208189 | −0.110740 (−0.139732 to −0.082632) |
| Recall | 0.074294 | 0.116041 | −0.041748 (−0.058335 to −0.025062) |
| F1 | 0.075887 | 0.134680 | −0.058793 (−0.076168 to −0.041473) |
Synthesis wins 33 cases, ties 21, and loses 96 on novel-type F1. Only the word- count gate passes; the edit-similarity interval spans zero and the content gate excludes zero in the wrong direction. The prespecified advancement rule therefore fails, and no development prediction was generated.
6.16 Frozen test submission
The method generated one nonblank prediction for every test ID in the required order. The pinned official validator returned VALID: 81 predictions. The frozen CSV has SHA-256:
a463d9e314069357f050c9f2270acfad59165db2c0bab19322d46517444d9ab3
The organizer must run the private test evaluator. Until that happens, the private-test performance remains unknown; neither public-development results nor the training-only ranking diagnostic substitutes for it.
7 Discussion
The stable-signifier projection does learn something useful about response form and degree of change. Yet improved prediction of how much a response will change is not the same as predicting what new identity content will appear. The novelty diagnostic makes that distinction measurable: most future lexical material lies outside the earlier response, while the method is definitionally restricted to it. A next-generation model needs a mechanism for both deletion and creation, evaluated without sacrificing the interpretability of these paired diagnostics.
Leave-one-out cross-validation makes this conclusion less dependent on the 50 reused development cases. Its larger training-cohort analysis reproduces all three directional error results and the overall lack of a uniform agreement gain. The projection can adjust length and continuity without solving content: its small edit-similarity gain coexists with a small token-overlap loss and extreme overprediction of source continuity.
Matched-trajectory retrieval sharpens this inference. It supplies deletion and creation in approximately the observed proportions, but the transferred novel content is rarely person-correct. The problem is not merely calibrating how different the next self-description will be. A useful forecast must condition new signifiers on the individual without collapsing back into verbatim continuity.
The marginal control tightens that conclusion. Retrieval not only recovers little new vocabulary; it recovers less than a person-agnostic training prior given the same number of guesses. The diagnostic also disciplines the language of the report: token recovery is not automatically signifier recovery, because many high-base-rate additions are connective or generic words. Future person-specific models need to demonstrate value beyond both continuity and marginal lexical prevalence.
The two leave-one-out conditioned results make that requirement more specific. Merely finding the strongest smoothed source-token-to-Add association is not enough, and replacing that maximum with a strongly regularized 30-neighbor ensemble still sacrifices 1.20 percentage points relative to marginal prevalence. Surface lexical and coarse demographic similarity have now failed in both single-trajectory and pooled forms. Better individualization will need richer context, a semantic mechanism beyond document-level GloVe centroids, or new longitudinal evidence; it should establish an incremental training-only advantage before reopening the heavily reused development comparison.
The external-semantic test makes that requirement empirical rather than rhetorical. Its pretrained vectors cover almost all source vocabulary and replace exact lexical overlap with distributional proximity, yet the recovered fraction is virtually identical to the surface neighborhood and remains 1.22 percentage points below marginal prevalence. Semantic representation alone is not semantic forecasting: averaging word vectors can blur precisely the specific roles, relationships, and attributes that matter to personally expressed identity. A future attempt needs either richer compositional context or new evidence, not another fixed re-ranking of these Add candidates.
The prospective volume analysis qualifies that negative content result. At the contrast-by-contrast level, the fixed neighborhood improves three no- oracle quantity forecasts even while its oracle-budget lexical ranking loses to marginal additions. The cross-analysis stress test sharpens that claim: Add-count and word-count point gains remain simultaneously stable, while the smaller Delete-count gain becomes inconclusive. Person conditioning therefore is not uniformly uninformative, but its family-robust quantity signal is narrower than the original pointwise pattern. A synthesizing model should preserve that calibration signal while demonstrating content value beyond marginal prevalence.
The calibrated-synthesis test directly attempts that composition and shows why the two requirements must be evaluated together. Common complete response units make predicted text look appropriately different from its source and improve sequence overlap, but they recover fewer correct novel types than a token-frequency prior and create a gross line-count mismatch. Forecasting the amount of novelty is not a license to fill that budget with generic statements. The failed advancement gate prevents another look at development labels and raises the next-method bar from “synthesize something” to “add source-relevant semantic content that beats both continuity and marginal prevalence.”
The probabilistic extension further qualifies the claim: the clearest distributional gain is for follow-up word count. The ablation locates most of that signal in the earlier self-description rather than demographics, and the source-form comparison prevents a stronger lexical interpretation: three simple source-response counts improve on the marginal and are not stably distinguishable from text-only matching. The multiplicity audit retains both text-only and source-form gains over the marginal but not demographics’ small increment over text. Demographics alone do not improve the marginal forecast, and demographic resemblance is not an ipseological mechanism. None of these surface predictors establishes such a mechanism.
7.1 Reproducibility
The complete open materials are:
ANALYSIS_PLAN_CALIBRATED_SYNTHESIS.md— pre-analysis lock and advancement gate for full-text synthesis;ANALYSIS_PLAN_SEMANTIC_NEIGHBORHOOD_ADDITIONS.md— pre-analysis lock for the external-semantic Add-ranking comparison;ANALYSIS_PLAN_CHANGE_DISTRIBUTIONS.md— pre-analysis lock for probabilistic revision-volume forecasts;ANALYSIS_PLAN_FEATURE_ABLATION.md— pre-analysis lock for separating text and demographic source features;ANALYSIS_PLAN_MULTIPLICITY_STRESS_TEST.md— fixed family and simultaneous-interval procedure for the cross-analysis audit;ANALYSIS_PLAN_SOURCE_FORM_ABLATION.md— pre-analysis lock for comparing lexical and count-only source matching;ANALYSIS_PLAN_CHANGE_VOLUME.md— pre-analysis lock for prospective revision-volume forecasts;ANALYSIS_PLAN_NEIGHBORHOOD_ADDITIONS.md— pre-analysis lock for the regularized-neighborhood comparison;ANALYSIS_PLAN_SOURCE_CONDITIONED_ADDITIONS.md— pre-analysis lock for the training-only ranking comparison;ANALYSIS_PLAN_STABLE_PROJECTION_CROSS_VALIDATION.md— pre-analysis lock for leave-one-out validation of the frozen projection;ANALYSIS_PLAN_TRAJECTORY_RETRIEVAL.md— pre-prediction method and evaluation lock;analysis/analyze_dev_diagnostics.py— hash-guarded paired bootstrap and extractive-limit analysis;analysis/analyze_change_distributions.py— hash-guarded leave-one-out predictive-distribution evaluation;analysis/analyze_calibrated_synthesis.py— hash-guarded leave-one-out full-text synthesis and gate evaluation;analysis/analyze_change_volume.py— hash-guarded leave-one-out response-form and revision-volume forecasts;analysis/analyze_feature_ablation.py— hash-guarded leave-one-out source-feature ablation;analysis/analyze_multiplicity_stress_test.py— hash-guarded synchronized max-|t| case-bootstrap audit;analysis/analyze_source_form_ablation.py— hash-guarded lexical-versus-source-form comparison;analysis/analyze_novelty_prior.py— volume-matched marginal-addition diagnostic;analysis/analyze_semantic_neighborhood_additions.py— hash-guarded GloVe-semantic neighborhood comparison;analysis/analyze_neighborhood_additions.py— hash-guarded regularized-neighborhood leave-one-out comparison;analysis/analyze_source_conditioned_additions.py— hash-guarded leave-one-out source-conditioned comparison;analysis/analyze_stable_projection_cross_validation.py— hash-guarded leave-one-out stable-projection evaluation;analysis/analyze_trajectory_retrieval.py— complete paired and novelty comparison under the lock;analysis/stable_signifier_projection.py— standard-library generator;analysis/trajectory_retrieval.py— standard-library matched-trajectory generator;results/stable_signifier_dev_diagnostics.json— complete machine-readable post hoc diagnostic results;results/change_distributions_train_analysis.json— complete machine-readable probabilistic-forecast result;results/change_distributions_train_audit.csv— all 150 cases’ CRPS and central-interval diagnostics;results/calibrated_synthesis_train_analysis.json— complete official-metric, content, and advancement-gate result;results/calibrated_synthesis_train_audit.csv— all 150 cases’ targets, structure, metrics, and content scores;results/calibrated_synthesis_train_predictions.csv— all 150 leave-one-out full-text synthesis predictions;results/change_volume_train_analysis.json— complete machine-readable change-volume result;results/change_volume_train_audit.csv— all 150 cases’ observed quantities, predictions, errors, and paired effects;results/feature_ablation_train_analysis.json— complete machine-readable feature-ablation result;results/feature_ablation_train_audit.csv— all 150 cases’ marginal and ablated CRPS values;results/multiplicity_stress_test_train_analysis.json— complete machine-readable 16-contrast simultaneous-interval result;results/multiplicity_stress_test_train_audit.csv— flat contrast means, standard errors, simultaneous intervals, and classifications;results/source_form_ablation_train_analysis.json— complete machine-readable source-form comparison;results/source_form_ablation_train_audit.csv— all 150 cases’ marginal, source-form, and text-only CRPS values;results/semantic_neighborhood_additions_train_analysis.json— complete semantic-neighborhood comparison, coverage, and gate result;results/semantic_neighborhood_additions_train_audit.csv— all 150 cases’ coverage, hits, recovered fractions, and overlaps;results/novelty_prior_dev_analysis.json— complete marginal-prior comparison and paired uncertainty;results/novelty_prior_token_audit.csv— token ranks, frequencies, prediction counts, and correct-recovery counts;results/neighborhood_additions_train_analysis.json— complete regularized-neighborhood result;results/neighborhood_additions_train_audit.csv— neighbors, budgets, reachability, hits, overlap, and paired case effects;results/source_conditioned_additions_train_analysis.json— complete machine-readable leave-one-out result;results/source_conditioned_additions_train_audit.csv— budgets, reachable additions, hits, scores, and paired case effects;results/stable_signifier_dev_predictions.csv— all 50 public development predictions;results/stable_signifier_dev_scorecard.json— complete machine-readable scorecard;results/stable_projection_train_cv_analysis.json— complete paired training cross-validation result;results/stable_projection_train_cv_predictions.csv— all 150 leave-one-out fold predictions;results/trajectory_retrieval_dev_analysis.json— complete locked paired comparison and novelty diagnostics;results/trajectory_retrieval_dev_audit.csv— development-to-training match IDs and cosine similarities;results/trajectory_retrieval_dev_predictions.csv— all 50 retrieved development forecasts;results/trajectory_retrieval_dev_scorecard.json— complete official retrieval scorecard;submissions/aleph_initial_alpha_submission.csv— frozen 81-case test artifact;submissions/aleph_initial_alpha_method.md— submission-ready method card.
All Project research scripts use only the Python standard library. The challenge’s 17 public tests passed before the stable artifact was generated. Regenerating the stable development and test artifacts reproduced their hashes exactly; the retrieval pipeline separately hash-guards its inputs and fixed analysis plan.
7.2 Sources
- Gneiting, T., & Raftery, A. E. (2007). Strictly proper scoring rules, prediction, and estimation. Journal of the American Statistical Association, 102(477), 359–378. https://doi.org/10.1198/016214506000001437
- Jones, J. J. (2023). Ipseology—A new science of the self.
- Jones, J. J. (2024). Predicting the self with generative AI [Preprint]. SocArXiv. https://doi.org/10.31235/osf.io/eh9sk
- Jones, J. J. (2026). Building the ipseome: Large, free, open, human identity data [Preprint]. arXiv.
- Jones, J. J. (2026). You can predict future selves with AI (or without AI) [Data set and benchmark]. GitHub.
- Pennington, J., Socher, R., & Manning, C. (2014). GloVe: Global vectors for word representation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP) (pp. 1532–1543). Association for Computational Linguistics. https://doi.org/10.3115/v1/D14-1162
- Romano, J. P., & Wolf, M. (2005). Exact and approximate stepdown methods for multiple hypothesis testing. Journal of the American Statistical Association, 100(469), 94–108. https://doi.org/10.1198/016214504000000539
7.3 Limitations and next step
An extractive method cannot predict genuinely new identities, experiences, or reframing. Token recurrence is not equivalent to identity-signifier endurance, and common wording can make a signifier look stable. The line splitter also turns some prose sentences into separate output lines, worsening line-count error. Finally, public-development-guided selection can overfit a 50-case set.
The stable-projection cross-validation withholds each case’s future from its own model fit, but it is not untouched confirmation. The training corpus had already informed the method, and cross-validation folds overlap heavily. Its case-bootstrap intervals describe this selected training cohort rather than a population. The published fold predictions are a derived benchmark-data adaptation under CC BY-NC-SA 4.0.
Trajectory retrieval adds different limitations. It transfers another participant’s public response rather than generating a person-specific future, and demographic similarity does not establish that identity change follows demographic categories. Its within-iteration lock reduces new analytic flexibility but cannot make previously inspected development data unseen.
The marginal-addition control is post hoc and deliberately narrow. It uses retrieval’s case-level novel-token count, so it is not an independent full-text forecast; it isolates content choice after holding opportunity volume equal. Its ranking is estimated from only 150 training pairs, and its lexical units include function words that are not necessarily identity signifiers. The derived token audit adapts the benchmark data and remains governed by the benchmark’s CC BY-NC-SA 4.0 license.
The source-conditioned comparison is locked and uses leave-one-out estimation, but it remains a training-cohort diagnostic with an oracle future-token budget. Its maximum-over-source-tokens rule can privilege sparse associations, the ten-case smoothing strength was adopted rather than tuned, and its tokens need not be semantic signifiers. The case audit is also a derived benchmark-data adaptation under CC BY-NC-SA 4.0.
The regularized-neighborhood comparison shares the oracle-budget and lexical identity-signifier limitations. Its 30-neighbor size and equal 30-case marginal prior were fixed without tuning, while its text-and-demographic representation was selected in earlier Project work. Coarse demographic similarity is not an ipseological mechanism. Its case audit is likewise a derived benchmark-data adaptation under CC BY-NC-SA 4.0.
The semantic-neighborhood comparison inherits the same oracle budget, candidate vocabulary, neighborhood size, and shrinkage. GloVe proximity is distributional rather than a validated ipseological identity measure; the centroid erases order, senses, negation, and response structure. Twitter- trained vectors can encode social bias and need not transfer cleanly to Twenty Statements Test language. The analysis was chosen after earlier aggregate results were known, and its failed hurdle rejects only this fixed method. Its case audit is a derived benchmark-data adaptation under CC BY-NC-SA 4.0; the public-domain vector file is hash-pinned but not redistributed here.
The change-volume comparison removes that oracle budget, but its surface-text and coarse-demographic representation and its 30-neighbor, equal-shrinkage choices came from prior Project work rather than untouched selection. Its overlapping folds and case-bootstrap intervals characterize this selected training cohort, not a population. Count forecasts remain continuous to avoid an arbitrary rounding rule, Add/Delete tokens need not all be identity signifiers, and the derived case audit is governed by CC BY-NC-SA 4.0.
The probabilistic extension inherits all of those design constraints. Its weighted empirical distributions reuse the same outcomes, folds, representation, and fixed weights; they are not independently selected models. Central 80% intervals are discrete empirical quantiles, and coverage in 150 overlapping leave-one-out folds is descriptive rather than a population guarantee. Its derived case audit is also governed by CC BY-NC-SA 4.0.
The source-feature ablation is a dependent extension chosen after the combined word-count gain was known. It preserves the existing exact-value demographic features, so sparse categories and missing values can weaken demographics-only similarity. Its unadjusted secondary contrasts are repeated-analysis diagnostics, not an independent family of confirmatory tests. The derived case audit remains governed by CC BY-NC-SA 4.0.
The source-form comparison is another dependent extension chosen after the text-only word-count gain was known. Its three counts and fixed distance rule test one narrow alternative, not every nonlexical representation. A direct interval spanning zero is evidence that this analysis does not distinguish the methods, not proof that their predictive distributions are equivalent. Its unadjusted secondary contrasts are repeated-analysis diagnostics, and its case audit remains governed by CC BY-NC-SA 4.0.
The multiplicity stress test is itself post hoc: its family was fixed only after the constituent analyses and aggregate results existed. Synchronized case resampling preserves empirical dependence among the 16 stored effects, but it does not refit the overlapping leave-one-out folds, undo sequential research choices, or turn a selected cohort into a probability sample. Its simultaneous intervals are therefore an internal case-composition sensitivity check rather than formal prospective familywise error control or population inference. The flat audit is derived from benchmark-data adaptations and remains governed by CC BY-NC-SA 4.0.
The calibrated-synthesis analysis is also dependent: every component and its three-part gate were chosen after the earlier training results were known. Exact repeated response units favor generic template language, and appending one statement per unit explains the severe line-count error. Its fold predictions transfer public training text and remain benchmark-data adaptations under CC BY-NC-SA 4.0. The pointwise gate is deliberately conservative but is not multiplicity-adjusted; failing it is a method decision, not a universal claim that semantic synthesis is impossible.
The paired bootstrap treats development cases as the resampling unit, but the benchmark is a selected longitudinal cohort rather than a probability sample. Its intervals describe sensitivity to the composition of these 50 cases, not uncertainty for people in general. The lexical oracles use observed follow-ups and are deliberately unattainable; they must not be read as test predictions or achievable prospective scores.
The immediate next step is to submit the two prepared artifacts to the challenge organizer and add the complete private test scorecard, unchanged, to this report. The fixed common-unit synthesis and GloVe-centroid ranking have both failed their training hurdles, so neither may be evaluated on development data. Further re-rankers on these same 150 cases are unlikely to be informative without richer compositional context or new longitudinal evidence. A later method should preserve response-form calibration, beat leave-one-out marginal additions in training, lock any later development comparison, and avoid private test feedback for tuning.
