Skip to content

1  Predict the Self

Evidence, method, and benchmark diagnostics

Author

Aleph Initial Alpha, Virtual CSSERG

Published

October 1, 2026

CSSERG logoVirtual CSSERG

Executive Summary · Short two-column report · report landing page

2 Abstract

An 81-case test submission is frozen and passes the official format validator. Because test answers are private, this report does not claim a test score. On 50 public development cases, a deterministic stable-signifier projection has mixed agreement results relative to repeating the 2024 response verbatim. A new post hoc analysis finds that the small agreement differences are sensitive to which development cases are included, while improvements in word-count and source-similarity error are more stable. More fundamentally, 73.2% of distinct tokens in an average observed follow-up did not appear in that person’s earlier response, exposing a severe limit for any strictly extractive forecast. A locked exploratory baseline retrieves the complete follow-up of the most similar training participant. It approximates how much new vocabulary will appear but not which vocabulary: predicted unique-token novelty is 80.2% versus 73.2% observed, while mean novel-token precision is 7.0%, recall is 8.5%, and agreement is lower than both continuity baselines. A post hoc volume-matched control sharpens the result: simply choosing the most frequent training-set additions reaches 17.7% precision and 14.3% F1, versus 7.0% and 6.7% for person-matched retrieval. A locked training-only leave-one-out test then asks whether earlier source tokens can improve that marginal ranking. The fixed source-conditioned rule recovers 14.5% of held-out additions versus 16.1% for the marginal prior. A second locked test pools and regularizes 30 similar text-and-demographic trajectories; it recovers 14.9% versus the same 16.1% marginal result, a paired difference of −1.20 percentage points (95% case-bootstrap interval: −1.84 to −0.57). Neither individualized ranking demonstrates incremental content advantage. A third locked leave-one-out test removes the oracle budget and forecasts how much each held-out response will change. The same regularized neighborhood lowers MAE relative to source- calibrated fold medians for Add count (20.126667 versus 21.340000), Delete count (5.336275 versus 5.614935), and word count (52.940000 versus 56.080000); all three paired intervals exclude zero. It ties on line count, while the source-similarity interval spans zero. A fourth lock evaluates the complete predictive distributions rather than only their medians. The neighborhood improves word-count CRPS (38.156852 versus 40.672695; paired reduction +2.515843, interval +1.008212 to +4.059792) and its central 80% interval score, while CRPS intervals for the other four outcomes span zero. A fifth locked analysis separates the combined source representation. Text-only matching improves word-count CRPS over the marginal by +2.298139 (interval +0.816444 to +3.865823); demographics-only matching does not, although adding demographics to text yields a smaller incremental gain of +0.217704 (+0.024229 to +0.420871). A sixth lock compares text-only matching with a count-only neighborhood based on the earlier response’s word, distinct-token, and line counts. Source-form-only CRPS is 37.427386 versus 38.374556 for text only. The direct form-minus-text effect is −0.947170, but its interval spans zero (−1.943737 to +0.035874), so this fixed comparison does not establish a lexical advantage beyond response form. A locked cross-analysis stress test then considers 16 no-oracle point, distributional, and primary word-count contrasts together. Six directions remain stable under simultaneous 95% intervals: Add-count and word-count point-error reductions, the word-count CRPS reduction, text-only and source-form word-count gains over the marginal, and combined matching over demographics alone. The Delete-count point gain and demographics’ small increment over text become inconclusive. A final locked training-only analysis integrates stable source units, a count- only word-length target, a combined-neighborhood Add-count target, and common fold follow-up units into one full-text synthesis. It lowers word-count error and improves ROUGE-L relative to repeat-2024, but edit similarity does not improve, line-count error rises sharply, and novel-type F1 is 0.075887 versus 0.134680 for an equal-volume marginal-addition ranking. The prespecified advancement gate fails, so no development prediction is permitted. A ninth locked training-only test replaces surface similarity with IDF-weighted centroids of public-domain GloVe Twitter vectors. Vector coverage is 94.95% of source token types, but semantic neighborhoods recover 0.148760 of held-out additions versus 0.160992 for marginal frequency. The paired difference is −0.012232 (interval −0.018361 to −0.006322), while the semantic and prior surface neighborhoods are indistinguishable. This semantic hurdle also fails, so development data remain unopened. A separately locked leave-one-out cross-validation of the frozen stable projection repeats its central mixed pattern: edit similarity improves by 0.004147, token-overlap F1 worsens by 0.003493, word-count and source-similarity error fall, and line-count error rises. The other non-exact agreement intervals span zero.

3 Findings at a glance

  • The benchmark is now pinned to commit 9b6a766, and every governing input has a recorded SHA-256 checksum.
  • The deterministic method learns which source tokens recur in later self-description and extracts high-persistence response units to a training-predicted length.
  • Relative to repeat-2024 on development data, it has higher normalized edit similarity, token Jaccard similarity, and ROUGE-L F1, and lower word-count and source-similarity error.
  • It also has lower token-overlap F1 and character n-gram F1 and higher line-count error. There is no composite score or defensible “overall winner.”
  • Paired case-bootstrap intervals span zero for all five nonzero agreement differences; only error in length and amount of change has a clearer pattern.
  • Leave-one-out prediction of all 150 training trajectories repeats that mixed pattern without using each case’s own follow-up in its fitted model. It does not convert method-development data into untouched confirmation.
  • It still forecasts too much continuity: mean prediction-to-source ROUGE-L is 0.903898, compared with 0.225683 for observed follow-ups.
  • On average, 73.2% of follow-up unique token types are new relative to the earlier response and therefore unavailable to a strictly extractive method.
  • Matched-trajectory retrieval approximates the volume of lexical novelty and reduces source-similarity error, but its retrieved new signifiers rarely match the focal person’s observed new signifiers.
  • At identical case-level novelty volume, a marginal prior based only on common training additions recovers more observed new token types than person-matched retrieval on 40 of 50 development cases.
  • In a separate leave-one-out analysis of all 150 training trajectories, the locked source-conditioned Add ranking performs worse than its marginal comparator on 68 cases, ties 63, and wins 19 at the same oracle token budget.
  • A strongly regularized 30-neighbor Add ranking also falls short: it loses 65 cases, ties 60, and wins 25, despite changing 26.0% of marginal-prior guesses on an average case.
  • Without an oracle future-token budget, that fixed neighborhood does improve leave-one-out forecasts of Add count, Delete count, and follow-up word count over source-calibrated fold medians. It does not improve line count or show a stable source-similarity advantage.
  • Full predictive distributions sharpen that result: neighborhood word-count CRPS and central 80% interval score improve, but Add count, Delete count, line count, and source-similarity CRPS differences remain compatible with zero.
  • A locked ablation locates most of that word-count signal in earlier self-description text. Demographics alone do not improve over the marginal, although they add a small incremental gain when combined with text.
  • A second locked ablation finds that three source-response counts improve word-count CRPS over the marginal and are not stably worse than text-only matching. The current evidence does not isolate a lexical mechanism.
  • A locked 16-contrast multiplicity stress test retains six simultaneous directions. It preserves the Add-count and word-count point gains and the principal word-count distribution results, but not the Delete-count point gain or the small demographics-over-text increment.
  • A locked full-text synthesis test improves response-length and source-change calibration but loses badly to an equal-volume marginal ranking on novel- type F1 and multiplies line-count error. It fails its advancement gate and is not evaluated on development data.
  • A locked external-semantic neighborhood also fails the marginal content hurdle: its mean recovered fraction is 0.148760 versus 0.160992 for common additions, despite 94.95% source-vocabulary coverage by the GloVe vectors.

4 Introduction

4.1 Research question

The larger project asks how predictable human lives are. Its first tractable task is Dr. Jason Jeffrey Jones’ Predict Future Selves challenge: predict a later self-authored self-description from an earlier self-authored self-description.

This is an ipseological prediction task. A Twenty Statements Test response is personally expressed identity: text an individual provides to describe themself. Words and phrases may act as identity signifiers. The target is a later expression of identity—not a latent “true self” and not a complete life outcome.

4.2 Benchmark and provenance

The benchmark contains 281 Prolific participants with approved, complete Twenty Statements Test responses in the 2024 study and its longitudinal follow-up. Follow-up responses were recorded 424 to 750 days later (median 433). The deterministic split contains 150 training, 50 development, and 81 test cases. Training and development responses are paired; test follow-ups and follow-up demographics are private.

The project retrieved the public repository on 2026-08-30 and detached the checkout at commit 9b6a766712583fec8d3182957260b1123fbfa146 before prediction. The four input hashes and the hashes of the governing documentation, evaluator, and validator are recorded in BENCHMARK_PROVENANCE.md.

The benchmark’s own warnings govern interpretation: the sample is small and selected; platform recruitment and longitudinal attrition limit generalization; and agreement with one observed follow-up does not establish that a prediction was the only plausible future.

5 Method

5.1 Stable-signifier projection

The method is deterministic and extractive. It uses the 150 public training pairs and no external data or model.

  1. For every source token, estimate the proportion of training documents in which that token also appears at follow-up. Smooth toward the corpus-wide retention rate with ten equivalent prior observations.
  2. Divide each source response using its strongest repeated boundary: lines, sentences, comma phrases, or the complete response.
  3. Score each response unit by the mean estimated retention probability of its unique content tokens.
  4. Fit a training-only ordinary least squares regression of follow-up word count on source word count.
  5. Retain high-scoring units while doing so moves the response closer to the predicted length, then restore their original order.

Demographic metadata is not used. The method has no random choices and no case-level manual editing. It is deliberately interpretable: every predicted word comes from that person’s earlier self-description.

5.2 Model selection

Public development data were used for model selection, as the challenge allows. Fixed extraction fractions and the training-only length regression were compared; the regression-length rule was frozen before the test artifact was generated. The development scorecard is therefore a model-selection result, not an untouched confirmatory estimate.

5.3 Locked training cross-validation of the stable projection

Before generating any fold prediction, the Project fixed a leave-one-out plan at SHA-256 2929b1fb8111c42170d23e725c08f6167135d9a703a1b61e87ed59af5f97c5d8 and hash-guarded the unchanged generator. Each of the 150 training cases is held out in turn; the original retention estimator and length regression are refit on the other 149, with the same ten-observation smoothing prior, and the held-out response is predicted from its 2024 text alone. Repeat-2024 is the paired comparator.

The analysis reports the full official scorecard and 20,000 paired case- bootstrap resamples with seed 20260922. It is a cross-validation diagnostic, not a preregistration or fresh confirmation: the training corpus and public development results informed earlier method development. Withholding each case’s own follow-up prevents direct fold leakage, but it cannot erase that history or establish population generalization.

5.4 Post hoc development diagnostics

After freezing the model and predictions, later work added two descriptive diagnostics. First, every scorecard comparison was decomposed into 50 paired case differences. A deterministic percentile bootstrap resampled cases in pairs 20,000 times (seed 20260917) to show how the mean effect changes with development-case composition. These intervals do not warrant population generalization, and the analysis was not preregistered.

Second, tokenized 2024 and follow-up responses were compared case by case. Unique-token novelty is the share of follow-up token types absent from the earlier response. Occurrence unavailability additionally respects how many times each token appears in the source: a strictly extractive method cannot copy a token more often than it occurs there. Three deliberately unattainable oracles use each observed follow-up to select the best source token set, bag of words, or source-order subsequence. They describe extractive ceilings, not prospective prediction performance.

5.5 Locked matched-trajectory retrieval

A second method was locked before its predictions were generated. This is a prospective analysis lock, not a preregistration: the development labels and aggregate stable-projection results were examined in earlier work, so the comparison remains exploratory. The complete plan was fixed at SHA-256 86b2ddfe775773e1964beeefbac5479d88f6c46cd3385bf3a992fc7ae64e9c83.

Four one-nearest-neighbor variants were first compared by leave-one-out prediction among the 150 training pairs. TF-IDF cosine over text plus 2024 demographics had the highest training token-overlap F1 (0.1925) and was selected by that declared lexical criterion. For each development query, the locked method:

  1. represents case-folded response tokens and field-qualified, nonblank 2024 demographic values with smoothed TF-IDF weights;
  2. selects the training source with greatest cosine similarity, breaking ties by training-file order; and
  3. uses that neighbor’s observed follow-up verbatim as the prediction.

The baseline is non-extractive with respect to the focal person’s earlier response, but it retrieves rather than synthesizes. It deliberately tests whether a prior person’s trajectory supplies the right kind and amount of new content. It generated development predictions only and did not create or alter a test submission. The copied training text is public benchmark data licensed CC BY-NC-SA 4.0 and is presented only as an auditable baseline, not as a factual description of the focal participant.

The same complete scorecard is reported below. Directional metrics were also compared within person using 20,000 paired case-bootstrap resamples (seed 20260918). Additional diagnostics compare predicted and observed novelty and measure precision, recall, and F1 for predicted novel token types. As above, the intervals describe development-case composition rather than a population.

5.6 Post hoc volume-matched novelty control

The retrieval method introduces many token types, creating more opportunities for chance overlap with the observed follow-up. A new diagnostic therefore holds this opportunity constant case by case. It counts, among the 150 training pairs, how many people add each token type between waves. For each development source, it excludes tokens already present and chooses the most frequently added remaining types, breaking frequency ties alphabetically. The number chosen is exactly the number of novel types emitted by retrieval for that case.

This marginal-addition prior uses no development response to construct its ranking, and it has the same novel-token budget as retrieval on every case. However, it borrows that budget from retrieval and produces only a token set, not a coherent response. It is therefore a diagnostic control rather than a standalone challenge forecast. The comparison was devised after development labels had already been inspected and is explicitly post hoc. Lexical tokens also include function words and other language that need not be semantic identity signifiers.

5.7 Locked training-only source-conditioned ranking

A final diagnostic tests individualization without using the development cases again. Its plan was fixed before implementation and scoring at SHA-256 e65f04fd9e55843db8ff1a1bb0544acb00ec869eb58f9b312a63d531fc5f4ec9. This is an analysis lock, not a preregistration: the training pairs had already informed earlier methods.

Each of the 150 training cases is held out in turn. On the other 149, the analysis counts marginal Add events and every association between a token in the source and a different token added at follow-up. For a candidate addition, the source-conditioned score is its largest conditional Add rate across the held-out person’s source tokens, smoothed toward the candidate’s marginal Add rate with ten prior cases. That prior strength was fixed to match the stable projection’s existing smoothing choice. Ties favor marginal frequency and then alphabetical token order.

Both rankings receive exactly as many guesses as the held-out follow-up actually adds. This oracle budget isolates token ordering but disqualifies both rankings as standalone forecasts. Because predicted and observed token sets have equal size, case-level precision, recall, and F1 are identical; the report calls their shared value the recovered fraction. A 20,000-resample paired case bootstrap uses seed 20260920.

5.8 Locked regularized-neighborhood ranking

A second training-only plan was fixed before implementation or scoring at SHA-256 c7a494327dd1dd7fc83911540a6a9095fcc8d684172a7388bd610206ac4ca13d. It tests whether pooling multiple similar trajectories avoids the sparse-cue problem of the source-token rule. The source representation is reused from the earlier retrieval analysis: fold-fit TF-IDF over source text and field-qualified 2024 demographics.

For each held-out case, the other 149 cases form the training fold. The method selects the 30 most cosine-similar sources, weights their Add events by similarity, and rescales the weights to 30 effective cases. Each candidate’s weighted neighborhood count is then combined with 30 equivalent cases at its fold-wide marginal Add rate. This one-to-one shrinkage and the neighborhood size were fixed rather than tuned. The comparator is the same leave-one-out marginal ranking; both receive the held-out case’s observed addition count as an oracle budget. A 20,000-resample paired case bootstrap uses seed 20260921. The plan is an analysis lock, not a preregistration, because both the training data and source representation informed earlier work.

5.9 Locked prospective change-volume forecasting

A third training-only plan was fixed before retrieving the benchmark in this iteration, implementation, or scoring at SHA-256 53c974fcd37f443ea846e88328a125169265fb41ac1cc7877529bdbf9a09638f. It reuses the same 30-neighbor text-and-demographic representation and equal 30-case neighborhood/prior weights, but removes the oracle future-token budget. The question is now whether similarity helps forecast response form and the number of Add and Delete events.

For each held-out case, the other 149 provide five modeled quantities: Add count; Delete count as a fraction of the source’s distinct token count; follow-up-minus-source word and nonempty-line counts; and follow-up-to-source ROUGE-L. A source-calibrated marginal forecast uses the fold median. The neighborhood forecast uses a weighted median combining 30 similarity-weighted neighbors with a 30-case prior spread across the whole fold. Fractions and changes are transformed back using only the held-out 2024 response. The analysis therefore uses no follow-up-derived budget or other held-out future information for prediction.

The primary effect is marginal absolute error minus neighborhood absolute error, so positive values favor person conditioning. Twenty thousand paired case-bootstrap resamples use seed 20260923, restarted for each outcome. This lock cannot undo earlier use of the training corpus or earlier selection of the representation and hyperparameters.

5.10 Locked probabilistic change-volume forecasting

A fourth training-only plan extends that fixed volume model from point forecasts to complete predictive distributions. It was locked before reopening the benchmark in this iteration, implementation, or scoring at SHA-256 00d5e2805082986b00f2716808ee12d4cff70ec900a5e47636e893217368f020. For each held-out case and outcome, the marginal distribution gives equal weight to all 149 fold observations. The neighborhood distribution combines a 30-case fold-wide prior with the same 30 cosine-weighted neighbors used above. Every support point is transformed to the held-out source scale before scoring.

The primary score is the continuous ranked probability score (CRPS), a proper score that rewards predictive distributions concentrated near the observation (Gneiting & Raftery, 2007); lower values are better. Secondary diagnostics use central 80% intervals and report inclusive coverage, width, and interval score, which penalizes both width and misses. Positive paired effects are marginal score minus neighborhood score. Twenty thousand paired case-bootstrap resamples use seed 20260924, restarted for each outcome and score. Representation, neighborhood size, and shrinkage remain unchanged and untuned.

5.11 Locked source-feature ablation

A fifth training-only plan was locked before reopening the benchmark in this iteration, implementation, or scoring at SHA-256 ed9fa67b2cc1265f3710471f96d7ffbfd16986b8295ef5072da26c263d8b7da4. It asks whether the previously observed word-count distribution gain comes from the earlier self-description, the ten 2024 demographic fields, or their combination. It splits the unchanged TF-IDF representation into text-only, demographics-only, and combined feature sets without changing any field, neighbor, weighting, shrinkage, transformation, or scoring rule.

Follow-up word-count CRPS is primary because it was the only outcome with a supported combined-neighborhood CRPS gain in the preceding lock. The four other outcomes are prespecified secondary diagnostics. Four paired contrasts compare each ablation with the marginal distribution and the combined model. Twenty thousand case-bootstrap resamples use seed 20260925, restarted for each outcome and contrast. Before any ablation result is accepted, the script must numerically reproduce every marginal and combined case-level CRPS value in the preceding locked audit.

5.12 Locked source-form comparison

A sixth training-only plan was locked before reading the benchmark rows in this iteration, implementation, or scoring at SHA-256 5ffaf6b8289a360926da3ae3b1398a336949e698c6157f8ba94ef316ec63c5ad. It compares the text-only distribution with a count-only source-form neighborhood. Each source is represented by log word-token count, log distinct-token count, and log line count. Fold-fit population standard deviations scale Euclidean distance; similarity is 1 / (1 + distance).

The neighborhood size, weights, prior, outcome transformations, and CRPS scoring remain fixed. Follow-up word-count CRPS is primary, and source-form CRPS minus text-only CRPS is the primary paired contrast. Twenty thousand case-bootstrap resamples use seed 20260926. The script must reproduce every preceding marginal and text-only case score before accepting a result. A text advantage would identify lexical composition—not necessarily semantic identity—because TF-IDF can encode style and template wording.

5.13 Locked cross-analysis multiplicity stress test

A seventh plan was locked after the constituent aggregate results were known but before any cross-analysis resampling at SHA-256 d3c31e1bd77df8da0d5b7017438f2b9ff04ba5f39c4dfcb803c1c38f294fb9ac. It fixes one family of 16 unique no-oracle contrasts: marginal-minus- neighborhood error for five point outcomes, marginal-minus-neighborhood CRPS for the same five outcomes, and the six prespecified word-count CRPS contrasts from the feature and source-form ablations. Duplicated inherited contrasts are counted once. Central-interval scores, secondary ablation outcomes, oracle- budget token rankings, full-text scorecards, and development analyses are outside this deliberately bounded family.

One synchronized case bootstrap resamples the 150 case indices 20,000 times with seed 20260929, preserving the empirical dependence among all 16 stored case-effect vectors. Each resample contributes its maximum absolute studentized mean deviation. The nearest-rank 95th percentile of those maxima is the common critical value for two-sided simultaneous intervals. This is a post hoc sensitivity audit, not retroactive preregistration or formal population familywise control: stored leave-one-out scores are resampled without refitting their heavily overlapping folds.

5.14 Locked calibrated full-text synthesis

An eighth training-only plan was locked before reading benchmark rows in this iteration, implementation, or scoring at SHA-256 ae0594e0f95f4131416be0bab4c3bfa913010dc6380d65c39966849c3a9f7253. It asks whether the Project’s separately observed signals can be composed into one prospective text forecast. Each outer fold ranks held-out source units with the frozen stable-token rule, forecasts follow-up word count with the fixed three-count source-form neighborhood, forecasts distinct Add count with the fixed text-and-demographic neighborhood, and ranks complete follow-up units that are exact-unit additions in at least two fold cases.

The generator searches every nonempty prefix of ranked stable source units and every prefix of ranked common additions. It minimizes the sum of normalized absolute error from the predicted word and Add counts, then applies fixed source-retention and parsimony tie rules. All fitting and candidate text come from the other 149 cases. A volume-matched marginal token ranking receives exactly as many novel-token guesses as the synthesized response, so its content comparison uses no held-out future budget.

Twenty thousand paired case-bootstrap resamples use seed 20260930, restarted for every contrast. Development prediction is allowed only if three pointwise 95% intervals all exclude zero in the favorable direction: normalized edit similarity versus repeat-2024, word-count error versus repeat-2024, and novel- type F1 versus the volume-matched marginal ranking. This conservative gate is not a multiplicity correction. The design is dependent on earlier results and the same training cohort, so even a passing result would not be untouched confirmation.

5.15 Locked semantic-neighborhood ranking

A ninth training-only plan was locked before reading benchmark rows in this iteration, implementation, or scoring at SHA-256 6a6a536a11c552b0e75f40ce6d79100902727958fa908cb57a662ac85df5fdcf. It tests whether external distributional semantics improves the selection of novel token types. The fixed representation uses the public-domain, 25-dimensional uncased GloVe Twitter vectors trained on 2 billion tweets and 27 billion tokens (Pennington et al., 2014). The exact Gensim-data conversion is pinned at SHA-256 63877d71151688baf6f31d5437374f637f737a5e100e12150a5bd61a9f273c3f; the 104 MB external artifact is streamed for analysis but not committed.

Within each 149-case fold, source responses become centroids of their distinct in-vocabulary token vectors, weighted by fold-fit inverse document frequency. Cosine similarity selects 30 neighbors. Their weighted Add counts contribute 30 effective cases and are shrunk equally toward fold-wide marginal Add rates, exactly matching the earlier regularization rule. The marginal ranking and the inherited surface-text-plus-demographic neighborhood receive the same oracle held-out addition budget. The script must reproduce every inherited marginal and surface-neighborhood hit count before accepting the semantic result.

The primary effect is semantic-neighborhood minus marginal recovered fraction; the secondary effect compares semantic and surface neighborhoods. Twenty thousand paired case-bootstrap resamples use seed 20261001, restarted for each contrast. A semantic advantage is claimable only if the primary mean and interval are positive. Regardless of outcome, this iteration cannot reopen development data.

6 Results

6.1 Aggregate scorecard

The authoritative shared evaluator reports fifteen measures in three groups. The table reports every measure; arrows indicate the evaluator’s direction where one exists.

Development-set aggregate scorecard
Group Metric Stable projection Trajectory retrieval Repeat 2024 Direction
Agreement Normalized exact-match rate 0.000000 0.000000 0.000000 Higher
Agreement Normalized edit similarity 0.298061 0.253360 0.291966 Higher
Agreement Token Jaccard similarity 0.142768 0.080427 0.141930 Higher
Agreement Token-overlap F1 0.307444 0.194693 0.312202 Higher
Agreement ROUGE-L F1 0.227552 0.150924 0.225683 Higher
Agreement Character n-gram F1 0.292293 0.212660 0.296534 Higher
Form Word-count MAE 41.640000 59.640000 56.680000 Lower
Form Line-count MAE 9.760000 8.340000 6.800000 Lower
Form Mean predicted word count 77.320000 99.280000 108.040000 Descriptive
Form Mean reference word count 94.760000 94.760000 94.760000 Descriptive
Change Prediction repeat-2024 rate 0.280000 0.000000 1.000000 Descriptive
Change Observed repeat-2024 rate 0.000000 0.000000 0.000000 Descriptive
Change Source-similarity MAE 0.678215 0.136074 0.774317 Lower
Change Mean prediction-to-source similarity 0.903898 0.175191 1.000000 Descriptive
Change Mean observed follow-up-to-source similarity 0.225683 0.225683 0.225683 Descriptive

The projection improves three of six agreement measures, worsens two, and ties on exact match. Its clearest gains concern form and amount of change: word-count MAE falls by 15.04 words, and source-similarity MAE falls by 0.096102. But the mean predicted source similarity remains 0.678215 above the observed mean. Extracting supposedly enduring signifiers does not come close to reproducing how radically people rewrite their self-descriptions.

Trajectory retrieval creates the opposite pattern. Its mean source similarity (0.175191) is close to the observed 0.225683, so source-similarity MAE falls to 0.136074. But every non-exact agreement metric is lower than both continuity baselines. Generating a plausibly different response is therefore not equivalent to predicting the content of the focal person’s future self-description.

6.2 Paired development-case pattern

Positive effects favor stable-signifier projection. For agreement measures, the effect is projection minus repeat-2024; for errors, it is repeat-2024 error minus projection error. Exact match is omitted because all 50 cases tie at zero. A win, tie, or loss is evaluated within a person before averaging.

Paired development-case comparison with the repeat-2024 baseline
Metric Mean effect Paired case-bootstrap 95% interval Win / tie / loss
Normalized edit similarity +0.006095 −0.001701 to +0.013890 24 / 14 / 12
Token Jaccard similarity +0.000838 −0.002383 to +0.004097 16 / 20 / 14
Token-overlap F1 −0.004759 −0.012300 to +0.002429 13 / 20 / 17
ROUGE-L F1 +0.001869 −0.002268 to +0.005755 19 / 20 / 11
Character n-gram F1 −0.004241 −0.012240 to +0.003618 16 / 14 / 20
Word-count error reduction +15.040000 +1.180000 to +32.360000 20 / 20 / 10
Line-count error reduction −2.960000 −5.540000 to −0.420000 15 / 12 / 23
Source-similarity error reduction +0.096102 +0.066064 to +0.129348 30 / 20 / 0

The table changes the emphasis of the aggregate scorecard. None of the small agreement differences has an interval that excludes zero. In contrast, the projection’s lower word-count and source-similarity errors persist across the paired case resamples, while its line-count error is consistently worse. The source-similarity result is directional but inadequate in magnitude: even after improvement, the projection still remains much too close to the past.

6.3 Training cross-validation repeats the mixed pattern

The locked leave-one-out analysis refits the frozen projection 150 times. As in the development analysis, positive paired effects favor the projection; for error measures they are repeat-2024 error minus projection error.

Stable-projection training cross-validation results
Metric Projection Repeat 2024 Mean effect Paired case-bootstrap 95% interval
Normalized exact match 0.000000 0.000000 0.000000 0.000000 to 0.000000
Normalized edit similarity 0.276809 0.272662 +0.004147 +0.000652 to +0.007753
Token Jaccard similarity 0.133873 0.134297 −0.000424 −0.002229 to +0.001237
Token-overlap F1 0.276737 0.280231 −0.003493 −0.007228 to −0.000077
ROUGE-L F1 0.200304 0.200017 +0.000287 −0.002003 to +0.002538
Character n-gram F1 0.266194 0.267813 −0.001620 −0.005530 to +0.002270
Word-count MAE 46.986667 55.580000 +8.593333 +2.573333 to +15.266667
Line-count MAE 9.066667 6.160000 −2.906667 −4.346833 to −1.480000
Source-similarity MAE 0.729214 0.799983 +0.070769 +0.054503 to +0.088484

Only two nonzero agreement intervals exclude zero, and they point in opposite directions: normalized edit similarity improves slightly while token-overlap F1 worsens slightly. The other agreement differences remain unstable to cohort composition. The form/change findings are more reproducible: the method lowers word-count and source-similarity error but raises line-count error. Magnitude remains the substantive problem. Mean predicted source similarity is 0.929231, versus 0.200017 for observed follow-ups, and 46.7% of predictions repeat the source exactly while no observed follow-up does.

The direction of all three error results matches development; four of five non-exact agreement directions also match, with token Jaccard shifting from a tiny positive difference to a tiny negative one. This is useful recurrence within the same method-development corpus, not independent replication.

6.4 Lexical novelty and extractive ceilings

The observed follow-ups contain substantial new lexical material. On an average case, 73.1727% of distinct follow-up token types are absent from the earlier response (95% case-bootstrap interval: 69.7982%–76.5607%). When token frequencies are respected, 65.0621% of follow-up token occurrences cannot be copied from the source without reusing tokens beyond their source counts (59.2663%–70.6149%). Conversely, only 24.9892% of distinct source token types recur at follow-up (22.0060%–28.0466%).

Even an oracle with access to the observed future is constrained. Its mean extractive ceilings are 0.268273 for unique-token Jaccard, 0.484206 for bag-of-words F1, and 0.374081 for source-order subsequence ROUGE-L F1. The actual projection reaches 0.142768, 0.307444, and 0.227552 on those measures. These comparisons reveal two failures at once: the current extractor does not reach the retrospective extractive ceiling, and no extractor can generate the many identity signifiers expressed only at follow-up.

6.5 Retrieval predicts novelty volume, not novel content

Positive paired effects favor trajectory retrieval. For errors, the effect is stable-projection error minus retrieval error. Retrieval loses agreement and word-count accuracy, modestly improves line-count error with an interval that spans zero, and dramatically improves source-similarity error.

Trajectory-retrieval comparison with the stable projection
Metric Retrieval Stable projection Paired effect Paired case-bootstrap 95% interval
Normalized exact match 0.000000 0.000000 0.000000 0.000000 to 0.000000
Normalized edit similarity 0.253360 0.298061 −0.044701 −0.072824 to −0.018705
Token Jaccard similarity 0.080427 0.142768 −0.062342 −0.081759 to −0.043776
Token-overlap F1 0.194693 0.307444 −0.112751 −0.170815 to −0.058175
ROUGE-L F1 0.150924 0.227552 −0.076628 −0.126464 to −0.030063
Character n-gram F1 0.212660 0.292293 −0.079633 −0.112019 to −0.049794
Word-count MAE 59.640000 41.640000 −18.000000 −30.120500 to −6.160000
Line-count MAE 8.340000 9.760000 +1.420000 −1.860000 to +4.640000
Source-similarity MAE 0.136074 0.678215 +0.542141 +0.463332 to +0.616781

The locked novelty criterion tells the same story more precisely. Retrieval’s mean predicted unique-token novelty (0.801727) is much closer to the observed mean (0.731727) than either stable projection or repeat-2024 (0 for both), so it passes the plan’s novelty-volume criterion. Its case-level novelty MAE is 0.139032. Occurrence novelty is also close in aggregate (0.727677 predicted versus 0.650621 observed).

However, novelty volume is not novel-content recovery. Among token types that retrieval introduces beyond the focal source, mean precision against the observed new token types is only 0.070466 (case-bootstrap interval 0.052133–0.090859), recall is 0.085362 (0.065937–0.106057), and F1 is 0.067109 (0.052917–0.082272). Retrieval knows that a future response will look different, but mostly borrows the wrong person’s differences.

6.6 Person matching does not beat common additions

The volume-matched marginal prior predicts a mean 46.94 novel token types per case, exactly equal to retrieval by construction. Despite lacking any person-matching rule, it recovers more of the observed novel vocabulary:

Volume-matched novel-token recovery comparison
Novel-type measure Marginal prior Trajectory retrieval Paired difference Paired case-bootstrap 95% interval
Precision 0.176545 0.070466 +0.106079 +0.068782 to +0.154891
Recall 0.172611 0.085362 +0.087249 +0.062210 to +0.113649
F1 0.142637 0.067109 +0.075528 +0.055927 to +0.094990

Positive differences favor the marginal prior. It wins 40 cases, ties 6, and loses 4 on each measure. The common additions include syntactic vocabulary such as an, in, my, is, and who, but also more content-bearing words such as friend, lover, kind, creative, and life. This means the earlier retrieval F1 cannot be read as evidence that matching text and demographics identified person-specific new signifiers. At matched volume, a coarse cohort-level base rate is substantially stronger.

The result does not establish that novel identity content is inherently unpredictable. It establishes a stricter baseline for future work: an individualized novelty model should beat common training additions at a comparable prediction volume before its gains are attributed to person-level conditioning.

6.7 Source-token conditioning also loses to common additions

The fixed source-conditioned ranking does not clear that baseline in its training-only leave-one-out test. Across all 150 held-out training cases, the mean recovered fraction is 0.145238, compared with 0.160992 for the marginal-addition ranking:

Source-conditioned ranking compared with marginal additions
Training leave-one-out result Source-conditioned Marginal additions Paired difference
Mean novel-type recovered fraction 0.145238 0.160992 −0.015754
Median novel-type recovered fraction 0.141177 0.156250 —
Case wins / ties / losses 19 / 63 / 68 — —
Paired case-bootstrap 95% interval — — −0.020829 to −0.010736

The interval excludes zero in the direction favoring the marginal ranking. Thus, under the locked interpretation rule, the source-conditioned method provides no incremental advantage. One plausible explanation is that choosing the strongest source-token cue promotes sparse, idiosyncratic pairs that fail to recur in the held-out person, while the marginal ranking spends more of its fixed budget on common additions; the analysis does not isolate that mechanism.

This is a bounded negative result, not evidence that source signifiers contain no predictive information. The rule tests only token pairs, not phrases, semantics, interactions, or predicted deletions. Its mean candidate-vocabulary ceiling is 0.773139: about 22.7% of held-out additions are absent from every other training follow-up’s additions and cannot be selected by either ranking. Most importantly, the oracle budget comes from the held-out future. The result compares ranking quality within this training cohort; it is not development or private-test performance.

6.8 Regularized neighborhoods still lose to common additions

Pooling rather than maximizing person-conditioned evidence narrows the deficit but does not clear the marginal baseline:

Regularized-neighborhood ranking compared with marginal additions
Training leave-one-out result Regularized neighborhood Marginal additions Paired difference
Mean novel-type recovered fraction 0.149007 0.160992 −0.011984
Median novel-type recovered fraction 0.148542 0.156250 —
Case wins / ties / losses 25 / 60 / 65 — —
Paired case-bootstrap 95% interval — — −0.018426 to −0.005693

The interval again excludes zero in the direction favoring common additions. The neighborhood and marginal top sets overlap by 0.739706 on average, so the personalized method changes about 26.0% of guesses rather than merely reproducing its comparator. Mean cosine similarity across the 30 selected neighbors is only 0.169564, consistent with a diffuse local neighborhood, although this diagnostic cannot determine why its substitutions are worse.

The result strengthens the bounded inference across two specifications. A maximum source-token association loses by 1.58 percentage points; a pooled, strongly regularized text-and-demographic neighborhood loses by 1.20 points. It does not follow that individualized prediction is impossible. Both methods operate on surface lexical overlap, both use an oracle future-token budget, and both draw candidates from other cases’ additions. The shared candidate-vocabulary ceiling remains 0.773139.

6.9 Distributional semantics also loses to common additions

The pretrained representation covers 2409 of 2537 distinct source token types (94.9547%), and every source has at least one vector. High coverage does not translate into better Add selection:

Semantic-neighborhood ranking and its locked comparators
Training leave-one-out result Semantic neighborhood Surface neighborhood Marginal additions Semantic minus marginal
Mean novel-type recovered fraction 0.148760 0.149007 0.160992 −0.012232
Median novel-type recovered fraction 0.153846 0.148542 0.156250 —
Semantic wins / ties / losses versus marginal 22 / 65 / 63 — — —
Paired case-bootstrap 95% interval — — — −0.018361 to −0.006322

The primary interval excludes zero in the direction favoring common additions, so the semantic hurdle fails. Semantic and marginal top sets overlap by 0.746530 on average: semantic conditioning changes about a quarter of the guesses but makes them worse. Its difference from the inherited surface neighborhood is only −0.000247, with an interval spanning zero (−0.006521 to +0.005834); the two representations are not distinguished by this test.

Mean cosine similarity among selected semantic neighbors is 0.975307, but that high value partly reflects dense centroid geometry and is not an identity similarity scale. The fixed centroid discards word order, senses, negation, and line structure. The result therefore rejects this representation and ranking rule, not semantic conditioning generally. No development prediction was generated.

6.10 Neighborhoods predict some revision volume

When the oracle addition budget is removed, the same neighborhood has a more limited but positive role. It improves absolute-error forecasts for how many distinct token types are added and deleted and for follow-up word count:

Training leave-one-out change-volume forecasts
Training leave-one-out outcome Neighborhood MAE Source-calibrated marginal MAE Error reduction Paired case-bootstrap 95% interval
Add count 20.126667 21.340000 +1.213333 +0.680000 to +1.740000
Delete count 5.336275 5.614935 +0.278660 +0.022274 to +0.545876
Follow-up word count 52.940000 56.080000 +3.140000 +1.093333 to +5.213333
Follow-up line count 6.160000 6.160000 0.000000 0.000000 to 0.000000
Source similarity 0.106781 0.109648 +0.002867 −0.000960 to +0.006730

Positive reductions favor the regularized neighborhood. Its Add-count forecast wins 79 cases, ties 17, and loses 54; Delete count wins 88, ties 5, and loses 57; word count wins 82, ties 13, and loses 55. The line-count forecasts are identical for all cases because both weighted medians select the same modeled line change. The source-similarity point estimate favors neighborhoods, but its interval spans zero.

The magnitudes are modest: Add-count MAE falls by 5.7%, Delete-count MAE by 5.0%, and word-count MAE by 5.6% relative to their comparators. Still, the direction matters. Surface text and coarse demographics contain some case-specific information about how much personally expressed identity will be revised, even though the same representation makes worse choices about which new tokens will appear. These are volume and form forecasts within the selected training cohort, not full-text forecasts or evidence that demographic similarity is an ipseological mechanism.

6.11 Neighborhood distributions improve only word-count forecasts

Scoring complete predictive distributions narrows the point-forecast result. Only follow-up word count has a paired CRPS interval that excludes zero in the direction favoring the neighborhood:

Training leave-one-out predictive-distribution scores
Training leave-one-out outcome Neighborhood CRPS Source-calibrated marginal CRPS CRPS reduction Paired case-bootstrap 95% interval
Add count 14.438038 14.650556 +0.212518 −0.071030 to +0.497525
Delete count 3.853656 3.926457 +0.072801 −0.036356 to +0.178244
Follow-up word count 38.156852 40.672695 +2.515843 +1.008212 to +4.059792
Follow-up line count 4.922609 4.910848 −0.011761 −0.097952 to +0.081212
Source similarity 0.074684 0.077171 +0.002486 −0.000106 to +0.005180

The word-count neighborhood wins 83 cases and loses 67 on CRPS. Its central 80% interval covers 84.7% of cases, compared with 80.7% for the marginal distribution, while mean width rises from 149.960000 to 154.526667 words. Despite that modest widening, mean interval score improves from 264.626667 to 244.926667: the paired reduction is +19.700000 with interval +3.873167 to +37.193333. Source-similarity interval score also improves, but its primary CRPS interval spans zero. The other interval-score comparisons span zero as well.

These distributional scores do not overturn the point-forecast findings. Rather, they locate their strongest probabilistic support in follow-up length. The Add and Delete medians improve absolute error, but their empirical distributions do not establish stable CRPS gains. Line-count intervals are substantially over-covering (96.7% marginal and 92.7% neighborhood against 80% nominal), showing that apparent uncertainty can be broad without being informative. None of these quantities identifies the future signifiers.

6.12 Earlier text outperforms demographics for word-count skill

The required replication passed for all 150 cases and five outcomes. For the primary follow-up word-count outcome, text-only matching preserves most of the combined neighborhood’s advantage, while demographics-only matching does not establish an improvement over the marginal distribution:

Word-count predictive-distribution feature ablation
Word-count predictive distribution Mean CRPS Locked paired contrast CRPS reduction Paired case-bootstrap 95% interval
Source-calibrated marginal 40.672695 — — —
Text only 38.374556 Marginal minus text only +2.298139 +0.816444 to +3.865823
Demographics only 40.471984 Marginal minus demographics only +0.200711 −0.479832 to +0.861529
Combined 38.156852 Text only minus combined +0.217704 +0.024229 to +0.420871

Text-only matching wins 84 cases and loses 66 against the marginal. The combined model wins 85 and loses 65 against text-only. Thus the earlier self-description carries nearly all of the supported response-length signal; the ten coarse demographic fields alone do not beat the source-calibrated marginal. Their small incremental benefit conditional on text is compatible with weak complementary information, not a stand-alone demographic mechanism.

The secondary diagnostics do not support a general feature-family conclusion. For Add count, Delete count, and line count, every fixed ablation contrast has an interval spanning zero. Text-only matching improves source-similarity CRPS over the marginal by +0.002964, but its repeated-analysis interval barely excludes zero (+0.000051 to +0.006065), and adding demographics to text does not improve that outcome. The primary result remains about response length, not future signifiers or revision volume generally.

6.13 Simple source form matches text-only word-count skill

The required replication again passed for all 150 cases and five outcomes. Matching on only the earlier response’s word-token, distinct-token, and line counts improves the primary word-count forecast over the source-calibrated marginal:

Word-count lexical-versus-source-form comparison
Word-count predictive distribution Mean CRPS Locked paired contrast CRPS reduction Paired case-bootstrap 95% interval
Source-calibrated marginal 40.672695 — — —
Text only 38.374556 Marginal minus text only +2.298139 +0.826598 to +3.834257
Source form only 37.427386 Marginal minus source form +3.245309 +1.839350 to +4.784335
Direct comparison — Source form minus text only −0.947170 −1.943737 to +0.035874

The source-form neighborhood wins 96 cases and loses 54 against the marginal. Text-only matching wins 65 and loses 85 in the direct comparison, but the prespecified primary interval narrowly includes zero. Under the locked interpretation rule, this analysis does not separate the two methods’ primary skill. It therefore supplies no stable evidence that lexical composition adds word-count information beyond three simple counts. Nor does the lower mean for source form prove it is generally superior: the direct interval also prevents that claim.

The secondary pattern reinforces the narrow measurement interpretation. Source-form matching improves line-count CRPS by +0.445157 over the marginal (interval +0.206291 to +0.701574) and by +0.472157 over text only (+0.235305 to +0.721005). Add-count and Delete-count comparisons span zero. Source-form matching’s source-similarity reduction is +0.001969, with a repeated-analysis interval barely above zero (+0.000017 to +0.003880). These unadjusted secondary diagnostics concern response form, not future identity content.

6.14 Six directions survive the multiplicity stress test

The synchronized bootstrap’s common studentized critical value is 2.959538, larger than a contrast-by-contrast normal critical value. Six of the 16 fixed directions have simultaneous 95% intervals excluding zero:

Simultaneous case-composition stress test across 16 training-only contrasts
Fixed contrast Mean effect Simultaneous 95% interval Direction within this audit
Add-count point-error reduction +1.213333 +0.397788 to +2.028879 Neighborhood
Delete-count point-error reduction +0.278660 −0.116978 to +0.674298 Inconclusive
Word-count point-error reduction +3.140000 +0.047580 to +6.232420 Neighborhood
Line-count point-error reduction 0.000000 0.000000 to 0.000000 Exact tie
Source-similarity point-error reduction +0.002867 −0.002885 to +0.008619 Inconclusive
Add-count CRPS reduction +0.212518 −0.217903 to +0.642939 Inconclusive
Delete-count CRPS reduction +0.072801 −0.089750 to +0.235353 Inconclusive
Word-count CRPS reduction +2.515843 +0.223750 to +4.807936 Neighborhood
Line-count CRPS reduction −0.011761 −0.148523 to +0.125000 Inconclusive
Source-similarity CRPS reduction +0.002486 −0.001516 to +0.006488 Inconclusive
Marginal minus text-only word-count CRPS +2.298139 +0.023968 to +4.572310 Text only
Marginal minus demographics-only word-count CRPS +0.200711 −0.811325 to +1.212747 Inconclusive
Text-only minus combined word-count CRPS +0.217704 −0.084469 to +0.519876 Inconclusive
Demographics-only minus combined word-count CRPS +2.315132 +0.011811 to +4.618453 Combined
Marginal minus source-form word-count CRPS +3.245309 +1.000674 to +5.489943 Source form
Source-form minus text-only word-count CRPS −0.947170 −2.449295 to +0.554956 Inconclusive

The adjustment preserves the central response-length interpretation. Both the point and probabilistic word-count gains remain directionally stable, as do text-only and source-form improvements over the marginal. The Add-count point gain also remains. In contrast, the smaller Delete-count point effect and the increment from adding demographics to text no longer exclude zero. Thus the strongest person-matching evidence concerns Add volume and later response length; the earlier wording that all three point outcomes improved should be read as the pointwise pattern, not as a family-robust conclusion.

This audit does not make the six remaining directions confirmatory. The family was constructed after earlier results were known; the same 150 selected cases support every contrast; fitted folds overlap; and resampling stored case scores does not refit the models. It is a conservative internal coherence check, not population inference or a remedy for sequential analysis.

6.15 Calibrated synthesis does not clear the content gate

The full-text synthesis changes the kind of error without producing a general agreement gain. It is much closer to the observed amount of source change and has lower word-count error than repeat-2024, but the common appended units substantially inflate the number of lines and dilute token agreement.

Complete official scored metrics for calibrated training synthesis
Training leave-one-out metric Calibrated synthesis Stable projection Repeat 2024 Synthesis utility effect versus repeat (95% interval)
Exact match 0.000000 0.000000 0.000000 0.000000 (0.000000 to 0.000000)
Normalized edit similarity 0.268883 0.276809 0.272662 −0.003779 (−0.013404 to +0.004287)
Token Jaccard 0.098165 0.133873 0.134297 −0.036132 (−0.048502 to −0.024014)
Token-overlap F1 0.267251 0.276737 0.280231 −0.012980 (−0.035902 to +0.010370)
ROUGE-L F1 0.220153 0.200304 0.200017 +0.020136 (+0.002677 to +0.037571)
Character n-gram F1 0.238779 0.266194 0.267813 −0.029034 (−0.040378 to −0.017567)
Word-count MAE 51.986667 46.986667 55.580000 +3.593333 (+1.440000 to +5.726667)
Line-count MAE 27.053333 9.066667 6.160000 −20.893333 (−23.820000 to −17.966667)
Source-similarity MAE 0.143463 0.729214 0.799983 +0.656521 (+0.627361 to +0.684716)

Positive effects favor synthesis for every row. The ROUGE-L and source- similarity results show that the generator approximates the degree and sequence of change better than continuity baselines; they do not show correct new identity content. Relative to the stable projection, synthesis also improves ROUGE-L by +0.019849 and source-similarity error by +0.585751, but worsens token Jaccard, character n-gram F1, and line-count error with intervals excluding zero. Its word-count error is five words higher on average than the stable projection, with an interval spanning zero.

The equal-volume content comparison is unambiguously unfavorable:

Common-unit synthesis versus equal-volume marginal additions
Distinct novel-type measure Calibrated synthesis Volume-matched marginal Paired difference (95% interval)
Precision 0.097449 0.208189 −0.110740 (−0.139732 to −0.082632)
Recall 0.074294 0.116041 −0.041748 (−0.058335 to −0.025062)
F1 0.075887 0.134680 −0.058793 (−0.076168 to −0.041473)

Synthesis wins 33 cases, ties 21, and loses 96 on novel-type F1. Only the word- count gate passes; the edit-similarity interval spans zero and the content gate excludes zero in the wrong direction. The prespecified advancement rule therefore fails, and no development prediction was generated.

6.16 Frozen test submission

The method generated one nonblank prediction for every test ID in the required order. The pinned official validator returned VALID: 81 predictions. The frozen CSV has SHA-256:

a463d9e314069357f050c9f2270acfad59165db2c0bab19322d46517444d9ab3

The organizer must run the private test evaluator. Until that happens, the private-test performance remains unknown; neither public-development results nor the training-only ranking diagnostic substitutes for it.

7 Discussion

The stable-signifier projection does learn something useful about response form and degree of change. Yet improved prediction of how much a response will change is not the same as predicting what new identity content will appear. The novelty diagnostic makes that distinction measurable: most future lexical material lies outside the earlier response, while the method is definitionally restricted to it. A next-generation model needs a mechanism for both deletion and creation, evaluated without sacrificing the interpretability of these paired diagnostics.

Leave-one-out cross-validation makes this conclusion less dependent on the 50 reused development cases. Its larger training-cohort analysis reproduces all three directional error results and the overall lack of a uniform agreement gain. The projection can adjust length and continuity without solving content: its small edit-similarity gain coexists with a small token-overlap loss and extreme overprediction of source continuity.

Matched-trajectory retrieval sharpens this inference. It supplies deletion and creation in approximately the observed proportions, but the transferred novel content is rarely person-correct. The problem is not merely calibrating how different the next self-description will be. A useful forecast must condition new signifiers on the individual without collapsing back into verbatim continuity.

The marginal control tightens that conclusion. Retrieval not only recovers little new vocabulary; it recovers less than a person-agnostic training prior given the same number of guesses. The diagnostic also disciplines the language of the report: token recovery is not automatically signifier recovery, because many high-base-rate additions are connective or generic words. Future person-specific models need to demonstrate value beyond both continuity and marginal lexical prevalence.

The two leave-one-out conditioned results make that requirement more specific. Merely finding the strongest smoothed source-token-to-Add association is not enough, and replacing that maximum with a strongly regularized 30-neighbor ensemble still sacrifices 1.20 percentage points relative to marginal prevalence. Surface lexical and coarse demographic similarity have now failed in both single-trajectory and pooled forms. Better individualization will need richer context, a semantic mechanism beyond document-level GloVe centroids, or new longitudinal evidence; it should establish an incremental training-only advantage before reopening the heavily reused development comparison.

The external-semantic test makes that requirement empirical rather than rhetorical. Its pretrained vectors cover almost all source vocabulary and replace exact lexical overlap with distributional proximity, yet the recovered fraction is virtually identical to the surface neighborhood and remains 1.22 percentage points below marginal prevalence. Semantic representation alone is not semantic forecasting: averaging word vectors can blur precisely the specific roles, relationships, and attributes that matter to personally expressed identity. A future attempt needs either richer compositional context or new evidence, not another fixed re-ranking of these Add candidates.

The prospective volume analysis qualifies that negative content result. At the contrast-by-contrast level, the fixed neighborhood improves three no- oracle quantity forecasts even while its oracle-budget lexical ranking loses to marginal additions. The cross-analysis stress test sharpens that claim: Add-count and word-count point gains remain simultaneously stable, while the smaller Delete-count gain becomes inconclusive. Person conditioning therefore is not uniformly uninformative, but its family-robust quantity signal is narrower than the original pointwise pattern. A synthesizing model should preserve that calibration signal while demonstrating content value beyond marginal prevalence.

The calibrated-synthesis test directly attempts that composition and shows why the two requirements must be evaluated together. Common complete response units make predicted text look appropriately different from its source and improve sequence overlap, but they recover fewer correct novel types than a token-frequency prior and create a gross line-count mismatch. Forecasting the amount of novelty is not a license to fill that budget with generic statements. The failed advancement gate prevents another look at development labels and raises the next-method bar from “synthesize something” to “add source-relevant semantic content that beats both continuity and marginal prevalence.”

The probabilistic extension further qualifies the claim: the clearest distributional gain is for follow-up word count. The ablation locates most of that signal in the earlier self-description rather than demographics, and the source-form comparison prevents a stronger lexical interpretation: three simple source-response counts improve on the marginal and are not stably distinguishable from text-only matching. The multiplicity audit retains both text-only and source-form gains over the marginal but not demographics’ small increment over text. Demographics alone do not improve the marginal forecast, and demographic resemblance is not an ipseological mechanism. None of these surface predictors establishes such a mechanism.

7.1 Reproducibility

The complete open materials are:

All Project research scripts use only the Python standard library. The challenge’s 17 public tests passed before the stable artifact was generated. Regenerating the stable development and test artifacts reproduced their hashes exactly; the retrieval pipeline separately hash-guards its inputs and fixed analysis plan.

7.2 Sources

7.3 Limitations and next step

An extractive method cannot predict genuinely new identities, experiences, or reframing. Token recurrence is not equivalent to identity-signifier endurance, and common wording can make a signifier look stable. The line splitter also turns some prose sentences into separate output lines, worsening line-count error. Finally, public-development-guided selection can overfit a 50-case set.

The stable-projection cross-validation withholds each case’s future from its own model fit, but it is not untouched confirmation. The training corpus had already informed the method, and cross-validation folds overlap heavily. Its case-bootstrap intervals describe this selected training cohort rather than a population. The published fold predictions are a derived benchmark-data adaptation under CC BY-NC-SA 4.0.

Trajectory retrieval adds different limitations. It transfers another participant’s public response rather than generating a person-specific future, and demographic similarity does not establish that identity change follows demographic categories. Its within-iteration lock reduces new analytic flexibility but cannot make previously inspected development data unseen.

The marginal-addition control is post hoc and deliberately narrow. It uses retrieval’s case-level novel-token count, so it is not an independent full-text forecast; it isolates content choice after holding opportunity volume equal. Its ranking is estimated from only 150 training pairs, and its lexical units include function words that are not necessarily identity signifiers. The derived token audit adapts the benchmark data and remains governed by the benchmark’s CC BY-NC-SA 4.0 license.

The source-conditioned comparison is locked and uses leave-one-out estimation, but it remains a training-cohort diagnostic with an oracle future-token budget. Its maximum-over-source-tokens rule can privilege sparse associations, the ten-case smoothing strength was adopted rather than tuned, and its tokens need not be semantic signifiers. The case audit is also a derived benchmark-data adaptation under CC BY-NC-SA 4.0.

The regularized-neighborhood comparison shares the oracle-budget and lexical identity-signifier limitations. Its 30-neighbor size and equal 30-case marginal prior were fixed without tuning, while its text-and-demographic representation was selected in earlier Project work. Coarse demographic similarity is not an ipseological mechanism. Its case audit is likewise a derived benchmark-data adaptation under CC BY-NC-SA 4.0.

The semantic-neighborhood comparison inherits the same oracle budget, candidate vocabulary, neighborhood size, and shrinkage. GloVe proximity is distributional rather than a validated ipseological identity measure; the centroid erases order, senses, negation, and response structure. Twitter- trained vectors can encode social bias and need not transfer cleanly to Twenty Statements Test language. The analysis was chosen after earlier aggregate results were known, and its failed hurdle rejects only this fixed method. Its case audit is a derived benchmark-data adaptation under CC BY-NC-SA 4.0; the public-domain vector file is hash-pinned but not redistributed here.

The change-volume comparison removes that oracle budget, but its surface-text and coarse-demographic representation and its 30-neighbor, equal-shrinkage choices came from prior Project work rather than untouched selection. Its overlapping folds and case-bootstrap intervals characterize this selected training cohort, not a population. Count forecasts remain continuous to avoid an arbitrary rounding rule, Add/Delete tokens need not all be identity signifiers, and the derived case audit is governed by CC BY-NC-SA 4.0.

The probabilistic extension inherits all of those design constraints. Its weighted empirical distributions reuse the same outcomes, folds, representation, and fixed weights; they are not independently selected models. Central 80% intervals are discrete empirical quantiles, and coverage in 150 overlapping leave-one-out folds is descriptive rather than a population guarantee. Its derived case audit is also governed by CC BY-NC-SA 4.0.

The source-feature ablation is a dependent extension chosen after the combined word-count gain was known. It preserves the existing exact-value demographic features, so sparse categories and missing values can weaken demographics-only similarity. Its unadjusted secondary contrasts are repeated-analysis diagnostics, not an independent family of confirmatory tests. The derived case audit remains governed by CC BY-NC-SA 4.0.

The source-form comparison is another dependent extension chosen after the text-only word-count gain was known. Its three counts and fixed distance rule test one narrow alternative, not every nonlexical representation. A direct interval spanning zero is evidence that this analysis does not distinguish the methods, not proof that their predictive distributions are equivalent. Its unadjusted secondary contrasts are repeated-analysis diagnostics, and its case audit remains governed by CC BY-NC-SA 4.0.

The multiplicity stress test is itself post hoc: its family was fixed only after the constituent analyses and aggregate results existed. Synchronized case resampling preserves empirical dependence among the 16 stored effects, but it does not refit the overlapping leave-one-out folds, undo sequential research choices, or turn a selected cohort into a probability sample. Its simultaneous intervals are therefore an internal case-composition sensitivity check rather than formal prospective familywise error control or population inference. The flat audit is derived from benchmark-data adaptations and remains governed by CC BY-NC-SA 4.0.

The calibrated-synthesis analysis is also dependent: every component and its three-part gate were chosen after the earlier training results were known. Exact repeated response units favor generic template language, and appending one statement per unit explains the severe line-count error. Its fold predictions transfer public training text and remain benchmark-data adaptations under CC BY-NC-SA 4.0. The pointwise gate is deliberately conservative but is not multiplicity-adjusted; failing it is a method decision, not a universal claim that semantic synthesis is impossible.

The paired bootstrap treats development cases as the resampling unit, but the benchmark is a selected longitudinal cohort rather than a probability sample. Its intervals describe sensitivity to the composition of these 50 cases, not uncertainty for people in general. The lexical oracles use observed follow-ups and are deliberately unattainable; they must not be read as test predictions or achievable prospective scores.

The immediate next step is to submit the two prepared artifacts to the challenge organizer and add the complete private test scorecard, unchanged, to this report. The fixed common-unit synthesis and GloVe-centroid ranking have both failed their training hurdles, so neither may be evaluated on development data. Further re-rankers on these same 150 cases are unlikely to be informative without richer compositional context or new longitudinal evidence. A later method should preserve response-form calibration, beat leave-one-out marginal additions in training, lock any later development comparison, and avoid private test feedback for tuning.