Networks that behave alike now can still learn differently when training continues.
Work on loss of plasticity and critical periods shows that the path to a state shapes what follows; it does not show whether the path carries information that a measurement of the state itself misses.
We ask when the training history of a network predicts its future learning better than its current state.
In a main study
small multilayer perceptrons were trained under three history regimes (42 histories), and future learning was measured at four checkpoints by a short probe: a copy of the network trained for 100 updates on a new task.
Before the prediction result was read, the protocol checked the probe.
It responded monotonically to a function-preserving rescaling of hidden units, repeated measurements agreed (intraclass correlation 0.940, [0.903, 0.997], in the least reliable class, mean of three repeats), and a re-initialisation of units was visible directly after it but not 100 to 200 updates later.
A history state of at most four dimensions did not improve on a calibrated model of the current state (gain -21.4%, 90% interval [-91.9, 8.1]; required in advance: 10%)
A companion screen on 1,560 synthetic regression runs
asked the same question for a target further away, the final error of the run.
There, history models forecast better than the current validation error after 12 of up to 240 epochs (compact state 30.3%, [15.8, 39.4], a contextual comparison) and were not distinguishable from it after 48.
In both studies the history was informative only while the current state was not yet informative about the target; this reading was formed after the results.