Recursive feature machines (RFMs) learn representations of data by alternating between fitting a predictor to a dataset and updating features of that predictor using the average gradient outer product (AGOP).
Connections between AGOPs and feature learning in neural networks motivate linear RFMs as a simple setting for analyzing how representations evolve during training.
Our Study
Here, we study the dynamics and statistics of linear RFM in noisy multi-output regression with isotropic sub-Gaussian input data and targets generated by a low-rank teacher matrix of dimension $d$.
We extend the known connection between linear RFM and iteratively reweighted least squares from the interpolating setting to ridge-regularized multi-output regression with noise.
We show that the learned feature matrix remains close to its infinite-data ideal counterpart at every iteration.
Namely, for $n$ samples, we show the error in the feature matrix decays as $O(\sqrt{d/n})$ with high probability.
Experiments
Experiments on real-world text and single-cell gene-expression data illustrate the features learned by this simple linear model.