首页 > AI前沿 > SPD-MetaFormer is what you need for small-data brain decoding

SPD-MetaFormer is what you need for small-data brain decoding

arXiv机器学习 2026-10-08 06:09 4 阅读 查看原文

Brain signal decoding is challenging because neural recordings are noisy and vary across individuals, while labeled data are often limited.

Recent attention-based models on the symmetric positive definite (SPD) manifold have nevertheless achieved strong performance using covariance and connectivity representations, yet the contribution of learned token weighting remains unclear.

We examine two representative architectures

MAtt (based on log-Euclidean geometry) and GBWAtt (based on generalized Bures--Wasserstein geometry), and find that their learned attention weights remain close to uniform after training.

We relate this behavior to bounded similarity parameterizations that, under the original softmax scaling, limit attention-weight contrast.

Moreover, replacing learned weights with uniform weights, throughout training and evaluation, has little effect on mean predictive performance while preserving each model's original aggregation geometry.

Motivated by these findings

we introduce SPD-MetaFormer, an attention-free architecture built on uniformly weighted Fréchet aggregation under log-Euclidean geometry.

Its backbone uses a geodesic residual to update a summary token and a shared spectral feedforward map to transform all tokens, followed by a learned weighted readout.

Token states remain SPD-valued until tangent-space classification.

Across three EEG benchmarks

SPD-MetaFormer achieves competitive results relative to published Euclidean and manifold baselines.

Separate matched reproductions test learned versus uniform weighting within MAtt and GBWAtt

These results suggest that, in the short-sequence and limited-data regimes studied, carefully designed SPD architectures can provide a simpler and effective alternative to adaptive manifold attention.