首页 > AI前沿 > TICDA: Tabular In-Context Data Attribution

TICDA: Tabular In-Context Data Attribution

arXiv机器学习 2026-10-06 16:55 5 阅读 查看原文

Tabular foundation models (TFMs) achieve strong predictive performance by conditioning on labeled demonstrations provided in context, without any parameter update.

Yet how individual demonstrations shape a given prediction remains poorly understood.

This gap matters in practice: the context is often assembled from whatever labeled data is available, potentially leading to the inclusion of mislabeled, redundant, or low-quality examples that degrade performance.

Standard data attribution methods do not transfer to the TFM setting: resampling-based approaches such as DemoShapley require a combinatorial number of forward passes, and gradient-based estimators such as influence functions require computing training point's effect on the model parameters, which in-context learning never updates.

We introduce TICDA

We introduce TICDA, a method that measures the influence of every demonstration in the context directly from linear surrogates trained on TFM latent embeddings, in a single forward pass and at negligible cost.

Performance Across Tasks

  • detecting labeling errors
  • curating context to preserve predictive accuracy while lowering inference cost
  • producing attribution scores that transfer across TFMs
  • supporting an acquisition strategy for efficient active learning