首页 > AI前沿 > CDMD: A Cross-Dataset Mixed-Type Diffusion Model for Tabular Data

CDMD: A Cross-Dataset Mixed-Type Diffusion Model for Tabular Data

arXiv机器学习 2026-09-30 14:56 7 阅读 查看原文

Generative models for tabular data are typically trained separately for each dataset, limiting knowledge transfer and requiring the storage of many specialized models.

In this paper, we introduce CDMD, a tabular diffusion model trained jointly across heterogeneous datasets with different schemas and variable numbers of numerical and categorical features.

Unlike existing cross-dataset tabular diffusion models that operate in continuous representation spaces, CDMD defines diffusion directly over the mixed-type feature space and is trained end-to-end.

To accommodate heterogeneous categorical domains, we introduce a schema-restricted reverse-process parameterization for masked diffusion models, in which the output space dynamically adapts to each feature's vocabulary.

We then compose numerical and categorical feature-level diffusion processes into a schema-dependent row-level process.

A shared schema-aware Transformer denoiser captures dependencies between features and parameterizes the reverse process across varying schemas.

On seven real-world datasets, a single jointly trained CDMD achieves the highest average generation quality among strong single-dataset and cross-dataset baselines, while using substantially fewer total parameters than the collection of separately trained models.

Furthermore, pre-training on a corpus of 337 datasets improves generation on previously unseen datasets under both limited target data and limited adaptation epochs.

These results demonstrate the potential of direct mixed-type diffusion for shared and transferable tabular data generation.

Our code is available at https://github.com/ketatam/cdmd.