Inverse design in engineering often runs into a simple problem.
Each labeled training sample must be produced through expensive simulation, so building a large dataset is slow and costly.
This study addresses that problem for partial inverse design
where only some design variables are specified and the rest must be inferred to reach a target performance value.
We propose CoNN-AL, a framework for data-efficient partial inverse design that adds stream-based active learning to the Cooperative Neural Network with Denoising Autoencoder (CoNN-DAE).
The model estimates predictive uncertainty through Monte Carlo dropout and uses it to decide, in real time, which incoming candidate samples are worth labeling, so the limited labeling budget is spent on the most informative designs.
We validate the framework on a real-world automotive glass run channel dataset
of more than 900,000 unique simulated designs.
With only 20,000 actively selected labels, about 2.3% of the training pool, CoNN-AL reaches R-squared values of 0.967 to 0.982 across all missing-variable levels, approaching the upper-bound models trained on far more data.
It reaches R-squared of at least 0.95 with 30 to 40% fewer labels than random sampling at the more difficult missing-variable levels and, at the most challenging level, is the only strategy in this study to reach R-squared of 0.98.
Together with this work, we publicly release the dataset to support future research on data-driven design.