Model optimizations help improve inference performance and accuracy of ML workflows.
However, relying on a single model to perform inference across all data batches often fails to maximize accuracy and thus overall performance.
In many cases, alternate models could perform better on specific subsets of data where a primary model underperforms.
Our experiments with real ML workflows indeed show that switching models improves workflow accuracy by up to 23%.
Yet, current systems lack the ability to adaptively switch between models based on performance, forcing users to manually test models in sequence.
We present FlexiFlow
FlexiFlow, a dataflow system that dynamically switches between alternate models when the current model exhibits low accuracy.
FlexiFlow learns to rank models using a novel multi-armed bandit approach that accounts for model runtimes, probability of passing user-defined assertions, and the computational structure of the ML workflow.
Standard Thompson sampling approach
We show that the standard Thompson sampling approach is insufficient for switching models in ML workflows.
In contrast, our proposed approaches are effective and scales to complex real-world ML workflows.
Experiments show that switching models at runtime while reusing intermediate results provides higher accuracy, but also 48% efficiency gain compared to sequential workflow runs.