Quand
Où
860 rue St Priest, Montpellier
Machine Learning in Montpellier, Theory & Practice
The use of deep neural networks in critical domains, such as aviation, is limited by our ability to ensure their correct behavior. Execution monitors are components aimed at identifying dangerous predictions and discarding them before they lead to catastrophic consequences. Several recent works on real-time monitoring have focused on detecting out-of-distribution (OOD) inputs, i.e., identifying inputs that differ from the training data. In this presentation, we will show that OOD detection is not a well-suited framework for designing effective execution monitors and that it is more relevant to evaluate monitors based on their ability to discard incorrect predictions. We call this paradigm out-of-model-scope (OMS) detection and discuss the conceptual differences with OOD. We will also present in-depth experiments to show that studying monitors in the OOD framework can be misleading: 1. very good OOD results can give a false sense of security, 2. an OOD-based comparison may not identify the best monitor for detecting errors.
