Spotting Expressivity Bottlenecks and Fixing Them Optimally

28 septembre 2023 @ 14 h 00 min –

Camelot Workshop, St Priest campus

Machine Learning in Montpellier, Theory & Practice

Reject option is a well established technique to abstain from prediction when the doubt in the label is too big or because of ressources limits. In this talk I will first present the problem of learning with reject option in its more classical use. I will then present two recent applications: the first one focuses on building an efficient active learning algorithm that uses rejection option learning arguments to iteratively construct the uncertain region; the second one deals with reducing energy consumption during inference with a deep neural network. In the latter application, we have built networks with early exits, and it appears that the reject option is particularly useful for calibrating thresholds for early existing. 3:30 PM – 4:00 PM *Coffee Break (2nd floor) 4 PM – 4:45 PM *Mathilde Mougeot, ENSIIE & ENS Paris-Saclay* Title: *Leveraging knowledge to design machine learning despite the lack of data. Abstract: In recent years, considerable progress has been made in the implementation of decision support procedures based on machine learning methods through the exploitation of very large databases and the use of learning algorithms. In the industrial environment, the databases available in research and development or in production are rarely so voluminous and the question arises as to whether in this context it is reasonable to use machine learning methods. This talk presents research work around transfer learning and hybrid models that use knowledge from related application domains or physics to implement efficient models with an economy of data. Several achievements in industrial collaborations will be presented that successfully use these learning models to design machine learning for industrial small data regimes and to develop powerful decision support tools even in cases where the initial data volume is limited. References – de Mathelin, A., Deheeger, F., Mougeot, M., Vayatis, N. (2023) From Theoretical to Practical Transfer Learning: the ADAPT library, Federated and Transfer Learning Springer book. – Nguyen, Khoa.T.N, Dairay, T., Meunier, R., Mougeot, M. (2023) Fixed-Budget Online Adaptive Learning for Physics-Informed Neural Networks. Towards Parameterized Problem Inference, Lecture Notes in Computer Science, Springer volume 14073. 4:45 PM – 5:30 PM *Stephane Chretien, University of Lyon 2* Title: *Relationship between sample size and architecture for the estimation of Sobolev functions using deep neural networks* Abstract: Beyond the many successes of Deep Learning based techniques in various branches of data analytics, medicine, business, engineering and the human sciences, a sound understanding of the generalisation properties of these techniques is still elusive. Central to these successes are the availability of huge datasets and the availability of huge computational ressources and some of the most recent trends have given paramount importance to the necessity of building huge neural networks with millions of parameters, and most often, of several orders of magnitude larger than the size of the training set. This set-up has however led to many surprises and counterintuitive discoveries. Overprametrisation was recently shown to favour connectivity in a weak sense of the set of stationary points, hence permitting stochastic gradient type methods to potentially reach good minimisers in several stages despite the wild nonconvexity of the training problem as demonstrated by Kuditipudi et al. Relating generalisation to stability, recent theoretical breakthroughs have been able to provide a better understanding of why generalisation cannot even happen without overparametrisation as shown by Bubeck et al. Following the ideas developed by Belkin, a substantial amount of work has also been undertaken in order to study the double descent phenomenon, and the associated benign overfitting property which holds for least norm estimators in linear and mildly non-linear regression, as well as in for certain kernel based methods. In the present paper, we aim at studying the generalisation properties of overparametrised deep neural networks using a novel approach based on Neuberger’s theorem.

Machine Learning in Montpellier, Theory & Practice