lun
mar
mer
jeu
ven
sam
dim
l
m
m
j
v
s
d
30
1
3
4
5
6
7
8
9
10
11
12
13
14
15
17
18
19
20
22
24
25
26
27
28
29
30
31
1
2
Interpretable machine learning models for predicting with missing values
Room 01.124, Building 5, St Priest campus Machine Learning in Montpellier, Theory & Practice Machine learning models are frequently used when inputs are missing during training or prediction, potentially leading to increased bias or impractical models without imputing unobserved variables. Imputing missing values is often inadequate and hard to interpret, especially with complex functions. This talk addresses the challenge of predicting with missing data at test time, highlighting the need for interpretable and practical models, crucial in critical sectors like healthcare. In this talk, I present two novel approaches: the Shared Pattern Sparsity Model (SPSM) for scenarios with recurrent missing data patterns, which promotes efficient data use and interpretability without reliance on imputation; and MINTY, a sparse linear rule model, regularized to minimize dependence on features with missing values. This model allows a trade-off between goodness of fit, interpretability, and robustness to missing values at test time. Additionally, I'll share early results from a project developing a predictive model for sequential risk scores sensitive to missing values across time steps. In collaboration with the Traumabase network, we'll conduct a user study with clinical professionals to assess the effectiveness of interpretable models in handling missing data. Machine Learning in Montpellier, Theory & Practice
Active Clustering with bandit feedback
Room 109, IMAG, Triolet campus Machine Learning in Montpellier, Theory & Practice We will present the recent Active Clustering Problem (ACP). In this problem, a set of items can be partitioned into groups where items within the same group are characterised by the same multi-dimensional vector. A learner obtains noisy observations of these vectors, and we consider an active setting where the learner chooses the order and the number of observations. The objective is to recover the hidden partition of the items, using as few requests as possible.In the presentation, I will explain the ACP and answer two questions. Can we improve upon the number of requests of the simple uniform sampling algorithm, using the benefits of active sampling ? Is there a fundamental computation-information gap for clustering in high-dimension with repeated measurements? Machine Learning in Montpellier, Theory & Practice
No Bluffing: Proving ML Model Trustworthiness
Room 02.124, Building 5, St Priest campus Machine Learning in Montpellier, Theory & Practice Over the past few years, we have seen significant efforts in building trustworthy ML. Many institutions are making privacy and fairness promises about their services. I start the presentation by arguing that institutions might not adhere to their claims of using trustworthy ML across various services intentionally or accidentally due to their interests in maximizing utility and minimizing costs/efforts, resulting in FairWashing /[NeurIPS2022]/ and PrivacyWashing. To address these risks, then, I present Confidential-PROFITT /[Oral ICLR2023] /and Confidential-DPproof /[Spotlight ICLR2024]/ frameworks that enable institutions to directly prove to any interested party through the execution of Zero Knowledge Proof protocols that they train ML models in a fair and privacy-preserving manner, respectively, while protecting the confidentiality of their data and model. [NeurIPS2022] Washing The Unwashable : On The (Im)possibility of Fairwashing Detection, https://openreview.net/pdf?id=3vmKQUctNy [Oral ICLR2023] Confidential-PROFITT: Confidential PROof of FaIr Training of Trees, https://openreview.net/pdf?id=iIfDQVyuFD [Spotlight ICLR2024] Confidential-DPproof: Confidential Proof of Differentially Private Training, https://openreview.net/pdf?id=PQY2v6VtGe Machine Learning in Montpellier, Theory & Practice
Multiply robust off-policy evaluation and learning under truncation by death
Room 01.124, Building 5, St Priest campus Machine Learning in Montpellier, Theory & Practice Typical off-policy evaluation (OPE) and off-policy learning (OPL) are not well-defined problems under "truncation by death", where the outcomeof interest is not defined after some events, such as death. The standard OPE no longer yields consistent estimators, and the standard OPL results in suboptimal policies. In this paper, we formulate OPE and OPL using principal stratification under "truncation by death". We propose a survivor value function for a subpopulation whose outcomes are always defined regardless of treatment conditions. We establish a novel identification strategy under principal ignorability, and derive the semiparametric efficiency bound of an OPE estimator. Then, we propose multiply robust estimators for OPE and OPL. We show that the proposed estimators are consistent and asymptotically normal even with flexible semi/nonparametric models for nuisance functions approximation. Moreover, under mild rate conditions of nuisance functions approximation, the estimators achieve the semiparametric efficiency bound. Finally, we conduct experiments to demonstrate the empirical performance of the proposed estimators. Machine Learning in Montpellier, Theory & Practice
Probabilistic graphical models and deep neural networks for remote sensing image analysis
Room 02.124, Building 5, St Priest campus Machine Learning in Montpellier, Theory & Practice Given the current advances in space missions for Earth observation, it is possible to have access to very-high-resolution and multimodal satellite imagery. The data acquired can be optical (e.g., panchromatic, multispectral, and hyperspectral images) or radar, with different synthetic aperture and various trade-offs between resolution and coverage. This offers great application potential in the field of remote sensing. An important role in this context is played by semantic segmentation whose purpose is to assign each pixel in an image to a semantic class, typically related to land cover or land use and with prominent applications in areas such as urban planning, precision agriculture, monitoring of forest species, natural disaster management, and climate change monitoring and mitigation. This presentation focuses on novel methods for the analysis of multimodal data aimed at fully exploiting all the available information, combining ideas from stochastic models and deep learning. On the one hand, deep learning is currently the dominant approach to image classification and segmentation. However, the performances of deep learning methods are remarkably influenced by the quantity and quality of the ground truth used for training. On the other hand, probabilistic graphical models have sparked major interest in the past few years, because of the ever-growing need for structured predictions. Depending on the underlying graph topology over which they are defined, they can effectively model spatial and multiresolution information. The idea is to develop approaches leveraging the advantages of these two major methodological families for the exploitation of multimodal remote sensing data and of the complementary information they convey. The experimental validations, conducted with multimodal multispectral, panchromatic, and radar satellite images, suggest the effectiveness of the proposed methods. Best, Cassio Machine Learning in Montpellier, Theory & Practice
Évènements du 29 avril 2024
Évènements du 2 mai 2024
Évènements du 16 mai 2024
Évènements du 21 mai 2024
Aucun évènement à afficher.
Aucun évènement à afficher.