Évènements passés
lun
mar
mer
jeu
ven
sam
dim
l
m
m
j
v
s
d
27
28
29
1
2
3
5
6
7
8
9
10
11
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
Room 109, Building 9, St Eloi campus
Machine Learning in Montpellier, Theory & Practice
Biodiversity is under severe pressure, as many different disturbance events threaten terrestrial and marine ecosystems with varying impacts. Therefore, habitat distribution modelling, which aims to quantify the statistical links between environmental covariates and an habitat’s occurrence, is increasingly relevant. Herein, we present two different approaches to guide investment, management and regulatory decisions. Firstly, a framework based on tabular data, which experiments with different network architectures, feature encodings, hyperparameter tuning and noise addition strategies to identify the optimal model for habitat classification based on plant species composition. Secondly, we introduce Pl@ntBERT, which leverages sophisticated natural language processes based on transformers (i.e., models with attention components able to learn contextual relations between categorical and numerical features). In particular, since they reinforce each other, the pipeline makes use of both masked language modelling and text classification. The first step helps to get a statistical understanding of the plant species composition (the language in which the model is trained in). Then, subsequent training is used to assign an habitat type to sentences describing vegetation plots. The fine-tuning of a pretrained foundation model on in-domain data shows significant upgrade. Notably, it clearly outperforms previous state-of-the-art methods by pushing the accuracy score on a large database containing millions of European samples. Finally, our results showcase that flora is a strong marker of habitat type and doesn't need to be coupled with environmental spatial data to train neural networks with high predictive power. Looking forward to seeing you.
Machine Learning in Montpellier, Theory & Practice
Room 02.124, Building 5, St Priest campus
Machine Learning in Montpellier, Theory & Practice
We consider a Multi-Armed Bandit problem with covering constraints, where the primary goal is to ensure that each arm receives a minimum expected reward while maximizing the total cumulative reward. In this scenario, the optimal policy then belongs to some unknown feasible set. Unlike much of the existing literature, we do not assume the presence of a safe policy or a feasibility margin, which hinders the exclusive use of conservative approaches. Consequently, we propose and analyze an algorithm that switches between pessimism and optimism in the face of uncertainty. We prove both precise problem-dependent and problem-independent bounds, demonstrating that our algorithm achieves the best of the two approaches– depending on the presence or absence of a feasibility margin – in terms of constraint violation guarantees. Furthermore, our results indicate that playing greedily on the constraints actually outperforms pessimism when considering long-term violations rather than violations on a per-round basis. Le jeudi 30 novembre 2023 à 10:22:44 UTC+1, a écrit : Dear all, On Monday, December 4 , at 2 pm Paris time we will have the following talk: - Dorian Baudry will give a talk entitled Multi-armed bandits with guaranteed revenue per arm You can join the talk on this link: [ https://isdm.umontpellier.fr/inria.webex.com/inria/j.php?MTID=m0d098926b482a7f3e957eab08dc706d3 | https://isdm.umontpellier.fr/inria.webex.com/inria/j.php?MTID=m0d098926b482a7f3e957eab08dc706d3 ]
Machine Learning in Montpellier, Theory & Practice
Room 02.124, Building 5, St Priest campus
Machine Learning in Montpellier, Theory & Practice
Recent AI systems make use of machine-learning algorithms, where the computing system learns from data (observations) in order to adjust its behavior. In fact, the availability of huge amounts of data has been the key enabler of modern AI-based technologies. This dependency on data, however, is also the Achille's heel of AI systems. The data, which can come from a wide variety of sources, is not always trustworthy. Some sources can provide erroneous or corrupted data. With current machine-learning algorithms, a single "bad"" source can lead the entire learning scheme to make critical mistakes. Moreover, to handle the huge amounts of data, machine-learning algorithms are often deployed over a large number of computing machines. Consequently, as this number increases, the likelihood of machine errors also increases. Indeed, software and hardware bugs are prevalent. Furthermore, machines can sometimes be hacked by malicious players, either directly or indirectly through viruses. Some of these players attempt to corrupt the entire learning procedure, merely for the pleasure of claiming to have destroyed an important system. Others attempt to influence the learning procedure for their own benefit. Building machine-learning schemes that are robust to these events is paramount to transitioning AI from being a mere spectacle capable of momentary feats to a dependable tool with guaranteed safety. In this talk, we cover some effective techniques for achieving such robustness. We present machine-learning algorithms that do not trust any individual data source or computing unit. We do assume, however, that a majority of data sources and machines are trustworthy; otherwise, no meaningful learning guarantee can be provided. The challenge arises from the fact that the identity of the trusted data sources and machines is a priori unknown."
Machine Learning in Montpellier, Theory & Practice