Quand
Où
Montpellier
Machine Learning in Montpellier, Theory & Practice
Biodiversity is under severe pressure, as many different disturbance events threaten terrestrial and marine ecosystems with varying impacts. Therefore, habitat distribution modelling, which aims to quantify the statistical links between environmental covariates and an habitat’s occurrence, is increasingly relevant. Herein, we present two different approaches to guide investment, management and regulatory decisions. Firstly, a framework based on tabular data, which experiments with different network architectures, feature encodings, hyperparameter tuning and noise addition strategies to identify the optimal model for habitat classification based on plant species composition. Secondly, we introduce Pl@ntBERT, which leverages sophisticated natural language processes based on transformers (i.e., models with attention components able to learn contextual relations between categorical and numerical features). In particular, since they reinforce each other, the pipeline makes use of both masked language modelling and text classification. The first step helps to get a statistical understanding of the plant species composition (the language in which the model is trained in). Then, subsequent training is used to assign an habitat type to sentences describing vegetation plots. The fine-tuning of a pretrained foundation model on in-domain data shows significant upgrade. Notably, it clearly outperforms previous state-of-the-art methods by pushing the accuracy score on a large database containing millions of European samples. Finally, our results showcase that flora is a strong marker of habitat type and doesn’t need to be coupled with environmental spatial data to train neural networks with high predictive power. Looking forward to seeing you.
