BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//wp-events-plugin.com//7.4.0.1//EN
TZID:Europe/Paris
X-WR-TIMEZONE:Europe/Paris
BEGIN:VEVENT
UID:200@isdm.umontpellier.fr
DTSTART;TZID=Europe/Paris:20251002T140000
DTEND;TZID=Europe/Paris:20251002T140000
DTSTAMP:20260825T131907Z
URL:https://isdm.umontpellier.fr/events/advances-in-pure-exploration-in-ba
 ndits-non-asymptotic-private-and-structured/
SUMMARY:Advances in Pure Exploration in Bandits: Non-Asymptotic\, Private\,
  and Structured
DESCRIPTION:Room 02.124\, Building 5\, St Priest campus\n\nMachine Learning
  in Montpellier\, Theory &amp\; Practice - Marc Jourdan (EPFL)\n\nIn pure 
 exploration problems for stochastic multi-armed bandits\, the goal is to a
 nswer a question about a set of unknown distributions (for example\, the e
 fficacy of a treatment) by strategically sampling from them\, while provid
 ing guarantees on the returned answer. The archetypal example is the best 
 arm identification problem\, where the task is to find the arm with the la
 rgest mean. Top Two algorithms\, which select the next arm to sample from 
 among a leader and a challenger\, have received significant attention in r
 ecent years due to their simplicity and interpretability. In this talk\, I
  will present recent advances on three complementary aspects of pure explo
 ration: achieving non-asymptotic guarantees\, ensuring differential privac
 y\, and leveraging problem structure. First\, we propose a Top Two algorit
 hm which has an asymptotically optimal expected sample complexity\, and al
 so provides anytime guarantees on the probability of misidentifying a suff
 iciently good arm. Second\, we show how the Top Two principle can be combi
 ned with differential privacy mechanisms\, leading to algorithms that pres
 erve near-optimal efficiency while ensuring privacy guarantees. Finally\, 
 we address structured pure exploration by overcoming the computational bot
 tleneck of Pareto set identification in linear bandits\, through a game-ba
 sed algorithm grounded in posterior sampling. These results not only deepe
 n our theoretical understanding but also enable more practical and privacy
 -aware bandit algorithms for real-world problems. Bio: I am a postdoctoral
  researcher at EPFL in the Theory of Machine Learning lab\, working with N
 icolas Flammarion on the theoretical foundations of post-training for Larg
 e Language Models. I earned my PhD in Computer Science from the University
  of Lille\, under the supervision of Emilie Kaufmann and Rémy Degenne wit
 hin the Inria Scool team. I studied pure exploration in stochastic bandits
  and helped establish the Top Two approach as a principled methodology wit
 h strong theoretical and empirical performance. Previously\, I graduated f
 rom École Polytechnique and ETH Zurich.\n\nMachine Learning in Montpellie
 r\, Theory &amp\; Practice
ATTACH;FMTTYPE=image/jpeg:https://isdm.umontpellier.fr/wp-content/uploads/
 2026/06/ml-mtp-gC78d5.png
CATEGORIES:ML MTP
LOCATION:Room 02.124 Building 5\, St Priest Campus\, Montpellier\, 34000\, 
 France
X-APPLE-STRUCTURED-LOCATION;VALUE=URI;X-ADDRESS=St Priest Campus\, Montpell
 ier\, 34000\, France;X-APPLE-RADIUS=100;X-TITLE=Room 02.124 Building 5:geo
 :0,0
END:VEVENT
BEGIN:VTIMEZONE
TZID:Europe/Paris
X-LIC-LOCATION:Europe/Paris
BEGIN:DAYLIGHT
DTSTART:20250330T030000
TZOFFSETFROM:+0100
TZOFFSETTO:+0200
TZNAME:CEST
END:DAYLIGHT
END:VTIMEZONE
END:VCALENDAR