BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//wp-events-plugin.com//7.4.0.1//EN
TZID:Europe/Paris
X-WR-TIMEZONE:Europe/Paris
BEGIN:VEVENT
UID:196@isdm.umontpellier.fr
DTSTART;TZID=Europe/Paris:20240321T140000
DTEND;TZID=Europe/Paris:20240321T140000
DTSTAMP:20260825T131610Z
URL:https://isdm.umontpellier.fr/events/adapting-newton-s-method-to-neural
 -networks-through-a-summary-of-higher-order-derivatives/
SUMMARY:Adapting Newton's Method to Neural Networks through a Summary of Hi
 gher-Order Derivatives
DESCRIPTION:Room 109\, IMAG\, Triolet campus\n\nMachine Learning in Montpel
 lier\, Theory &amp\; Practice\n\nWe consider a gradient-based optimization
  method applied to a function [image: \\mathcal{L}] of a vector of variabl
 es [image: {\\theta}]\, in the case where [image: {\\theta}] is represente
 d as a tuple of tensors [image: (\\mathbf{T}_1\, \\cdots\, \\mathbf{T}_S)]
 . This framework encompasses many common use-cases\, such as training neur
 al networks by gradient descent. First\, we propose a computationally inex
 pensive technique providing higher-order information on [image: \\mathcal{
 L}]\, especially about the interactions between the tensors [image: \\math
 bf{T}_s]\, based on automatic differentiation and computational tricks. Se
 cond\, we use this technique at order 2 to build a second-order optimizati
 on method which is suitable\, among other things\, for training deep neura
 l networks of various architectures. This second-order method leverages th
 e partition structure of [image: \\theta] into tensors [image: (\\mathbf{T
 }_1\, \\cdots\, \\mathbf{T}_S)]\, in such a way that it requires neither t
 he computation of the Hessian of [image: \\mathcal{L}] according to [image
 : \\theta]\, nor any approximation of it. The key part consists in computi
 ng a smaller matrix interpretable as a ``Hessian according to the partitio
 n’’\, which can be computed exactly and efficiently. In contrast to ma
 ny existing practical second-order methods used in neural networks\, which
  performs a diagonal or block-diagonal approximation of the Hessian or its
  inverse\, the method we propose does not neglect interactions between lay
 ers. GitHub: \\url{https://isdm.umontpellier.fr/github.com/p-wol/GroupedNe
 wton%7D. arXiv https://isdm.umontpellier.fr/arxiv.org/abs/2312.03885\n\nMa
 chine Learning in Montpellier\, Theory &amp\; Practice
ATTACH;FMTTYPE=image/jpeg:https://isdm.umontpellier.fr/wp-content/uploads/
 2026/06/ml-mtp-gC78d5.png
CATEGORIES:ML MTP
LOCATION:Triolet Campus- IMAG - Room 109\, Place Eugène Bataillon\, Montpe
 llier\, 
X-APPLE-STRUCTURED-LOCATION;VALUE=URI;X-ADDRESS=Place Eugène Bataillon\, M
 ontpellier\, ;X-APPLE-RADIUS=100;X-TITLE=Triolet Campus- IMAG - Room 109:g
 eo:0,0
END:VEVENT
BEGIN:VTIMEZONE
TZID:Europe/Paris
X-LIC-LOCATION:Europe/Paris
BEGIN:STANDARD
DTSTART:20231029T020000
TZOFFSETFROM:+0200
TZOFFSETTO:+0100
TZNAME:CET
END:STANDARD
END:VTIMEZONE
END:VCALENDAR