BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//jEvents 2.0 for Joomla//EN
CALSCALE:GREGORIAN
METHOD:PUBLISH
BEGIN:VTIMEZONE
TZID:America/New_York
BEGIN:STANDARD
DTSTART:20220228T104500
RDATE:20220313T030000
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:America/New_York EST
END:STANDARD
BEGIN:STANDARD
DTSTART:20221106T010000
RDATE:20230312T030000
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:America/New_York EST
END:STANDARD
BEGIN:STANDARD
DTSTART:20231105T010000
RDATE:20240310T030000
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:America/New_York EST
END:STANDARD
BEGIN:STANDARD
DTSTART:20241103T010000
RDATE:20250309T030000
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:America/New_York EST
END:STANDARD
BEGIN:STANDARD
DTSTART:20251102T010000
RDATE:20260308T030000
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:America/New_York EST
END:STANDARD
BEGIN:STANDARD
DTSTART:20261101T010000
RDATE:20270314T030000
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:America/New_York EST
END:STANDARD
BEGIN:STANDARD
DTSTART:20271107T010000
RDATE:20280312T030000
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:America/New_York EST
END:STANDARD
BEGIN:STANDARD
DTSTART:20281105T010000
RDATE:20290311T030000
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:America/New_York EST
END:STANDARD
BEGIN:DAYLIGHT
DTSTART:20220313T030000
RDATE:20221106T010000
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:America/New_York EDT
END:DAYLIGHT
BEGIN:DAYLIGHT
DTSTART:20230312T030000
RDATE:20231105T010000
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:America/New_York EDT
END:DAYLIGHT
BEGIN:DAYLIGHT
DTSTART:20240310T030000
RDATE:20241103T010000
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:America/New_York EDT
END:DAYLIGHT
BEGIN:DAYLIGHT
DTSTART:20250309T030000
RDATE:20251102T010000
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:America/New_York EDT
END:DAYLIGHT
BEGIN:DAYLIGHT
DTSTART:20260308T030000
RDATE:20261101T010000
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:America/New_York EDT
END:DAYLIGHT
BEGIN:DAYLIGHT
DTSTART:20270314T030000
RDATE:20271107T010000
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:America/New_York EDT
END:DAYLIGHT
BEGIN:DAYLIGHT
DTSTART:20280312T030000
RDATE:20281105T010000
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:America/New_York EDT
END:DAYLIGHT
END:VTIMEZONE
BEGIN:VEVENT
UID:52eaaa35066ce2ff84f4b4fb951807cb
CATEGORIES:Mathematical Physics Seminar
CREATED:20230223T091431
SUMMARY:Stochastic learning dynamics and generalization in neural networks: Can statistical physicists help understand AI?
LOCATION:Zoom
DESCRIPTION:Abstract: Despite the great success of deep learning, it remains largely a 
 black box. For example, the main search engine in deep neural networks is b
 ased on the Stochastic Gradient Descent (SGD) algorithm, however, little is
  known about how SGD finds ``good" solutions (low generalization error) in 
 the high-dimensional weight space. In this talk, we will first give a gener
 al overview of SGD followed by a more detailed description of our recent wo
 rk [1,2] on the SGD learning dynamics, the loss function landscape, and the
 ir relationship.\nMore specifically, our study shows that SGD dynamics foll
 ows a low-dimensional drift-diffusion motion in the weight space and the lo
 ss function is flat in most directions with large values of flatness (small
  curvatures). Furthermore, our study reveals a robust inverse relation betw
 een the weight variance in SGD and the landscape flatness opposite to the f
 luctuation-response relation in equilibrium systems. We develop a statistic
 al theory of SGD based on properties of the ensemble of minibatch loss func
 tions and show that the noise strength in SGD depends inversely on the land
 scape flatness, which explains the inverse variance-flatness relation. Our 
 study suggests that SGD serves as an ``intelligent" annealing strategy wher
 e the effective anisotropic “temperature” self-adjusts according to the los
 s landscape in order to find the flat minima that is found to be more gener
 alizable. Finally, we discuss an application of these insights for reducing
  catastrophic forgetting for sequential multiple tasks learning.\nTime perm
 its, we will discuss a more recent work on trying to understand why flat so
 lutions are more generalizable and whether there are other measures for bet
 ter generalization based on an exact duality relation we found between neur
 on activity and network weight [3].\n[1] “The inverse variance-flatness rel
 ation in Stochastic-Gradient-Descent is critical for finding flat minima”, 
 Y. Feng and Y. Tu, PNAS, 118 (9), 2021.\n[2] “Phases of learning dynamics i
 n artificial neural networks: in the absence and presence of mislabeled dat
 a”, Y. Feng and Y. Tu, Machine Learning: Science and Technology (MLST), Jul
 y 19, 2021.  (https://nam02.safelinks.protection.outlook.com/?url=https%3A%
 2F%2Fiopscience.iop.org%2Farticle%2F10.1088%2F2632-2153%2Fabf5b9%2Fpdf&amp;
 data=05%7C01%7Cavishag.klatzkin%40math.rutgers.edu%7C3ef3d49675d148459db108
 db14e1d3d5%7Cb92d2b234d35447093ff69aca6632ffe%7C1%7C0%7C638126732403174435%
 7CUnknown%7CTWFpbGZsb3d8eyJWIjoiMC4wLjAwMDAiLCJQIjoiV2luMzIiLCJBTiI6Ik1haWw
 iLCJXVCI6Mn0%3D%7C3000%7C%7C%7C&amp;sdata=HJ1XRmeSbHOKK4%2BijlVj2E45qzoDhiv
 K50OEO8G%2F5mI%3D&amp;reserved=0)https://iopscience.iop.org/article/10.1088
 /2632-2153/abf5b9/pdf\n[3] “The activity-weight duality in feed forward neu
 ral networks: The geometric determinants of generalization”, Y. Feng and Y.
  Tu,  (https://nam02.safelinks.protection.outlook.com/?url=https%3A%2F%2Far
 xiv.org%2Fabs%2F2203.10736&amp;data=05%7C01%7Cavishag.klatzkin%40math.rutge
 rs.edu%7C3ef3d49675d148459db108db14e1d3d5%7Cb92d2b234d35447093ff69aca6632ff
 e%7C1%7C0%7C638126732403174435%7CUnknown%7CTWFpbGZsb3d8eyJWIjoiMC4wLjAwMDAi
 LCJQIjoiV2luMzIiLCJBTiI6Ik1haWwiLCJXVCI6Mn0%3D%7C3000%7C%7C%7C&amp;sdata=AM
 zjTsrcRneqIEvE1s%2FD%2FdqUMdqT2i3WduaIbMplGyk%3D&amp;reserved=0)https://arx
 iv.org/abs/2203.10736\n
X-ALT-DESC;FMTTYPE=text/html:<p style="text-align: left;">Abstract: Despite the great success of deep le
 arning, it remains largely a black box. For example, the main search engine
  in deep neural networks is based on the Stochastic Gradient Descent (SGD) 
 algorithm, however, little is known about how SGD finds ``good" solutions (
 low generalization error) in the high-dimensional weight space. In this tal
 k, we will first give a general overview of SGD followed by a more detailed
  description of our recent work [1,2] on the SGD learning dynamics, the los
 s function landscape, and their relationship.</p><p style="text-align: left
 ;">More specifically, our study shows that SGD dynamics follows a low-dimen
 sional drift-diffusion motion in the weight space and the loss function is 
 flat in most directions with large values of flatness (small curvatures). F
 urthermore, our study reveals a robust inverse relation between the weight 
 variance in SGD and the landscape flatness opposite to the fluctuation-resp
 onse relation in equilibrium systems. We develop a statistical theory of SG
 D based on properties of the ensemble of minibatch loss functions and show 
 that the noise strength in SGD depends inversely on the landscape flatness,
  which explains the inverse variance-flatness relation. Our study suggests 
 that SGD serves as an ``intelligent" annealing strategy where the effective
  anisotropic “temperature” self-adjusts according to the loss landscape in 
 order to find the flat minima that is found to be more generalizable. Final
 ly, we discuss an application of these insights for reducing catastrophic f
 orgetting for sequential multiple tasks learning.</p><p style="text-align: 
 left;">Time permits, we will discuss a more recent work on trying to unders
 tand why flat solutions are more generalizable and whether there are other 
 measures for better generalization based on an exact duality relation we fo
 und between neuron activity and network weight [3].</p><p style="text-align
 : left;">[1] “The inverse variance-flatness relation in Stochastic-Gradient
 -Descent is critical for finding flat minima”, Y. Feng and Y. Tu, PNAS, 118
  (9), 2021.</p><p style="text-align: left;">[2] “Phases of learning dynamic
 s in artificial neural networks: in the absence and presence of mislabeled 
 data”, Y. Feng and Y. Tu, Machine Learning: Science and Technology (MLST), 
 July 19, 2021. <a href="https://nam02.safelinks.protection.outlook.com/?url
 =https%3A%2F%2Fiopscience.iop.org%2Farticle%2F10.1088%2F2632-2153%2Fabf5b9%
 2Fpdf&amp;data=05%7C01%7Cavishag.klatzkin%40math.rutgers.edu%7C3ef3d49675d1
 48459db108db14e1d3d5%7Cb92d2b234d35447093ff69aca6632ffe%7C1%7C0%7C638126732
 403174435%7CUnknown%7CTWFpbGZsb3d8eyJWIjoiMC4wLjAwMDAiLCJQIjoiV2luMzIiLCJBT
 iI6Ik1haWwiLCJXVCI6Mn0%3D%7C3000%7C%7C%7C&amp;sdata=HJ1XRmeSbHOKK4%2BijlVj2
 E45qzoDhivK50OEO8G%2F5mI%3D&amp;reserved=0"></a><a href="https://iopscience
 .iop.org/article/10.1088/2632-2153/abf5b9/pdf">https://iopscience.iop.org/a
 rticle/10.1088/2632-2153/abf5b9/pdf</a></p><p style="text-align: left;">[3]
  “The activity-weight duality in feed forward neural networks: The geometri
 c determinants of generalization”, Y. Feng and Y. Tu, <a href="https://nam0
 2.safelinks.protection.outlook.com/?url=https%3A%2F%2Farxiv.org%2Fabs%2F220
 3.10736&amp;data=05%7C01%7Cavishag.klatzkin%40math.rutgers.edu%7C3ef3d49675
 d148459db108db14e1d3d5%7Cb92d2b234d35447093ff69aca6632ffe%7C1%7C0%7C6381267
 32403174435%7CUnknown%7CTWFpbGZsb3d8eyJWIjoiMC4wLjAwMDAiLCJQIjoiV2luMzIiLCJ
 BTiI6Ik1haWwiLCJXVCI6Mn0%3D%7C3000%7C%7C%7C&amp;sdata=AMzjTsrcRneqIEvE1s%2F
 D%2FdqUMdqT2i3WduaIbMplGyk%3D&amp;reserved=0"></a><a href="https://arxiv.or
 g/abs/2203.10736">https://arxiv.org/abs/2203.10736</a></p>
CONTACT:Yuhai Tu - IBM T. J. Watson Research Center
DTSTAMP:20260829T124407
DTSTART;TZID=America/New_York:20230301T104500
DTEND;TZID=America/New_York:20230301T114500
SEQUENCE:0
TRANSP:OPAQUE
END:VEVENT
END:VCALENDAR