BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//jEvents 2.0 for Joomla//EN
CALSCALE:GREGORIAN
METHOD:PUBLISH
BEGIN:VTIMEZONE
TZID:America/New_York
BEGIN:STANDARD
DTSTART:20241111T140000
RDATE:20250309T030000
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:America/New_York EST
END:STANDARD
BEGIN:STANDARD
DTSTART:20251102T010000
RDATE:20260308T030000
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:America/New_York EST
END:STANDARD
BEGIN:STANDARD
DTSTART:20261101T010000
RDATE:20270314T030000
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:America/New_York EST
END:STANDARD
BEGIN:STANDARD
DTSTART:20271107T010000
RDATE:20280312T030000
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:America/New_York EST
END:STANDARD
BEGIN:STANDARD
DTSTART:20281105T010000
RDATE:20290311T030000
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:America/New_York EST
END:STANDARD
BEGIN:DAYLIGHT
DTSTART:20250309T030000
RDATE:20251102T010000
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:America/New_York EDT
END:DAYLIGHT
BEGIN:DAYLIGHT
DTSTART:20260308T030000
RDATE:20261101T010000
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:America/New_York EDT
END:DAYLIGHT
BEGIN:DAYLIGHT
DTSTART:20270314T030000
RDATE:20271107T010000
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:America/New_York EDT
END:DAYLIGHT
BEGIN:DAYLIGHT
DTSTART:20280312T030000
RDATE:20281105T010000
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:America/New_York EDT
END:DAYLIGHT
END:VTIMEZONE
BEGIN:VEVENT
UID:9519b16b3776984204dbf6dff1f7be94
CATEGORIES:Lean Seminar
CREATED:20251107T195749
SUMMARY:DeRL: Diverse‑Exploration Reinforcement Learning for Large Language Models
LOCATION:CoRE 431
DESCRIPTION:<p><span style="color: #242424; font-family: 'Segoe UI', 'Segoe UI Web (Wes
 t European)', -apple-system, 'system-ui', Roboto, 'Helvetica Neue', sans-se
 rif; font-size: 15px; font-style: normal; font-weight: 400; letter-spacing:
  normal; orphans: 2; text-align: start; text-indent: 0px; text-transform: n
 one; widows: 2; word-spacing: 0px; white-space: normal; background-color: #
 ffffff; float: none;">Current reinforcement-learning (RL) pipelines for lar
 ge language models (LLMs) that tackle mathematical reasoning and formal the
 orem proving tend to over-exploit a few high-probability chain-of-thought (
 CoT) sequences. Because rewards are granted solely for producing correct an
 swers, the policy quickly converges on those paths, neglecting the rich spa
 ce of alternative proofs and solution strategies that math problems usually
  have. We address this limitation with Diverse-Exploration RL (DeRL), a sim
 ple yet effective modification to the reward function and the RL prompts. D
 uring training, the model is explicitly instructed to solve each problem wi
 thout relying on its previously generated CoT. Next, an auxiliary LLM judge
  verifies the approach dissimilarity between the new LLM output and the pre
 vious CoT. Combined with the correctness metric, this new reward encourages
  exploration of novel reasoning paths while preserving accuracy. We test De
 RL on both natural-language math questions with boxed answers and formal th
 eorem proving problems in Lean. Across the MATH benchmark and Leanabell dat
 aset, DeRL yields more than 10% relative gain compared to the PPO baseline 
 for the <a href="mailto:Pass@1 metric.">Pass@1 metric.</a> DeRL also consis
 tently yields better results for the <a href="mailto:Pass@N metric.">Pass@N
  metric.</a> Our findings demonstrate that incorporating diversity-aware re
 wards facilitates broader exploration and enhances reasoning capabilities o
 f LLMs, indicating a promising direction for improving current reinforcemen
 t learning pipelines.</span></p><div style="border: 0px; font-style: normal
 ; font-weight: 400; font-size: 15px; line-height: inherit; font-family: 'Se
 goe UI', 'Segoe UI Web (West European)', -apple-system, 'system-ui', Roboto
 , 'Helvetica Neue', sans-serif; margin: 0px; padding: 0px; vertical-align: 
 baseline; color: #242424; letter-spacing: normal; orphans: 2; text-align: s
 tart; text-indent: 0px; text-transform: none; widows: 2; word-spacing: 0px;
  white-space: normal; background-color: #ffffff;">&nbsp;</div><div style="b
 order: 0px; font-style: normal; font-weight: 400; font-size: 15px; line-hei
 ght: inherit; font-family: 'Segoe UI', 'Segoe UI Web (West European)', -app
 le-system, 'system-ui', Roboto, 'Helvetica Neue', sans-serif; margin: 0px; 
 padding: 0px; vertical-align: baseline; color: #242424; letter-spacing: nor
 mal; orphans: 2; text-align: start; text-indent: 0px; text-transform: none;
  widows: 2; word-spacing: 0px; white-space: normal; background-color: #ffff
 ff;">I'm also going to mention other data augmentation techniques in Lean a
 nd reasoning in general.</div>
CONTACT:Chenyang An
DTSTAMP:20260830T044556
DTSTART;TZID=America/New_York:20251112T140000
DTEND;TZID=America/New_York:20251112T150000
SEQUENCE:0
TRANSP:OPAQUE
END:VEVENT
END:VCALENDAR