BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//jEvents 2.0 for Joomla//EN
CALSCALE:GREGORIAN
METHOD:PUBLISH
BEGIN:VTIMEZONE
TZID:America/New_York
BEGIN:STANDARD
DTSTART:20241111T140000
RDATE:20250309T030000
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:America/New_York EST
END:STANDARD
BEGIN:STANDARD
DTSTART:20251102T010000
RDATE:20260308T030000
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:America/New_York EST
END:STANDARD
BEGIN:STANDARD
DTSTART:20261101T010000
RDATE:20270314T030000
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:America/New_York EST
END:STANDARD
BEGIN:STANDARD
DTSTART:20271107T010000
RDATE:20280312T030000
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:America/New_York EST
END:STANDARD
BEGIN:STANDARD
DTSTART:20281105T010000
RDATE:20290311T030000
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:America/New_York EST
END:STANDARD
BEGIN:DAYLIGHT
DTSTART:20250309T030000
RDATE:20251102T010000
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:America/New_York EDT
END:DAYLIGHT
BEGIN:DAYLIGHT
DTSTART:20260308T030000
RDATE:20261101T010000
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:America/New_York EDT
END:DAYLIGHT
BEGIN:DAYLIGHT
DTSTART:20270314T030000
RDATE:20271107T010000
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:America/New_York EDT
END:DAYLIGHT
BEGIN:DAYLIGHT
DTSTART:20280312T030000
RDATE:20281105T010000
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:America/New_York EDT
END:DAYLIGHT
END:VTIMEZONE
BEGIN:VEVENT
UID:9519b16b3776984204dbf6dff1f7be94
CATEGORIES:Lean Seminar
CREATED:20251107T195749
SUMMARY:DeRL: Diverse‑Exploration Reinforcement Learning for Large Language Models
LOCATION:CoRE 431
DESCRIPTION:Current reinforcement-learning (RL) pipelines for large language models (LL
 Ms) that tackle mathematical reasoning and formal theorem proving tend to o
 ver-exploit a few high-probability chain-of-thought (CoT) sequences. Becaus
 e rewards are granted solely for producing correct answers, the policy quic
 kly converges on those paths, neglecting the rich space of alternative proo
 fs and solution strategies that math problems usually have. We address this
  limitation with Diverse-Exploration RL (DeRL), a simple yet effective modi
 fication to the reward function and the RL prompts. During training, the mo
 del is explicitly instructed to solve each problem without relying on its p
 reviously generated CoT. Next, an auxiliary LLM judge verifies the approach
  dissimilarity between the new LLM output and the previous CoT. Combined wi
 th the correctness metric, this new reward encourages exploration of novel 
 reasoning paths while preserving accuracy. We test DeRL on both natural-lan
 guage math questions with boxed answers and formal theorem proving problems
  in Lean. Across the MATH benchmark and Leanabell dataset, DeRL yields more
  than 10% relative gain compared to the PPO baseline for the Pass@1 metric.
  DeRL also consistently yields better results for the Pass@N metric. (mailt
 o:Pass@N metric.) Our findings demonstrate that incorporating diversity-awa
 re rewards facilitates broader exploration and enhances reasoning capabilit
 ies of LLMs, indicating a promising direction for improving current reinfor
 cement learning pipelines.\n I'm also going to mention other data augmentat
 ion techniques in Lean and reasoning in general.
X-ALT-DESC;FMTTYPE=text/html:<p><span style="color: #242424; font-family: 'Segoe UI', 'Segoe UI Web (Wes
 t European)', -apple-system, 'system-ui', Roboto, 'Helvetica Neue', sans-se
 rif; font-size: 15px; font-style: normal; font-weight: 400; letter-spacing:
  normal; orphans: 2; text-align: start; text-indent: 0px; text-transform: n
 one; widows: 2; word-spacing: 0px; white-space: normal; background-color: #
 ffffff; float: none;">Current reinforcement-learning (RL) pipelines for lar
 ge language models (LLMs) that tackle mathematical reasoning and formal the
 orem proving tend to over-exploit a few high-probability chain-of-thought (
 CoT) sequences. Because rewards are granted solely for producing correct an
 swers, the policy quickly converges on those paths, neglecting the rich spa
 ce of alternative proofs and solution strategies that math problems usually
  have. We address this limitation with Diverse-Exploration RL (DeRL), a sim
 ple yet effective modification to the reward function and the RL prompts. D
 uring training, the model is explicitly instructed to solve each problem wi
 thout relying on its previously generated CoT. Next, an auxiliary LLM judge
  verifies the approach dissimilarity between the new LLM output and the pre
 vious CoT. Combined with the correctness metric, this new reward encourages
  exploration of novel reasoning paths while preserving accuracy. We test De
 RL on both natural-language math questions with boxed answers and formal th
 eorem proving problems in Lean. Across the MATH benchmark and Leanabell dat
 aset, DeRL yields more than 10% relative gain compared to the PPO baseline 
 for the <a href="mailto:Pass@1 metric.">Pass@1 metric.</a> DeRL also consis
 tently yields better results for the <a href="mailto:Pass@N metric.">Pass@N
  metric.</a> Our findings demonstrate that incorporating diversity-aware re
 wards facilitates broader exploration and enhances reasoning capabilities o
 f LLMs, indicating a promising direction for improving current reinforcemen
 t learning pipelines.</span></p><div style="border: 0px; font-style: normal
 ; font-weight: 400; font-size: 15px; line-height: inherit; font-family: 'Se
 goe UI', 'Segoe UI Web (West European)', -apple-system, 'system-ui', Roboto
 , 'Helvetica Neue', sans-serif; margin: 0px; padding: 0px; vertical-align: 
 baseline; color: #242424; letter-spacing: normal; orphans: 2; text-align: s
 tart; text-indent: 0px; text-transform: none; widows: 2; word-spacing: 0px;
  white-space: normal; background-color: #ffffff;">&nbsp;</div><div style="b
 order: 0px; font-style: normal; font-weight: 400; font-size: 15px; line-hei
 ght: inherit; font-family: 'Segoe UI', 'Segoe UI Web (West European)', -app
 le-system, 'system-ui', Roboto, 'Helvetica Neue', sans-serif; margin: 0px; 
 padding: 0px; vertical-align: baseline; color: #242424; letter-spacing: nor
 mal; orphans: 2; text-align: start; text-indent: 0px; text-transform: none;
  widows: 2; word-spacing: 0px; white-space: normal; background-color: #ffff
 ff;">I'm also going to mention other data augmentation techniques in Lean a
 nd reasoning in general.</div>
CONTACT:Chenyang An
DTSTAMP:20260830T130048
DTSTART;TZID=America/New_York:20251112T140000
DTEND;TZID=America/New_York:20251112T150000
SEQUENCE:0
TRANSP:OPAQUE
END:VEVENT
END:VCALENDAR