Legal Mathematical Reasoning with LLMs: Procedural Alignment through Two-Stage Reinforcement Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Kepu, Xie, Guofu, Yu, Weijie, Xu, Mingyue, Tang, Xu, Li, Yaxin, Xu, Jun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Beyond Guilt: Legal Judgment Prediction with Trichotomous Reasoning
von: Zhang, Kepu, et al.
Veröffentlicht: (2024)
von: Zhang, Kepu, et al.
Veröffentlicht: (2024)
An Explicit Syllogistic Legal Reasoning Framework for Large Language Models
von: Zhang, Kepu, et al.
Veröffentlicht: (2025)
von: Zhang, Kepu, et al.
Veröffentlicht: (2025)
CitaLaw: Enhancing LLM with Citations in Legal Domain
von: Zhang, Kepu, et al.
Veröffentlicht: (2024)
von: Zhang, Kepu, et al.
Veröffentlicht: (2024)
PrLM: Learning Explicit Reasoning for Personalized RAG via Contrastive Reward Optimization
von: Zhang, Kepu, et al.
Veröffentlicht: (2025)
von: Zhang, Kepu, et al.
Veröffentlicht: (2025)
Logic Rules as Explanations for Legal Case Retrieval
von: Sun, Zhongxiang, et al.
Veröffentlicht: (2024)
von: Sun, Zhongxiang, et al.
Veröffentlicht: (2024)
Legal$Δ$: Enhancing Legal Reasoning in LLMs via Reinforcement Learning with Chain-of-Thought Guided Information Gain
von: Dai, Xin, et al.
Veröffentlicht: (2025)
von: Dai, Xin, et al.
Veröffentlicht: (2025)
Sandwich Reasoning: An Answer-Reasoning-Answer Approach for Low-Latency Query Correction
von: Zhang, Chen, et al.
Veröffentlicht: (2026)
von: Zhang, Chen, et al.
Veröffentlicht: (2026)
Effective In-Context Example Selection through Data Compression
von: Sun, Zhongxiang, et al.
Veröffentlicht: (2024)
von: Sun, Zhongxiang, et al.
Veröffentlicht: (2024)
Emergent Hierarchical Reasoning in LLMs through Reinforcement Learning
von: Wang, Haozhe, et al.
Veröffentlicht: (2025)
von: Wang, Haozhe, et al.
Veröffentlicht: (2025)
On the Decision-Making Abilities in Role-Playing using Large Language Models
von: Shen, Chenglei, et al.
Veröffentlicht: (2024)
von: Shen, Chenglei, et al.
Veröffentlicht: (2024)
Can LLMs Time Travel? Enhancing Temporal Consistency in Legal Agentic Search through Reinforcement Learning
von: Fan, Wei, et al.
Veröffentlicht: (2026)
von: Fan, Wei, et al.
Veröffentlicht: (2026)
Atomic Thinking of LLMs: Decoupling and Exploring Mathematical Reasoning Abilities
von: Kuang, Jiayi, et al.
Veröffentlicht: (2025)
von: Kuang, Jiayi, et al.
Veröffentlicht: (2025)
Critique-RL: Training Language Models for Critiquing through Two-Stage Reinforcement Learning
von: Xi, Zhiheng, et al.
Veröffentlicht: (2025)
von: Xi, Zhiheng, et al.
Veröffentlicht: (2025)
IRIS: Interleaved Reinforcement with Incremental Staged Curriculum for Cross-Lingual Mathematical Reasoning
von: Gupta, Navya, et al.
Veröffentlicht: (2026)
von: Gupta, Navya, et al.
Veröffentlicht: (2026)
Unilaw-R1: A Large Language Model for Legal Reasoning with Reinforcement Learning and Iterative Inference
von: Cai, Hua, et al.
Veröffentlicht: (2025)
von: Cai, Hua, et al.
Veröffentlicht: (2025)
GraphDancer: Training LLMs to Explore and Reason over Graphs via Two-Stage Curriculum Post-Training
von: Bai, Yuyang, et al.
Veröffentlicht: (2026)
von: Bai, Yuyang, et al.
Veröffentlicht: (2026)
RLKD: Distilling LLMs' Reasoning via Reinforcement Learning
von: Xu, Shicheng, et al.
Veröffentlicht: (2025)
von: Xu, Shicheng, et al.
Veröffentlicht: (2025)
Trigger$^3$: Refining Query Correction via Adaptive Model Selector
von: Zhang, Kepu, et al.
Veröffentlicht: (2024)
von: Zhang, Kepu, et al.
Veröffentlicht: (2024)
HalluClean: A Unified Framework to Combat Hallucinations in LLMs
von: Zhao, Yaxin, et al.
Veröffentlicht: (2025)
von: Zhao, Yaxin, et al.
Veröffentlicht: (2025)
FalseReject: A Resource for Improving Contextual Safety and Mitigating Over-Refusals in LLMs via Structured Reasoning
von: Zhang, Zhehao, et al.
Veröffentlicht: (2025)
von: Zhang, Zhehao, et al.
Veröffentlicht: (2025)
MAPS: Motivation-Aware Personalized Search via LLM-Driven Consultation Alignment
von: Qin, Weicong, et al.
Veröffentlicht: (2025)
von: Qin, Weicong, et al.
Veröffentlicht: (2025)
DSPC: Dual-Stage Progressive Compression Framework for Efficient Long-Context Reasoning
von: Gao, Yaxin, et al.
Veröffentlicht: (2025)
von: Gao, Yaxin, et al.
Veröffentlicht: (2025)
Advancing Multimodal Reasoning: From Optimized Cold Start to Staged Reinforcement Learning
von: Chen, Shuang, et al.
Veröffentlicht: (2025)
von: Chen, Shuang, et al.
Veröffentlicht: (2025)
CAPO: Towards Enhancing LLM Reasoning through Generative Credit Assignment
von: Xie, Guofu, et al.
Veröffentlicht: (2025)
von: Xie, Guofu, et al.
Veröffentlicht: (2025)
Enhancing the Traditional Chinese Medicine Capabilities of Large Language Model through Reinforcement Learning from AI Feedback
von: Yu, Song, et al.
Veröffentlicht: (2024)
von: Yu, Song, et al.
Veröffentlicht: (2024)
PRIME: A Process-Outcome Alignment Benchmark for Verifiable Reasoning in Mathematics and Engineering
von: Wang, Xiangfeng, et al.
Veröffentlicht: (2026)
von: Wang, Xiangfeng, et al.
Veröffentlicht: (2026)
Gavel: Agent Meets Checklist for Evaluating LLMs on Long-Context Legal Summarization
von: Dou, Yao, et al.
Veröffentlicht: (2026)
von: Dou, Yao, et al.
Veröffentlicht: (2026)
Teaching LLMs According to Their Aptitude: Adaptive Reasoning for Mathematical Problem Solving
von: Xu, Xin, et al.
Veröffentlicht: (2025)
von: Xu, Xin, et al.
Veröffentlicht: (2025)
MathHay: An Automated Benchmark for Long-Context Mathematical Reasoning in LLMs
von: Wang, Lei, et al.
Veröffentlicht: (2024)
von: Wang, Lei, et al.
Veröffentlicht: (2024)
UserLM-R1: Modeling Human Reasoning in User Language Models with Multi-Reward Reinforcement Learning
von: Zhang, Feng, et al.
Veröffentlicht: (2026)
von: Zhang, Feng, et al.
Veröffentlicht: (2026)
How Difficulty-Aware Staged Reinforcement Learning Enhances LLMs' Reasoning Capabilities: A Preliminary Experimental Study
von: Ji, Yunjie, et al.
Veröffentlicht: (2025)
von: Ji, Yunjie, et al.
Veröffentlicht: (2025)
KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning
von: Xu, Hongling, et al.
Veröffentlicht: (2025)
von: Xu, Hongling, et al.
Veröffentlicht: (2025)
A2R: An Asymmetric Two-Stage Reasoning Framework for Parallel Reasoning
von: Wang, Ziqi, et al.
Veröffentlicht: (2025)
von: Wang, Ziqi, et al.
Veröffentlicht: (2025)
Enabling Discriminative Reasoning in LLMs for Legal Judgment Prediction
von: Deng, Chenlong, et al.
Veröffentlicht: (2024)
von: Deng, Chenlong, et al.
Veröffentlicht: (2024)
Learning How and What to Memorize: Cognition-Inspired Two-Stage Optimization for Evolving Memory
von: Xu, Derong, et al.
Veröffentlicht: (2026)
von: Xu, Derong, et al.
Veröffentlicht: (2026)
Does Learning Mathematical Problem-Solving Generalize to Broader Reasoning?
von: Zhou, Ruochen, et al.
Veröffentlicht: (2025)
von: Zhou, Ruochen, et al.
Veröffentlicht: (2025)
Topic-Guided Reinforcement Learning with LLMs for Enhancing Multi-Document Summarization
von: Li, Chuyuan, et al.
Veröffentlicht: (2025)
von: Li, Chuyuan, et al.
Veröffentlicht: (2025)
LegalDuet: Learning Fine-grained Representations for Legal Judgment Prediction via a Dual-View Contrastive Learning
von: Xu, Buqiang, et al.
Veröffentlicht: (2024)
von: Xu, Buqiang, et al.
Veröffentlicht: (2024)
Self-Error-Instruct: Generalizing from Errors for LLMs Mathematical Reasoning
von: Yu, Erxin, et al.
Veröffentlicht: (2025)
von: Yu, Erxin, et al.
Veröffentlicht: (2025)
Towards Advanced Mathematical Reasoning for LLMs via First-Order Logic Theorem Proving
von: Cao, Chuxue, et al.
Veröffentlicht: (2025)
von: Cao, Chuxue, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Beyond Guilt: Legal Judgment Prediction with Trichotomous Reasoning
von: Zhang, Kepu, et al.
Veröffentlicht: (2024) -
An Explicit Syllogistic Legal Reasoning Framework for Large Language Models
von: Zhang, Kepu, et al.
Veröffentlicht: (2025) -
CitaLaw: Enhancing LLM with Citations in Legal Domain
von: Zhang, Kepu, et al.
Veröffentlicht: (2024) -
PrLM: Learning Explicit Reasoning for Personalized RAG via Contrastive Reward Optimization
von: Zhang, Kepu, et al.
Veröffentlicht: (2025) -
Logic Rules as Explanations for Legal Case Retrieval
von: Sun, Zhongxiang, et al.
Veröffentlicht: (2024)