Evaluating the Role of Verifiers in Test-Time Scaling for Legal Reasoning Tasks
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Romano, Davide, Schwarz, Jonathan, Giofré, Daniele |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Evaluating Test-Time Scaling LLMs for Legal Reasoning: OpenAI o1, DeepSeek-R1, and Beyond
par: Hu, Yinghao, et autres
Publié: (2025)
par: Hu, Yinghao, et autres
Publié: (2025)
Knowledge Graph-Assisted LLM Post-Training for Enhanced Legal Reasoning
par: Song, Dezhao, et autres
Publié: (2026)
par: Song, Dezhao, et autres
Publié: (2026)
InternBootcamp Technical Report: Boosting LLM Reasoning with Verifiable Task Scaling
par: Li, Peiji, et autres
Publié: (2025)
par: Li, Peiji, et autres
Publié: (2025)
xVerify: Efficient Answer Verifier for Reasoning Model Evaluations
par: Chen, Ding, et autres
Publié: (2025)
par: Chen, Ding, et autres
Publié: (2025)
Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning
par: Fatemi, Bahare, et autres
Publié: (2024)
par: Fatemi, Bahare, et autres
Publié: (2024)
Learning a Continue-Thinking Token for Enhanced Test-Time Scaling
par: Ringel, Liran, et autres
Publié: (2025)
par: Ringel, Liran, et autres
Publié: (2025)
Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models
par: Chang, Kaiyan, et autres
Publié: (2025)
par: Chang, Kaiyan, et autres
Publié: (2025)
Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning
par: Setlur, Amrith, et autres
Publié: (2024)
par: Setlur, Amrith, et autres
Publié: (2024)
Test-Time Scaling in Reasoning Models Is Not Effective for Knowledge-Intensive Tasks Yet
par: Zhao, James Xu, et autres
Publié: (2025)
par: Zhao, James Xu, et autres
Publié: (2025)
Linguistic Generalizability of Test-Time Scaling in Mathematical Reasoning
par: Son, Guijin, et autres
Publié: (2025)
par: Son, Guijin, et autres
Publié: (2025)
Scaling Flaws of Verifier-Guided Search in Mathematical Reasoning
par: Yu, Fei, et autres
Publié: (2025)
par: Yu, Fei, et autres
Publié: (2025)
On Verifiable Legal Reasoning: A Multi-Agent Framework with Formalized Knowledge Representations
par: Sadowski, Albert, et autres
Publié: (2025)
par: Sadowski, Albert, et autres
Publié: (2025)
Tasks and Roles in Legal AI: Data Curation, Annotation, and Verification
par: Koenecke, Allison, et autres
Publié: (2025)
par: Koenecke, Allison, et autres
Publié: (2025)
ATLAS: Adaptive Test-Time Latent Steering with External Verifiers for Enhancing LLMs Reasoning
par: Nguyen, Tuc, et autres
Publié: (2026)
par: Nguyen, Tuc, et autres
Publié: (2026)
On the Role of Feedback in Test-Time Scaling of Agentic AI Workflows
par: Chakraborty, Souradip, et autres
Publié: (2025)
par: Chakraborty, Souradip, et autres
Publié: (2025)
Evaluating Legal Reasoning Traces with Legal Issue Tree Rubrics
par: Lee, Jinu, et autres
Publié: (2025)
par: Lee, Jinu, et autres
Publié: (2025)
Accurate Legal Reasoning at Scale: Neuro-Symbolic Offloading and Structural Auditability for Robust Legal Adjudication
par: Sójka, Stanisław, et autres
Publié: (2026)
par: Sójka, Stanisław, et autres
Publié: (2026)
eagerlearners at SemEval2024 Task 5: The Legal Argument Reasoning Task in Civil Procedure
par: Sabzevari, Hoorieh, et autres
Publié: (2024)
par: Sabzevari, Hoorieh, et autres
Publié: (2024)
SoftCoT++: Test-Time Scaling with Soft Chain-of-Thought Reasoning
par: Xu, Yige, et autres
Publié: (2025)
par: Xu, Yige, et autres
Publié: (2025)
Chronos: Learning Temporal Dynamics of Reasoning Chains for Test-Time Scaling
par: Zhang, Kai, et autres
Publié: (2026)
par: Zhang, Kai, et autres
Publié: (2026)
Rethinking Agentic Workflows: Evaluating Inference-Based Test-Time Scaling Strategies in Text2SQL Tasks
par: Guo, Jiajing, et autres
Publié: (2025)
par: Guo, Jiajing, et autres
Publié: (2025)
Stepwise Reasoning Checkpoint Analysis: A Test Time Scaling Method to Enhance LLMs' Reasoning
par: Wang, Zezhong, et autres
Publié: (2025)
par: Wang, Zezhong, et autres
Publié: (2025)
Korean Canonical Legal Benchmark: Toward Knowledge-Independent Evaluation of LLMs' Legal Reasoning Capabilities
par: Oh, Hongseok, et autres
Publié: (2025)
par: Oh, Hongseok, et autres
Publié: (2025)
rStar-Coder: Scaling Competitive Code Reasoning with a Large-Scale Verified Dataset
par: Liu, Yifei, et autres
Publié: (2025)
par: Liu, Yifei, et autres
Publié: (2025)
Crosslingual Reasoning through Test-Time Scaling
par: Yong, Zheng-Xin, et autres
Publié: (2025)
par: Yong, Zheng-Xin, et autres
Publié: (2025)
On the Role of Temperature Sampling in Test-Time Scaling
par: Wu, Yuheng, et autres
Publié: (2025)
par: Wu, Yuheng, et autres
Publié: (2025)
VerifierQ: Enhancing LLM Test Time Compute with Q-Learning-based Verifiers
par: Qi, Jianing, et autres
Publié: (2024)
par: Qi, Jianing, et autres
Publié: (2024)
Test-Time Scaling of Reasoning Models for Machine Translation
par: Li, Zihao, et autres
Publié: (2025)
par: Li, Zihao, et autres
Publié: (2025)
Logical Reasoning with Outcome Reward Models for Test-Time Scaling
par: Thatikonda, Ramya Keerthy, et autres
Publié: (2025)
par: Thatikonda, Ramya Keerthy, et autres
Publié: (2025)
SynLogic: Synthesizing Verifiable Reasoning Data at Scale for Learning Logical Reasoning and Beyond
par: Liu, Junteng, et autres
Publié: (2025)
par: Liu, Junteng, et autres
Publié: (2025)
Parallel Test-Time Scaling for Latent Reasoning Models
par: You, Runyang, et autres
Publié: (2025)
par: You, Runyang, et autres
Publié: (2025)
From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling
par: Chen, Zhengyu, et autres
Publié: (2025)
par: Chen, Zhengyu, et autres
Publié: (2025)
TrimR: Verifier-based Training-Free Thinking Compression for Efficient Test-Time Scaling
par: Lin, Weizhe, et autres
Publié: (2025)
par: Lin, Weizhe, et autres
Publié: (2025)
LeMAJ (Legal LLM-as-a-Judge): Bridging Legal Reasoning and LLM Evaluation
par: Enguehard, Joseph, et autres
Publié: (2025)
par: Enguehard, Joseph, et autres
Publié: (2025)
Trust but Verify! A Survey on Verification Design for Test-time Scaling
par: Venktesh, V, et autres
Publié: (2025)
par: Venktesh, V, et autres
Publié: (2025)
MatryoshkaThinking: Recursive Test-Time Scaling Enables Efficient Reasoning
par: Chen, Hongwei, et autres
Publié: (2025)
par: Chen, Hongwei, et autres
Publié: (2025)
Towards Thinking-Optimal Scaling of Test-Time Compute for LLM Reasoning
par: Yang, Wenkai, et autres
Publié: (2025)
par: Yang, Wenkai, et autres
Publié: (2025)
Forest-of-Thought: Scaling Test-Time Compute for Enhancing LLM Reasoning
par: Bi, Zhenni, et autres
Publié: (2024)
par: Bi, Zhenni, et autres
Publié: (2024)
SciArena: An Open Evaluation Platform for Non-Verifiable Scientific Literature-Grounded Tasks
par: Zhao, Yilun, et autres
Publié: (2025)
par: Zhao, Yilun, et autres
Publié: (2025)
Magis-Bench: Evaluating LLMs on Magistrate-Level Legal Tasks
par: Pires, Ramon, et autres
Publié: (2026)
par: Pires, Ramon, et autres
Publié: (2026)
Documents similaires
-
Evaluating Test-Time Scaling LLMs for Legal Reasoning: OpenAI o1, DeepSeek-R1, and Beyond
par: Hu, Yinghao, et autres
Publié: (2025) -
Knowledge Graph-Assisted LLM Post-Training for Enhanced Legal Reasoning
par: Song, Dezhao, et autres
Publié: (2026) -
InternBootcamp Technical Report: Boosting LLM Reasoning with Verifiable Task Scaling
par: Li, Peiji, et autres
Publié: (2025) -
xVerify: Efficient Answer Verifier for Reasoning Model Evaluations
par: Chen, Ding, et autres
Publié: (2025) -
Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning
par: Fatemi, Bahare, et autres
Publié: (2024)