Evaluating the Role of Verifiers in Test-Time Scaling for Legal Reasoning Tasks
Fuente:
arXiv
Salvato in:
| Autori principali: | Romano, Davide, Schwarz, Jonathan, Giofré, Daniele |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Evaluating Test-Time Scaling LLMs for Legal Reasoning: OpenAI o1, DeepSeek-R1, and Beyond
di: Hu, Yinghao, et al.
Pubblicazione: (2025)
di: Hu, Yinghao, et al.
Pubblicazione: (2025)
Knowledge Graph-Assisted LLM Post-Training for Enhanced Legal Reasoning
di: Song, Dezhao, et al.
Pubblicazione: (2026)
di: Song, Dezhao, et al.
Pubblicazione: (2026)
InternBootcamp Technical Report: Boosting LLM Reasoning with Verifiable Task Scaling
di: Li, Peiji, et al.
Pubblicazione: (2025)
di: Li, Peiji, et al.
Pubblicazione: (2025)
xVerify: Efficient Answer Verifier for Reasoning Model Evaluations
di: Chen, Ding, et al.
Pubblicazione: (2025)
di: Chen, Ding, et al.
Pubblicazione: (2025)
Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning
di: Fatemi, Bahare, et al.
Pubblicazione: (2024)
di: Fatemi, Bahare, et al.
Pubblicazione: (2024)
Learning a Continue-Thinking Token for Enhanced Test-Time Scaling
di: Ringel, Liran, et al.
Pubblicazione: (2025)
di: Ringel, Liran, et al.
Pubblicazione: (2025)
Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models
di: Chang, Kaiyan, et al.
Pubblicazione: (2025)
di: Chang, Kaiyan, et al.
Pubblicazione: (2025)
Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning
di: Setlur, Amrith, et al.
Pubblicazione: (2024)
di: Setlur, Amrith, et al.
Pubblicazione: (2024)
Test-Time Scaling in Reasoning Models Is Not Effective for Knowledge-Intensive Tasks Yet
di: Zhao, James Xu, et al.
Pubblicazione: (2025)
di: Zhao, James Xu, et al.
Pubblicazione: (2025)
Linguistic Generalizability of Test-Time Scaling in Mathematical Reasoning
di: Son, Guijin, et al.
Pubblicazione: (2025)
di: Son, Guijin, et al.
Pubblicazione: (2025)
Scaling Flaws of Verifier-Guided Search in Mathematical Reasoning
di: Yu, Fei, et al.
Pubblicazione: (2025)
di: Yu, Fei, et al.
Pubblicazione: (2025)
On Verifiable Legal Reasoning: A Multi-Agent Framework with Formalized Knowledge Representations
di: Sadowski, Albert, et al.
Pubblicazione: (2025)
di: Sadowski, Albert, et al.
Pubblicazione: (2025)
Tasks and Roles in Legal AI: Data Curation, Annotation, and Verification
di: Koenecke, Allison, et al.
Pubblicazione: (2025)
di: Koenecke, Allison, et al.
Pubblicazione: (2025)
ATLAS: Adaptive Test-Time Latent Steering with External Verifiers for Enhancing LLMs Reasoning
di: Nguyen, Tuc, et al.
Pubblicazione: (2026)
di: Nguyen, Tuc, et al.
Pubblicazione: (2026)
On the Role of Feedback in Test-Time Scaling of Agentic AI Workflows
di: Chakraborty, Souradip, et al.
Pubblicazione: (2025)
di: Chakraborty, Souradip, et al.
Pubblicazione: (2025)
Evaluating Legal Reasoning Traces with Legal Issue Tree Rubrics
di: Lee, Jinu, et al.
Pubblicazione: (2025)
di: Lee, Jinu, et al.
Pubblicazione: (2025)
Accurate Legal Reasoning at Scale: Neuro-Symbolic Offloading and Structural Auditability for Robust Legal Adjudication
di: Sójka, Stanisław, et al.
Pubblicazione: (2026)
di: Sójka, Stanisław, et al.
Pubblicazione: (2026)
eagerlearners at SemEval2024 Task 5: The Legal Argument Reasoning Task in Civil Procedure
di: Sabzevari, Hoorieh, et al.
Pubblicazione: (2024)
di: Sabzevari, Hoorieh, et al.
Pubblicazione: (2024)
SoftCoT++: Test-Time Scaling with Soft Chain-of-Thought Reasoning
di: Xu, Yige, et al.
Pubblicazione: (2025)
di: Xu, Yige, et al.
Pubblicazione: (2025)
Chronos: Learning Temporal Dynamics of Reasoning Chains for Test-Time Scaling
di: Zhang, Kai, et al.
Pubblicazione: (2026)
di: Zhang, Kai, et al.
Pubblicazione: (2026)
Rethinking Agentic Workflows: Evaluating Inference-Based Test-Time Scaling Strategies in Text2SQL Tasks
di: Guo, Jiajing, et al.
Pubblicazione: (2025)
di: Guo, Jiajing, et al.
Pubblicazione: (2025)
Stepwise Reasoning Checkpoint Analysis: A Test Time Scaling Method to Enhance LLMs' Reasoning
di: Wang, Zezhong, et al.
Pubblicazione: (2025)
di: Wang, Zezhong, et al.
Pubblicazione: (2025)
Korean Canonical Legal Benchmark: Toward Knowledge-Independent Evaluation of LLMs' Legal Reasoning Capabilities
di: Oh, Hongseok, et al.
Pubblicazione: (2025)
di: Oh, Hongseok, et al.
Pubblicazione: (2025)
rStar-Coder: Scaling Competitive Code Reasoning with a Large-Scale Verified Dataset
di: Liu, Yifei, et al.
Pubblicazione: (2025)
di: Liu, Yifei, et al.
Pubblicazione: (2025)
Crosslingual Reasoning through Test-Time Scaling
di: Yong, Zheng-Xin, et al.
Pubblicazione: (2025)
di: Yong, Zheng-Xin, et al.
Pubblicazione: (2025)
On the Role of Temperature Sampling in Test-Time Scaling
di: Wu, Yuheng, et al.
Pubblicazione: (2025)
di: Wu, Yuheng, et al.
Pubblicazione: (2025)
VerifierQ: Enhancing LLM Test Time Compute with Q-Learning-based Verifiers
di: Qi, Jianing, et al.
Pubblicazione: (2024)
di: Qi, Jianing, et al.
Pubblicazione: (2024)
Test-Time Scaling of Reasoning Models for Machine Translation
di: Li, Zihao, et al.
Pubblicazione: (2025)
di: Li, Zihao, et al.
Pubblicazione: (2025)
Logical Reasoning with Outcome Reward Models for Test-Time Scaling
di: Thatikonda, Ramya Keerthy, et al.
Pubblicazione: (2025)
di: Thatikonda, Ramya Keerthy, et al.
Pubblicazione: (2025)
SynLogic: Synthesizing Verifiable Reasoning Data at Scale for Learning Logical Reasoning and Beyond
di: Liu, Junteng, et al.
Pubblicazione: (2025)
di: Liu, Junteng, et al.
Pubblicazione: (2025)
Parallel Test-Time Scaling for Latent Reasoning Models
di: You, Runyang, et al.
Pubblicazione: (2025)
di: You, Runyang, et al.
Pubblicazione: (2025)
From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling
di: Chen, Zhengyu, et al.
Pubblicazione: (2025)
di: Chen, Zhengyu, et al.
Pubblicazione: (2025)
TrimR: Verifier-based Training-Free Thinking Compression for Efficient Test-Time Scaling
di: Lin, Weizhe, et al.
Pubblicazione: (2025)
di: Lin, Weizhe, et al.
Pubblicazione: (2025)
LeMAJ (Legal LLM-as-a-Judge): Bridging Legal Reasoning and LLM Evaluation
di: Enguehard, Joseph, et al.
Pubblicazione: (2025)
di: Enguehard, Joseph, et al.
Pubblicazione: (2025)
Trust but Verify! A Survey on Verification Design for Test-time Scaling
di: Venktesh, V, et al.
Pubblicazione: (2025)
di: Venktesh, V, et al.
Pubblicazione: (2025)
MatryoshkaThinking: Recursive Test-Time Scaling Enables Efficient Reasoning
di: Chen, Hongwei, et al.
Pubblicazione: (2025)
di: Chen, Hongwei, et al.
Pubblicazione: (2025)
Towards Thinking-Optimal Scaling of Test-Time Compute for LLM Reasoning
di: Yang, Wenkai, et al.
Pubblicazione: (2025)
di: Yang, Wenkai, et al.
Pubblicazione: (2025)
Forest-of-Thought: Scaling Test-Time Compute for Enhancing LLM Reasoning
di: Bi, Zhenni, et al.
Pubblicazione: (2024)
di: Bi, Zhenni, et al.
Pubblicazione: (2024)
SciArena: An Open Evaluation Platform for Non-Verifiable Scientific Literature-Grounded Tasks
di: Zhao, Yilun, et al.
Pubblicazione: (2025)
di: Zhao, Yilun, et al.
Pubblicazione: (2025)
Magis-Bench: Evaluating LLMs on Magistrate-Level Legal Tasks
di: Pires, Ramon, et al.
Pubblicazione: (2026)
di: Pires, Ramon, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Evaluating Test-Time Scaling LLMs for Legal Reasoning: OpenAI o1, DeepSeek-R1, and Beyond
di: Hu, Yinghao, et al.
Pubblicazione: (2025) -
Knowledge Graph-Assisted LLM Post-Training for Enhanced Legal Reasoning
di: Song, Dezhao, et al.
Pubblicazione: (2026) -
InternBootcamp Technical Report: Boosting LLM Reasoning with Verifiable Task Scaling
di: Li, Peiji, et al.
Pubblicazione: (2025) -
xVerify: Efficient Answer Verifier for Reasoning Model Evaluations
di: Chen, Ding, et al.
Pubblicazione: (2025) -
Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning
di: Fatemi, Bahare, et al.
Pubblicazione: (2024)