SCI-Verifier: Scientific Verifier with Thinking
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zheng, Shenghe, Huang, Chenyu, Yu, Fangchen, Yao, Junchi, Ye, Jingqi, Chen, Tao, Luo, Yun, Ding, Ning, BAI, LEI, Cui, Ganqu, Ye, Peng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Scaling Physical Reasoning with the PHYSICS Dataset
von: Zheng, Shenghe, et al.
Veröffentlicht: (2025)
von: Zheng, Shenghe, et al.
Veröffentlicht: (2025)
RLPR: Extrapolating RLVR to General Domains without Verifiers
von: Yu, Tianyu, et al.
Veröffentlicht: (2025)
von: Yu, Tianyu, et al.
Veröffentlicht: (2025)
Scientific QA System with Verifiable Answers
von: Ljajić, Adela, et al.
Veröffentlicht: (2024)
von: Ljajić, Adela, et al.
Veröffentlicht: (2024)
InternBootcamp Technical Report: Boosting LLM Reasoning with Verifiable Task Scaling
von: Li, Peiji, et al.
Veröffentlicht: (2025)
von: Li, Peiji, et al.
Veröffentlicht: (2025)
Think Longer to Explore Deeper: Learn to Explore In-Context via Length-Incentivized Reinforcement Learning
von: Wang, Futing, et al.
Veröffentlicht: (2026)
von: Wang, Futing, et al.
Veröffentlicht: (2026)
Teaching Thinking Models to Reason with Tools: A Full-Pipeline Recipe for Tool-Integrated Reasoning
von: Cheng, Qianjia, et al.
Veröffentlicht: (2026)
von: Cheng, Qianjia, et al.
Veröffentlicht: (2026)
P1: Mastering Physics Olympiads with Reinforcement Learning
von: Chen, Jiacheng, et al.
Veröffentlicht: (2025)
von: Chen, Jiacheng, et al.
Veröffentlicht: (2025)
xVerify: Efficient Answer Verifier for Reasoning Model Evaluations
von: Chen, Ding, et al.
Veröffentlicht: (2025)
von: Chen, Ding, et al.
Veröffentlicht: (2025)
Verifying the Verifiers: Unveiling Pitfalls and Potentials in Fact Verifiers
von: Seo, Wooseok, et al.
Veröffentlicht: (2025)
von: Seo, Wooseok, et al.
Veröffentlicht: (2025)
Steering the Verifiability of Multimodal AI Hallucinations
von: Pang, Jianhong, et al.
Veröffentlicht: (2026)
von: Pang, Jianhong, et al.
Veröffentlicht: (2026)
CiteAudit: You Cited It, But Did You Read It? A Benchmark for Verifying Scientific References in the LLM Era
von: Shi, Kaiwen, et al.
Veröffentlicht: (2026)
von: Shi, Kaiwen, et al.
Veröffentlicht: (2026)
End-to-end Compositional Verification of Program Safety through Verified and Verifying Compilation
von: Wu, Jinhua, et al.
Veröffentlicht: (2025)
von: Wu, Jinhua, et al.
Veröffentlicht: (2025)
Towards Verified Compilation of Floating-point Optimization in Scientific Computing Programs
von: Tekriwal, Mohit, et al.
Veröffentlicht: (2025)
von: Tekriwal, Mohit, et al.
Veröffentlicht: (2025)
TextualVerifier: Verify TextGrad Step-by-Step
von: Situmorang, Eugenius Mario, et al.
Veröffentlicht: (2025)
von: Situmorang, Eugenius Mario, et al.
Veröffentlicht: (2025)
DocScope: Benchmarking Verifiable Reasoning for Trustworthy Long-Document Understanding
von: Feng, Xiang, et al.
Veröffentlicht: (2026)
von: Feng, Xiang, et al.
Veröffentlicht: (2026)
RustCompCert: A Verified and Verifying Compiler for a Sequential Subset of Rust
von: Wu, Jinhua, et al.
Veröffentlicht: (2026)
von: Wu, Jinhua, et al.
Veröffentlicht: (2026)
Verifiable Format Control for Large Language Model Generations
von: Wang, Zhaoyang, et al.
Veröffentlicht: (2025)
von: Wang, Zhaoyang, et al.
Veröffentlicht: (2025)
PhysicsMinions: Winning Gold Medals in the Latest Physics Olympiads with a Coevolutionary Multimodal Multi-Agent System
von: Yu, Fangchen, et al.
Veröffentlicht: (2025)
von: Yu, Fangchen, et al.
Veröffentlicht: (2025)
VERINA: Benchmarking Verifiable Code Generation
von: Ye, Zhe, et al.
Veröffentlicht: (2025)
von: Ye, Zhe, et al.
Veröffentlicht: (2025)
Generalizing Verifiable Instruction Following
von: Pyatkin, Valentina, et al.
Veröffentlicht: (2025)
von: Pyatkin, Valentina, et al.
Veröffentlicht: (2025)
LLatrieval: LLM-Verified Retrieval for Verifiable Generation
von: Li, Xiaonan, et al.
Veröffentlicht: (2023)
von: Li, Xiaonan, et al.
Veröffentlicht: (2023)
LLMEval-Logic: A Solver-Verified Chinese Benchmark for Logical Reasoning of LLMs with Adversarial Hardening
von: Zhang, Ming, et al.
Veröffentlicht: (2026)
von: Zhang, Ming, et al.
Veröffentlicht: (2026)
Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable Rewards
von: Liu, Xiaoyuan, et al.
Veröffentlicht: (2025)
von: Liu, Xiaoyuan, et al.
Veröffentlicht: (2025)
AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs
von: Pezeshkpour, Pouya, et al.
Veröffentlicht: (2026)
von: Pezeshkpour, Pouya, et al.
Veröffentlicht: (2026)
UltraIF: Advancing Instruction Following from the Wild
von: An, Kaikai, et al.
Veröffentlicht: (2025)
von: An, Kaikai, et al.
Veröffentlicht: (2025)
DeepSeekMath-V2: Towards Self-Verifiable Mathematical Reasoning
von: Shao, Zhihong, et al.
Veröffentlicht: (2025)
von: Shao, Zhihong, et al.
Veröffentlicht: (2025)
SCI-IDEA: Context-Aware Scientific Ideation Using Token and Sentence Embeddings
von: Keya, Farhana, et al.
Veröffentlicht: (2025)
von: Keya, Farhana, et al.
Veröffentlicht: (2025)
Asking LLMs to Verify First is Almost Free Lunch
von: Wu, Shiguang, et al.
Veröffentlicht: (2025)
von: Wu, Shiguang, et al.
Veröffentlicht: (2025)
RLVER: Reinforcement Learning with Verifiable Emotion Rewards for Empathetic Agents
von: Wang, Peisong, et al.
Veröffentlicht: (2025)
von: Wang, Peisong, et al.
Veröffentlicht: (2025)
From Verifiable Dot to Reward Chain: Harnessing Verifiable Reference-based Rewards for Reinforcement Learning of Open-ended Generation
von: Jiang, Yuxin, et al.
Veröffentlicht: (2026)
von: Jiang, Yuxin, et al.
Veröffentlicht: (2026)
Learning to Self-Verify Makes Language Models Better Reasoners
von: Chen, Yuxin, et al.
Veröffentlicht: (2026)
von: Chen, Yuxin, et al.
Veröffentlicht: (2026)
Toward Verifiable Misinformation Detection: A Multi-Tool LLM Agent Framework
von: Cui, Zikun, et al.
Veröffentlicht: (2025)
von: Cui, Zikun, et al.
Veröffentlicht: (2025)
VeriEquivBench: An Equivalence Score for Ground-Truth-Free Evaluation of Formally Verifiable Code
von: Zeng, Lingfei, et al.
Veröffentlicht: (2025)
von: Zeng, Lingfei, et al.
Veröffentlicht: (2025)
Scaling Agentic Verifier for Competitive Coding
von: Ma, Zeyao, et al.
Veröffentlicht: (2026)
von: Ma, Zeyao, et al.
Veröffentlicht: (2026)
Draft-OPD: On-Policy Distillation for Speculative Draft Models
von: Lei, Haodi, et al.
Veröffentlicht: (2026)
von: Lei, Haodi, et al.
Veröffentlicht: (2026)
SciArena: An Open Evaluation Platform for Non-Verifiable Scientific Literature-Grounded Tasks
von: Zhao, Yilun, et al.
Veröffentlicht: (2025)
von: Zhao, Yilun, et al.
Veröffentlicht: (2025)
TrimR: Verifier-based Training-Free Thinking Compression for Efficient Test-Time Scaling
von: Lin, Weizhe, et al.
Veröffentlicht: (2025)
von: Lin, Weizhe, et al.
Veröffentlicht: (2025)
Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning
von: Tong, Jingqi, et al.
Veröffentlicht: (2025)
von: Tong, Jingqi, et al.
Veröffentlicht: (2025)
A Unified Study of LoRA Variants: Taxonomy, Review, Codebase, and Empirical Evaluation
von: He, Haonan, et al.
Veröffentlicht: (2026)
von: He, Haonan, et al.
Veröffentlicht: (2026)
Verifier-Backed Hard Problem Generation for Mathematical Reasoning
von: Lai, Yuhang, et al.
Veröffentlicht: (2026)
von: Lai, Yuhang, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Scaling Physical Reasoning with the PHYSICS Dataset
von: Zheng, Shenghe, et al.
Veröffentlicht: (2025) -
RLPR: Extrapolating RLVR to General Domains without Verifiers
von: Yu, Tianyu, et al.
Veröffentlicht: (2025) -
Scientific QA System with Verifiable Answers
von: Ljajić, Adela, et al.
Veröffentlicht: (2024) -
InternBootcamp Technical Report: Boosting LLM Reasoning with Verifiable Task Scaling
von: Li, Peiji, et al.
Veröffentlicht: (2025) -
Think Longer to Explore Deeper: Learn to Explore In-Context via Length-Incentivized Reinforcement Learning
von: Wang, Futing, et al.
Veröffentlicht: (2026)