Detecting RLVR Training Data via Structural Convergence of Reasoning
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhang, Hongbo, Yang, Yue, Yan, Jianhao, Bao, Guangsheng, Zhang, Yue |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
How Likely Do LLMs with CoT Mimic Human Reasoning?
por: Bao, Guangsheng, et al.
Publicado: (2024)
por: Bao, Guangsheng, et al.
Publicado: (2024)
Direct Value Optimization: Improving Chain-of-Thought Reasoning in LLMs with Refined Values
por: Zhang, Hongbo, et al.
Publicado: (2025)
por: Zhang, Hongbo, et al.
Publicado: (2025)
Glimpse: Enabling White-Box Methods to Use Proprietary Models for Zero-Shot LLM-Generated Text Detection
por: Bao, Guangsheng, et al.
Publicado: (2024)
por: Bao, Guangsheng, et al.
Publicado: (2024)
CycleResearcher: Improving Automated Research via Automated Review
por: Weng, Yixuan, et al.
Publicado: (2024)
por: Weng, Yixuan, et al.
Publicado: (2024)
RefuteBench: Evaluating Refuting Instruction-Following for Large Language Models
por: Yan, Jianhao, et al.
Publicado: (2024)
por: Yan, Jianhao, et al.
Publicado: (2024)
Decoupling Content and Expression: Two-Dimensional Detection of AI-Generated Text
por: Bao, Guangsheng, et al.
Publicado: (2025)
por: Bao, Guangsheng, et al.
Publicado: (2025)
Correlation or Causation: Analyzing the Causal Structures of LLM and LRM Reasoning Process
por: FU, Zhizhang, et al.
Publicado: (2025)
por: FU, Zhizhang, et al.
Publicado: (2025)
Potential and Challenges of Model Editing for Social Debiasing
por: Yan, Jianhao, et al.
Publicado: (2024)
por: Yan, Jianhao, et al.
Publicado: (2024)
Exploring Multi-Temperature Strategies for Token- and Rollout-Level Control in RLVR
por: Zhuang, Haomin, et al.
Publicado: (2025)
por: Zhuang, Haomin, et al.
Publicado: (2025)
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels
por: Yan, Jianhao, et al.
Publicado: (2024)
por: Yan, Jianhao, et al.
Publicado: (2024)
From Query to Counsel: Structured Reasoning with a Multi-Agent Framework and Dataset for Legal Consultation
por: Lu, Mingfei, et al.
Publicado: (2026)
por: Lu, Mingfei, et al.
Publicado: (2026)
Learning to Reason under Off-Policy Guidance
por: Yan, Jianhao, et al.
Publicado: (2025)
por: Yan, Jianhao, et al.
Publicado: (2025)
Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes
por: Jiao, Rui, et al.
Publicado: (2025)
por: Jiao, Rui, et al.
Publicado: (2025)
Federated Learning with Layer Skipping: Efficient Training of Large Language Models for Healthcare NLP
por: Zhang, Lihong, et al.
Publicado: (2025)
por: Zhang, Lihong, et al.
Publicado: (2025)
Know Your Needs Better: Towards Structured Understanding of Marketer Demands with Analogical Reasoning Augmented LLMs
por: Wang, Junjie, et al.
Publicado: (2024)
por: Wang, Junjie, et al.
Publicado: (2024)
A State-of-the-Art SQL Reasoning Model using RLVR
por: Ali, Alnur, et al.
Publicado: (2025)
por: Ali, Alnur, et al.
Publicado: (2025)
Med-RLVR: Emerging Medical Reasoning from a 3B base model via reinforcement Learning
por: Zhang, Sheng, et al.
Publicado: (2025)
por: Zhang, Sheng, et al.
Publicado: (2025)
Supervised Knowledge Makes Large Language Models Better In-context Learners
por: Yang, Linyi, et al.
Publicado: (2023)
por: Yang, Linyi, et al.
Publicado: (2023)
Satirical News Detection with Semantic Feature Extraction and Game-theoretic Rough Sets
por: Zhou, Yue, et al.
Publicado: (2020)
por: Zhou, Yue, et al.
Publicado: (2020)
Paraphrase and Solve: Exploring and Exploiting the Impact of Surface Form on Mathematical Reasoning in Large Language Models
por: Zhou, Yue, et al.
Publicado: (2024)
por: Zhou, Yue, et al.
Publicado: (2024)
Selective Attention Federated Learning: Improving Privacy and Efficiency for Clinical Text Classification
por: Li, Yue, et al.
Publicado: (2025)
por: Li, Yue, et al.
Publicado: (2025)
The Death of Feature Engineering? BERT with Linguistic Features on SQuAD 2.0
por: Li, Jiawei, et al.
Publicado: (2024)
por: Li, Jiawei, et al.
Publicado: (2024)
JudgeRLVR: Judge First, Generate Second for Efficient Reasoning
por: Duo, Jiangshan, et al.
Publicado: (2026)
por: Duo, Jiangshan, et al.
Publicado: (2026)
Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards
por: Da, Jeff, et al.
Publicado: (2025)
por: Da, Jeff, et al.
Publicado: (2025)
Towards Safe Reasoning in Large Reasoning Models via Corrective Intervention
por: Zhang, Yichi, et al.
Publicado: (2025)
por: Zhang, Yichi, et al.
Publicado: (2025)
Reasoning Depth and Environment Complexity: A Controlled Study of RLVR Data Allocation across Logical Reasoning Tasks
por: Zhu, Yihua, et al.
Publicado: (2026)
por: Zhu, Yihua, et al.
Publicado: (2026)
Recycling Failures: Salvaging Exploration in RLVR via Fine-Grained Off-Policy Guidance
por: Ren, Yanwei, et al.
Publicado: (2026)
por: Ren, Yanwei, et al.
Publicado: (2026)
Dissecting Failure Dynamics in Large Language Model Reasoning
por: Zhu, Wei, et al.
Publicado: (2026)
por: Zhu, Wei, et al.
Publicado: (2026)
Fast-DetectGPT: Efficient Zero-Shot Detection of Machine-Generated Text via Conditional Probability Curvature
por: Bao, Guangsheng, et al.
Publicado: (2023)
por: Bao, Guangsheng, et al.
Publicado: (2023)
Balancing Faithfulness and Performance in Reasoning via Multi-Listener Soft Execution
por: Sivakumaran, Nithin, et al.
Publicado: (2026)
por: Sivakumaran, Nithin, et al.
Publicado: (2026)
Logical Reasoning in Large Language Models: A Survey
por: Liu, Hanmeng, et al.
Publicado: (2025)
por: Liu, Hanmeng, et al.
Publicado: (2025)
CLAUSE: Agentic Neuro-Symbolic Knowledge Graph Reasoning via Dynamic Learnable Context Engineering
por: Zhao, Yang, et al.
Publicado: (2025)
por: Zhao, Yang, et al.
Publicado: (2025)
Towards Data Contamination Detection for Modern Large Language Models: Limitations, Inconsistencies, and Oracle Challenges
por: Samuel, Vinay, et al.
Publicado: (2024)
por: Samuel, Vinay, et al.
Publicado: (2024)
Break the Chain: Large Language Models Can be Shortcut Reasoners
por: Ding, Mengru, et al.
Publicado: (2024)
por: Ding, Mengru, et al.
Publicado: (2024)
GAUSS: Benchmarking Structured Mathematical Skills for Large Language Models
por: Zhang, Yue, et al.
Publicado: (2025)
por: Zhang, Yue, et al.
Publicado: (2025)
Absolute Zero: Reinforced Self-play Reasoning with Zero Data
por: Zhao, Andrew, et al.
Publicado: (2025)
por: Zhao, Andrew, et al.
Publicado: (2025)
Speculative Thinking: Enhancing Small-Model Reasoning with Large Model Guidance at Inference Time
por: Yang, Wang, et al.
Publicado: (2025)
por: Yang, Wang, et al.
Publicado: (2025)
Logic Agent: Enhancing Validity with Logic Rule Invocation
por: Liu, Hanmeng, et al.
Publicado: (2024)
por: Liu, Hanmeng, et al.
Publicado: (2024)
SalaMAnder: Shapley-based Mathematical Expression Attribution and Metric for Chain-of-Thought Reasoning
por: Xin, Yue, et al.
Publicado: (2025)
por: Xin, Yue, et al.
Publicado: (2025)
When AI Settles Down: Late-Stage Stability as a Signature of AI-Generated Text Detection
por: Sun, Ke, et al.
Publicado: (2026)
por: Sun, Ke, et al.
Publicado: (2026)
Ejemplares similares
-
How Likely Do LLMs with CoT Mimic Human Reasoning?
por: Bao, Guangsheng, et al.
Publicado: (2024) -
Direct Value Optimization: Improving Chain-of-Thought Reasoning in LLMs with Refined Values
por: Zhang, Hongbo, et al.
Publicado: (2025) -
Glimpse: Enabling White-Box Methods to Use Proprietary Models for Zero-Shot LLM-Generated Text Detection
por: Bao, Guangsheng, et al.
Publicado: (2024) -
CycleResearcher: Improving Automated Research via Automated Review
por: Weng, Yixuan, et al.
Publicado: (2024) -
RefuteBench: Evaluating Refuting Instruction-Following for Large Language Models
por: Yan, Jianhao, et al.
Publicado: (2024)