DC-W2S: Dual-Consensus Weak-to-Strong Training for Reliable Process Reward Modeling in Biological Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Chan, Chi-Min, Hajiramezanali, Ehsan, Li, Xiner, De Brouwer, Edward, Edwards, Carl, Xue, Wei, Han, Sirui, Guo, Yike, Scalia, Gabriele |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RAG-Enhanced Collaborative LLM Agents for Drug Discovery
by: Lee, Namkyeong, et al.
Published: (2025)
by: Lee, Namkyeong, et al.
Published: (2025)
MolCap-Arena: A Comprehensive Captioning Benchmark on Language-Enhanced Molecular Property Prediction
by: Edwards, Carl, et al.
Published: (2024)
by: Edwards, Carl, et al.
Published: (2024)
AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents
by: De Brouwer, Edward, et al.
Published: (2026)
by: De Brouwer, Edward, et al.
Published: (2026)
Iterative Distillation for Reward-Guided Fine-Tuning of Diffusion Models in Biomolecular Design
by: Su, Xingyu, et al.
Published: (2025)
by: Su, Xingyu, et al.
Published: (2025)
Adding Conditional Control to Diffusion Models with Reinforcement Learning
by: Zhao, Yulai, et al.
Published: (2024)
by: Zhao, Yulai, et al.
Published: (2024)
Cell Morphology-Guided Small Molecule Generation with GFlowNets
by: Lu, Stephen Zhewen, et al.
Published: (2024)
by: Lu, Stephen Zhewen, et al.
Published: (2024)
Bridging Model-Based Optimization and Generative Modeling via Conservative Fine-Tuning of Diffusion Models
by: Uehara, Masatoshi, et al.
Published: (2024)
by: Uehara, Masatoshi, et al.
Published: (2024)
Generate in Reconstruction Space, Match in Semantic Space: Transport Geometry for One-Step Generation
by: Van Assel, Hugues, et al.
Published: (2026)
by: Van Assel, Hugues, et al.
Published: (2026)
Feedback Efficient Online Fine-Tuning of Diffusion Models
by: Uehara, Masatoshi, et al.
Published: (2024)
by: Uehara, Masatoshi, et al.
Published: (2024)
Fine-Tuning of Continuous-Time Diffusion Models as Entropy-Regularized Control
by: Uehara, Masatoshi, et al.
Published: (2024)
by: Uehara, Masatoshi, et al.
Published: (2024)
Not Just the Destination, But the Journey: Reasoning Traces Causally Shape Generalization Behaviors
by: Wen, Pengcheng, et al.
Published: (2026)
by: Wen, Pengcheng, et al.
Published: (2026)
Dynamic Search for Inference-Time Alignment in Diffusion Models
by: Li, Xiner, et al.
Published: (2025)
by: Li, Xiner, et al.
Published: (2025)
Weak-to-Strong Reasoning
by: Yang, Yuqing, et al.
Published: (2024)
by: Yang, Yuqing, et al.
Published: (2024)
Incentivizing Strong Reasoning from Weak Supervision
by: Yuan, Yige, et al.
Published: (2025)
by: Yuan, Yige, et al.
Published: (2025)
What, Whether and How? Unveiling Process Reward Models for Thinking with Images Reasoning
by: Zhou, Yujin, et al.
Published: (2026)
by: Zhou, Yujin, et al.
Published: (2026)
Reliable Weak-to-Strong Monitoring of LLM Agents
by: Kale, Neil, et al.
Published: (2025)
by: Kale, Neil, et al.
Published: (2025)
Weak-for-Strong: Training Weak Meta-Agent to Harness Strong Executors
by: Nie, Fan, et al.
Published: (2025)
by: Nie, Fan, et al.
Published: (2025)
VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning
by: Lu, Xingyu, et al.
Published: (2026)
by: Lu, Xingyu, et al.
Published: (2026)
E-valuator: Reliable Agent Verifiers with Sequential Hypothesis Testing
by: Sadhuka, Shuvom, et al.
Published: (2025)
by: Sadhuka, Shuvom, et al.
Published: (2025)
Weak-to-Strong Preference Optimization: Stealing Reward from Weak Aligned Model
by: Zhu, Wenhong, et al.
Published: (2024)
by: Zhu, Wenhong, et al.
Published: (2024)
Improving Weak-to-Strong Generalization with Reliability-Aware Alignment
by: Guo, Yue, et al.
Published: (2024)
by: Guo, Yue, et al.
Published: (2024)
Toward the Identifiability of Comparative Deep Generative Models
by: Lopez, Romain, et al.
Published: (2024)
by: Lopez, Romain, et al.
Published: (2024)
When Slower Isn't Truer: Inverse Scaling Law of Truthfulness in Multimodal Reasoning
by: Fang, Sitong, et al.
Published: (2025)
by: Fang, Sitong, et al.
Published: (2025)
Multi-Agent Collaborative Intelligence: Dual-Dial Control for Reliable LLM Reasoning
by: Chang, Edward Y., et al.
Published: (2025)
by: Chang, Edward Y., et al.
Published: (2025)
Prohibicionismo, grupos sociales "a riesgo" y autoritarismo institucional: la censura social hacia los "microtraficantes"
by: Paolo Scalia
Published: (2005)
by: Paolo Scalia
Published: (2005)
Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models
by: Amiri-Margavi, Alireza, et al.
Published: (2024)
by: Amiri-Margavi, Alireza, et al.
Published: (2024)
Importance Weighting Can Help Large Language Models Self-Improve
by: Jiang, Chunyang, et al.
Published: (2024)
by: Jiang, Chunyang, et al.
Published: (2024)
Can Post-Training Transform LLMs into Causal Reasoners?
by: Chen, Junqi, et al.
Published: (2026)
by: Chen, Junqi, et al.
Published: (2026)
Mixed layer depth values for the n = 1968 modern dinocyst database
by: Wu, Xiner
Published: (2025)
by: Wu, Xiner
Published: (2025)
Logical Reasoning with Outcome Reward Models for Test-Time Scaling
by: Thatikonda, Ramya Keerthy, et al.
Published: (2025)
by: Thatikonda, Ramya Keerthy, et al.
Published: (2025)
ActorMind: Emulating Human Actor Reasoning for Speech Role-Playing
by: Chen, Xi, et al.
Published: (2026)
by: Chen, Xi, et al.
Published: (2026)
A Review of High Temperature Durability and Laser Cladding Crack Suppression Techniques for Nickel‐Based Alloys: Mechanisms and Strategies
by: Xiner Li, et al.
Published: (2025)
by: Xiner Li, et al.
Published: (2025)
Learning Weakly Communicating Average-Reward CMDPs: Strong Duality and Improved Regret
by: Yu, Kihyun, et al.
Published: (2026)
by: Yu, Kihyun, et al.
Published: (2026)
Reward-Shifted Speculative Sampling Is An Efficient Test-Time Weak-to-Strong Aligner
by: Li, Bolian, et al.
Published: (2025)
by: Li, Bolian, et al.
Published: (2025)
LRAS: Advanced Legal Reasoning with Agentic Search
by: Zhou, Yujin, et al.
Published: (2026)
by: Zhou, Yujin, et al.
Published: (2026)
Improving Diffusion Generalization with Weak-to-Strong Segmented Guidance
by: Yuan, Liangyu, et al.
Published: (2026)
by: Yuan, Liangyu, et al.
Published: (2026)
A Hashgraph-Inspired Consensus Mechanism for Reliable Multi-Model Reasoning
by: Ogunsina, Kolawole E., et al.
Published: (2025)
by: Ogunsina, Kolawole E., et al.
Published: (2025)
ThinkPatterns-21k: A Systematic Study on the Impact of Thinking Patterns in LLMs
by: Wen, Pengcheng, et al.
Published: (2025)
by: Wen, Pengcheng, et al.
Published: (2025)
Glance-or-Gaze: Incentivizing LMMs to Adaptively Focus Search via Reinforcement Learning
by: Bai, Hongbo, et al.
Published: (2026)
by: Bai, Hongbo, et al.
Published: (2026)
R2-KG: General-Purpose Dual-Agent Framework for Reliable Reasoning on Knowledge Graphs
by: Jo, Sumin, et al.
Published: (2025)
by: Jo, Sumin, et al.
Published: (2025)
Similar Items
-
RAG-Enhanced Collaborative LLM Agents for Drug Discovery
by: Lee, Namkyeong, et al.
Published: (2025) -
MolCap-Arena: A Comprehensive Captioning Benchmark on Language-Enhanced Molecular Property Prediction
by: Edwards, Carl, et al.
Published: (2024) -
AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents
by: De Brouwer, Edward, et al.
Published: (2026) -
Iterative Distillation for Reward-Guided Fine-Tuning of Diffusion Models in Biomolecular Design
by: Su, Xingyu, et al.
Published: (2025) -
Adding Conditional Control to Diffusion Models with Reinforcement Learning
by: Zhao, Yulai, et al.
Published: (2024)