Act-Adaptive Margin: Dynamically Calibrating Reward Models for Subjective Ambiguity
Fuente:
arXiv
Guardado en:
| Autores principales: | Fang, Feiteng, Chen, Dingwei, Huang, Xiang, Lin, Ting-En, Wu, Yuchuan, Liu, Xiong, Ye, Xinge, Liu, Ziqiang, Zhang, Haonan, Zhu, Liang, Alinejad-Rokny, Hamid, Yang, Min, Li, Yongbin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
CPO: Addressing Reward Ambiguity in Role-playing Dialogue via Comparative Policy Optimization
por: Ye, Xinge, et al.
Publicado: (2025)
por: Ye, Xinge, et al.
Publicado: (2025)
Expanding before Inferring: Enhancing Factuality in Large Language Models through Premature Layers Interpolation
por: Chen, Dingwei, et al.
Publicado: (2025)
por: Chen, Dingwei, et al.
Publicado: (2025)
Lower Layers Matter: Alleviating Hallucination via Multi-Layer Fusion Contrastive Decoding with Truthfulness Refocused
por: Chen, Dingwei, et al.
Publicado: (2024)
por: Chen, Dingwei, et al.
Publicado: (2024)
OpenOmni: Advancing Open-Source Omnimodal Large Language Models with Progressive Multimodal Alignment and Real-Time Self-Aware Emotional Speech Synthesis
por: Luo, Run, et al.
Publicado: (2025)
por: Luo, Run, et al.
Publicado: (2025)
Reverse Preference Optimization for Complex Instruction Following
por: Huang, Xiang, et al.
Publicado: (2025)
por: Huang, Xiang, et al.
Publicado: (2025)
PLOT: Enhancing Preference Learning via Optimal Transport
por: Zhu, Liang, et al.
Publicado: (2026)
por: Zhu, Liang, et al.
Publicado: (2026)
OmniCharacter: Towards Immersive Role-Playing Agents with Seamless Speech-Language Personality Interaction
por: Zhang, Haonan, et al.
Publicado: (2025)
por: Zhang, Haonan, et al.
Publicado: (2025)
CollectiveSFT: Scaling Large Language Models for Chinese Medical Benchmark with Collective Instructions in Healthcare
por: Zhu, Jingwei, et al.
Publicado: (2024)
por: Zhu, Jingwei, et al.
Publicado: (2024)
How chromatin interactions shed light on interpreting non-coding genomic variants: opportunities and future direc-tions
por: Liang, Yuheng, et al.
Publicado: (2024)
por: Liang, Yuheng, et al.
Publicado: (2024)
ToolRM: Towards Agentic Tool-Use Reward Modeling
por: Li, Renhao, et al.
Publicado: (2025)
por: Li, Renhao, et al.
Publicado: (2025)
CLaSp: In-Context Layer Skip for Self-Speculative Decoding
por: Chen, Longze, et al.
Publicado: (2025)
por: Chen, Longze, et al.
Publicado: (2025)
Small Language Model as Data Prospector for Large Language Model
por: Ni, Shiwen, et al.
Publicado: (2024)
por: Ni, Shiwen, et al.
Publicado: (2024)
Enhancing Monte Carlo Dropout Performance for Uncertainty Quantification
por: Asgharnezhad, Hamzeh, et al.
Publicado: (2025)
por: Asgharnezhad, Hamzeh, et al.
Publicado: (2025)
Interpretable graph-based models on multimodal biomedical data integration: A technical review and benchmarking
por: Sadeghi, Alireza, et al.
Publicado: (2025)
por: Sadeghi, Alireza, et al.
Publicado: (2025)
AgentCourt: Simulating Court with Adversarial Evolvable Lawyer Agents
por: Chen, Guhong, et al.
Publicado: (2024)
por: Chen, Guhong, et al.
Publicado: (2024)
SC-Arena: A Natural Language Benchmark for Single-Cell Reasoning with Knowledge-Augmented Evaluation
por: Zhao, Jiahao, et al.
Publicado: (2026)
por: Zhao, Jiahao, et al.
Publicado: (2026)
InteractWeb-Bench: Can Multimodal Agent Escape Blind Execution in Interactive Website Generation?
por: Wang, Qiyao, et al.
Publicado: (2026)
por: Wang, Qiyao, et al.
Publicado: (2026)
RxSafeBench: Identifying Medication Safety Issues of Large Language Models in Simulated Consultation
por: Zhao, Jiahao, et al.
Publicado: (2025)
por: Zhao, Jiahao, et al.
Publicado: (2025)
PatRe: A Full-Stage Office Action and Rebuttal Generation Benchmark for Patent Examination
por: Wang, Qiyao, et al.
Publicado: (2026)
por: Wang, Qiyao, et al.
Publicado: (2026)
ETAGE: Enhanced Test Time Adaptation with Integrated Entropy and Gradient Norms for Robust Model Performance
por: Shamsi, Afshar, et al.
Publicado: (2024)
por: Shamsi, Afshar, et al.
Publicado: (2024)
Agentic Reinforcement Learning with Implicit Step Rewards
por: Liu, Xiaoqian, et al.
Publicado: (2025)
por: Liu, Xiaoqian, et al.
Publicado: (2025)
SemanticST: Spatially Informed Semantic Graph Learning for Clustering, Integration, and Scalable Analysis of Spatial Transcriptomics
por: Zahedi, Roxana, et al.
Publicado: (2025)
por: Zahedi, Roxana, et al.
Publicado: (2025)
Training Superior Sparse Autoencoders for Instruct Models
por: Li, Jiaming, et al.
Publicado: (2025)
por: Li, Jiaming, et al.
Publicado: (2025)
Advancing Medical Image Segmentation with Mini-Net: A Lightweight Solution Tailored for Efficient Segmentation of Medical Images
por: Javed, Syed, et al.
Publicado: (2024)
por: Javed, Syed, et al.
Publicado: (2024)
HiST: Histological Images Reconstruct Tumor Spatial Transcriptomics via MultiScale Fusion Deep Learning
por: Wei Li, et al.
Publicado: (2026)
por: Wei Li, et al.
Publicado: (2026)
P-GenRM: Personalized Generative Reward Model with Test-time User-based Scaling
por: Zhang, Pinyi, et al.
Publicado: (2026)
por: Zhang, Pinyi, et al.
Publicado: (2026)
Automatic Paper Reviewing with Heterogeneous Graph Reasoning over LLM-Simulated Reviewer-Author Debates
por: Li, Shuaimin, et al.
Publicado: (2025)
por: Li, Shuaimin, et al.
Publicado: (2025)
Improving Factual Consistency of News Summarization by Contrastive Preference Optimization
por: Feng, Huawen, et al.
Publicado: (2023)
por: Feng, Huawen, et al.
Publicado: (2023)
CLinNET: An Interpretable and Uncertainty‐Aware Deep Learning Framework for Multi‐Modal Clinical Genomics
por: Ivan Bakhshayeshi, et al.
Publicado: (2026)
por: Ivan Bakhshayeshi, et al.
Publicado: (2026)
STORYTELLER: An Enhanced Plot-Planning Framework for Coherent and Cohesive Story Generation
por: Li, Jiaming, et al.
Publicado: (2025)
por: Li, Jiaming, et al.
Publicado: (2025)
Beyond Quantity: Trajectory Diversity Scaling for Code Agents
por: Chen, Guhong, et al.
Publicado: (2026)
por: Chen, Guhong, et al.
Publicado: (2026)
Aligning Logits Generatively for Principled Black-Box Knowledge Distillation
por: Ma, Jing, et al.
Publicado: (2022)
por: Ma, Jing, et al.
Publicado: (2022)
A Diagnostic Model for Acute Lymphoblastic Leukemia Using Metaheuristics and Deep Learning Methods
por: Rahmani, Amir Masoud, et al.
Publicado: (2024)
por: Rahmani, Amir Masoud, et al.
Publicado: (2024)
Focal Modulation and Bidirectional Feature Fusion Network for Medical Image Segmentation
por: Safdar, Moin, et al.
Publicado: (2025)
por: Safdar, Moin, et al.
Publicado: (2025)
FlowPIE: Test-Time Scientific Idea Evolution with Flow-Guided Literature Exploration
por: Wang, Qiyao, et al.
Publicado: (2026)
por: Wang, Qiyao, et al.
Publicado: (2026)
AlphaDPO: Adaptive Reward Margin for Direct Preference Optimization
por: Wu, Junkang, et al.
Publicado: (2024)
por: Wu, Junkang, et al.
Publicado: (2024)
Who values competent minds and who likes warm hearts? The role of right‐wing authoritarianism and social dominance orientation in shaping voter preferences for political candidates
por: Feiteng Long, et al.
Publicado: (2025)
por: Feiteng Long, et al.
Publicado: (2025)
Ambiguity-aware Point Cloud Segmentation by Adaptive Margin Contrastive Learning
por: Chen, Yang, et al.
Publicado: (2025)
por: Chen, Yang, et al.
Publicado: (2025)
AutoPatent: A Multi-Agent Framework for Automatic Patent Generation
por: Wang, Qiyao, et al.
Publicado: (2024)
por: Wang, Qiyao, et al.
Publicado: (2024)
Reward Modeling from Natural Language Human Feedback
por: Wang, Zongqi, et al.
Publicado: (2026)
por: Wang, Zongqi, et al.
Publicado: (2026)
Ejemplares similares
-
CPO: Addressing Reward Ambiguity in Role-playing Dialogue via Comparative Policy Optimization
por: Ye, Xinge, et al.
Publicado: (2025) -
Expanding before Inferring: Enhancing Factuality in Large Language Models through Premature Layers Interpolation
por: Chen, Dingwei, et al.
Publicado: (2025) -
Lower Layers Matter: Alleviating Hallucination via Multi-Layer Fusion Contrastive Decoding with Truthfulness Refocused
por: Chen, Dingwei, et al.
Publicado: (2024) -
OpenOmni: Advancing Open-Source Omnimodal Large Language Models with Progressive Multimodal Alignment and Real-Time Self-Aware Emotional Speech Synthesis
por: Luo, Run, et al.
Publicado: (2025) -
Reverse Preference Optimization for Complex Instruction Following
por: Huang, Xiang, et al.
Publicado: (2025)