Voting with the Graph: Stable RLAIF via Topological Consistency Maximization
Fuente:
arXiv
Salvato in:
| Autori principali: | Liu, Boyin, Zhang, Zhuo, Huang, Sen, Xie, Lipeng, Fu, Qingxu, Chen, Haoran, YU, LI, Hu, Tianyi, Liu, Zhaoyang, Ding, Bolin, Zhao, Dongbin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SeeUPO: Sequence-Level Agentic-RL with Convergence Guarantees
di: Hu, Tianyi, et al.
Pubblicazione: (2026)
di: Hu, Tianyi, et al.
Pubblicazione: (2026)
Auto-Rubric: Learning From Implicit Weights to Explicit Rubrics for Reward Modeling
di: Xie, Lipeng, et al.
Pubblicazione: (2025)
di: Xie, Lipeng, et al.
Pubblicazione: (2025)
Investor Composition and the Liquidity Component in the U.S. Corporate Bond Market
di: JIAN LI, et al.
Pubblicazione: (2026)
di: JIAN LI, et al.
Pubblicazione: (2026)
A coupled theory of tropical climatology: warm pool, cold tongue and walker circulation
di: Zhengyu Liu and Boyin Huang
Pubblicazione: (1997)
di: Zhengyu Liu and Boyin Huang
Pubblicazione: (1997)
On the EPA's Radar: The Role of Financial Reports in Environmental Regulatory Oversight
di: BIN LI, et al.
Pubblicazione: (2024)
di: BIN LI, et al.
Pubblicazione: (2024)
Boosting Continuous Control with Consistency Policy
di: Chen, Yuhui, et al.
Pubblicazione: (2023)
di: Chen, Yuhui, et al.
Pubblicazione: (2023)
Geological Record of Late Eocene to Early Oligocene Lithosphere Delamination along the Jinshajiang–Red River Tectonic Zone
di: Zhiqi YU, et al.
Pubblicazione: (2025)
di: Zhiqi YU, et al.
Pubblicazione: (2025)
Unreal-MAP: Unreal-Engine-Based General Platform for Multi-Agent Reinforcement Learning
di: Hu, Tianyi, et al.
Pubblicazione: (2025)
di: Hu, Tianyi, et al.
Pubblicazione: (2025)
Generalizing Consistency Policy to Visual RL with Prioritized Proximal Experience Regularization
di: Li, Haoran, et al.
Pubblicazione: (2024)
di: Li, Haoran, et al.
Pubblicazione: (2024)
CPIG: Leveraging Consistency Policy with Intention Guidance for Multi-agent Exploration
di: Fu, Yuqian, et al.
Pubblicazione: (2024)
di: Fu, Yuqian, et al.
Pubblicazione: (2024)
RLAIF-SPA: Structured AI Feedback for Semantic-Prosodic Alignment in Speech Synthesis
di: Yang, Qing, et al.
Pubblicazione: (2025)
di: Yang, Qing, et al.
Pubblicazione: (2025)
Videos are Sample-Efficient Supervisions: Behavior Cloning from Videos via Latent Representations
di: Liu, Xin, et al.
Pubblicazione: (2025)
di: Liu, Xin, et al.
Pubblicazione: (2025)
Why Does RLAIF Work At All?
di: Young, Robin
Pubblicazione: (2026)
di: Young, Robin
Pubblicazione: (2026)
Balanced Collaborative Exploration via Distributed Topological Graph Voronoi Partition
di: Ding, Tianyi, et al.
Pubblicazione: (2025)
di: Ding, Tianyi, et al.
Pubblicazione: (2025)
P‐97: A Volumetric 3D Display System Based on Coded‐Multiplane PDLC
di: Anran LI, et al.
Pubblicazione: (2025)
di: Anran LI, et al.
Pubblicazione: (2025)
ConRFT: A Reinforced Fine-tuning Method for VLA Models via Consistency Policy
di: Chen, Yuhui, et al.
Pubblicazione: (2025)
di: Chen, Yuhui, et al.
Pubblicazione: (2025)
AgentEvolver: Towards Efficient Self-Evolving Agent System
di: Zhai, Yunpeng, et al.
Pubblicazione: (2025)
di: Zhai, Yunpeng, et al.
Pubblicazione: (2025)
Coevolving with the Other You: Fine-Tuning LLM with Sequential Cooperative Multi-Agent Reinforcement Learning
di: Ma, Hao, et al.
Pubblicazione: (2024)
di: Ma, Hao, et al.
Pubblicazione: (2024)
Target Enclosing Control for Nonholonomic Multi-Agent Systems with Connectivity Maintenance and Collision Avoidance
di: Zheng, Boyin, et al.
Pubblicazione: (2025)
di: Zheng, Boyin, et al.
Pubblicazione: (2025)
ConsisLoRA: Enhancing Content and Style Consistency for LoRA-based Style Transfer
di: Chen, Bolin, et al.
Pubblicazione: (2025)
di: Chen, Bolin, et al.
Pubblicazione: (2025)
Self-Clustering Hierarchical Multi-Agent Reinforcement Learning with Extensible Cooperation Graph
di: Fu, Qingxu, et al.
Pubblicazione: (2024)
di: Fu, Qingxu, et al.
Pubblicazione: (2024)
SWA-LDM: Toward Stealthy Watermarks for Latent Diffusion Models
di: Yang, Zhonghao, et al.
Pubblicazione: (2025)
di: Yang, Zhonghao, et al.
Pubblicazione: (2025)
20‐1: A novel liquid crystal planar display structure based on fringe field effect
di: An-ran LI, et al.
Pubblicazione: (2024)
di: An-ran LI, et al.
Pubblicazione: (2024)
Provably Convergent Plug-and-play Proximal Block Coordinate Descent Method for Hyperspectral Anomaly Detection
di: Liu, Xiaoxia, et al.
Pubblicazione: (2024)
di: Liu, Xiaoxia, et al.
Pubblicazione: (2024)
FM3Q: Factorized Multi-Agent MiniMax Q-Learning for Two-Team Zero-Sum Markov Game
di: Hu, Guangzheng, et al.
Pubblicazione: (2024)
di: Hu, Guangzheng, et al.
Pubblicazione: (2024)
Exploring a Design Framework for Children's Agency through Participatory Design
di: Yang, Boyin, et al.
Pubblicazione: (2026)
di: Yang, Boyin, et al.
Pubblicazione: (2026)
Optimal Persistence Reveals Hidden Topology in Complex Energy Landscapes
di: Zhenpeng, LI
Pubblicazione: (2026)
di: Zhenpeng, LI
Pubblicazione: (2026)
Research on the Application of Electromagnetic Method in the Exploration of Altered Rock‐type Gold Deposits in the East Kunlun Metallogenic Belt
di: Ji'en DONG, et al.
Pubblicazione: (2024)
di: Ji'en DONG, et al.
Pubblicazione: (2024)
Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes
di: Fu, Yuqian, et al.
Pubblicazione: (2026)
di: Fu, Yuqian, et al.
Pubblicazione: (2026)
Offline RLAIF: Piloting VLM Feedback for RL via SFO
di: Beck, Jacob
Pubblicazione: (2025)
di: Beck, Jacob
Pubblicazione: (2025)
Applying RLAIF for Code Generation with API-usage in Lightweight LLMs
di: Dutta, Sujan, et al.
Pubblicazione: (2024)
di: Dutta, Sujan, et al.
Pubblicazione: (2024)
Mirror-Consistency: Harnessing Inconsistency in Majority Voting
di: Huang, Siyuan, et al.
Pubblicazione: (2024)
di: Huang, Siyuan, et al.
Pubblicazione: (2024)
Respecting Temporal-Causal Consistency: Entity-Event Knowledge Graphs for Retrieval-Augmented Generation
di: Zhang, Ze Yu, et al.
Pubblicazione: (2025)
di: Zhang, Ze Yu, et al.
Pubblicazione: (2025)
NeuronsGym: A Hybrid Framework and Benchmark for Robot Tasks with Sim2Real Policy Learning
di: Li, Haoran, et al.
Pubblicazione: (2023)
di: Li, Haoran, et al.
Pubblicazione: (2023)
Curriculum-RLAIF: Curriculum Alignment with Reinforcement Learning from AI Feedback
di: Lin, Jiaye, et al.
Pubblicazione: (2025)
di: Lin, Jiaye, et al.
Pubblicazione: (2025)
Ranked Voting based Self-Consistency of Large Language Models
di: Wang, Weiqin, et al.
Pubblicazione: (2025)
di: Wang, Weiqin, et al.
Pubblicazione: (2025)
Energy Efficiency Maximization for Movable Antenna Communication Systems
di: Ding, Jingze, et al.
Pubblicazione: (2025)
di: Ding, Jingze, et al.
Pubblicazione: (2025)
Wavelet Packet-Based Diffusion Model for Ground Motion Generation with Multi-Conditional Energy and Spectral Matching
di: Ding, Yi, et al.
Pubblicazione: (2026)
di: Ding, Yi, et al.
Pubblicazione: (2026)
On-Chip Trace Detection of Cd2+ and Pb2+ of Deep Seawater Using CMOS-Integrated Low-Noise Transimpedance Amplifiers
di: YU, Yiming
Pubblicazione: (2026)
di: YU, Yiming
Pubblicazione: (2026)
THE VISUALIZATION AND ANALYSIS OF URBAN FACILITY POIS USING NETWORK KERNEL DENSITY ESTIMATION CONSTRAINED BY MULTI-FACTORS
di: WENHAO YU
Pubblicazione: (2014)
di: WENHAO YU
Pubblicazione: (2014)
Documenti analoghi
-
SeeUPO: Sequence-Level Agentic-RL with Convergence Guarantees
di: Hu, Tianyi, et al.
Pubblicazione: (2026) -
Auto-Rubric: Learning From Implicit Weights to Explicit Rubrics for Reward Modeling
di: Xie, Lipeng, et al.
Pubblicazione: (2025) -
Investor Composition and the Liquidity Component in the U.S. Corporate Bond Market
di: JIAN LI, et al.
Pubblicazione: (2026) -
A coupled theory of tropical climatology: warm pool, cold tongue and walker circulation
di: Zhengyu Liu and Boyin Huang
Pubblicazione: (1997) -
On the EPA's Radar: The Role of Financial Reports in Environmental Regulatory Oversight
di: BIN LI, et al.
Pubblicazione: (2024)