MaPPO: Maximum a Posteriori Preference Optimization with Prior Knowledge
Fuente:
arXiv
Guardado en:
| Autores principales: | Lan, Guangchen, Zhang, Sipeng, Wang, Tianle, Zhang, Yuwei, Zhang, Daoan, Wei, Xinpeng, Pan, Xiaoman, Zhang, Hongming, Han, Dong-Jun, Brinton, Christopher G. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Bridging SFT and DPO for Diffusion Model Alignment with Self-Sampling Preference Optimization
por: Zhang, Daoan, et al.
Publicado: (2024)
por: Zhang, Daoan, et al.
Publicado: (2024)
Reinforcement Learning for Scalable and Trustworthy Intelligent Systems
por: Lan, Guangchen
Publicado: (2026)
por: Lan, Guangchen
Publicado: (2026)
Contextual Integrity in LLMs via Reasoning and Reinforcement Learning
por: Lan, Guangchen, et al.
Publicado: (2025)
por: Lan, Guangchen, et al.
Publicado: (2025)
HIP Network: Historical Information Passing Network for Extrapolation Reasoning on Temporal Knowledge Graph
por: He, Yongquan, et al.
Publicado: (2024)
por: He, Yongquan, et al.
Publicado: (2024)
Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy
por: Lan, Guangchen, et al.
Publicado: (2026)
por: Lan, Guangchen, et al.
Publicado: (2026)
MCP: A Control-Theoretic Orchestration Framework for Synergistic Efficiency and Interpretability in Multimodal Large Language Models
por: Zhang, Luyan
Publicado: (2025)
por: Zhang, Luyan
Publicado: (2025)
KSHSeek: Data-Driven Approaches to Mitigating and Detecting Knowledge-Shortcut Hallucinations in Generative Models
por: Liu, Zhongxin, et al.
Publicado: (2025)
por: Liu, Zhongxin, et al.
Publicado: (2025)
Synergy over Discrepancy: A Partition-Based Approach to Multi-Domain LLM Fine-Tuning
por: Ye, Hua, et al.
Publicado: (2025)
por: Ye, Hua, et al.
Publicado: (2025)
Listwise Direct Preference Optimization with Multi-Dimensional Preference Mixing
por: Sun, Yuhui, et al.
Publicado: (2025)
por: Sun, Yuhui, et al.
Publicado: (2025)
Unleashing LLMs in Bayesian Optimization: Preference-Guided Framework for Scientific Discovery
por: Yuan, Xinzhe, et al.
Publicado: (2026)
por: Yuan, Xinzhe, et al.
Publicado: (2026)
Mitigating Cross-Lingual Cultural Inconsistencies in LLMs via Consensus-Driven Preference Optimisation
por: Resck, Lucas, et al.
Publicado: (2026)
por: Resck, Lucas, et al.
Publicado: (2026)
Induce, Align, Predict: Zero-Shot Stance Detection via Cognitive Inductive Reasoning
por: Zhang, Bowen, et al.
Publicado: (2025)
por: Zhang, Bowen, et al.
Publicado: (2025)
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
por: Fadli, Samih
Publicado: (2025)
por: Fadli, Samih
Publicado: (2025)
TRiMS: Real-Time Tracking of Minimal Sufficient Length for Efficient Reasoning via RL
por: Bian, Tingcheng, et al.
Publicado: (2026)
por: Bian, Tingcheng, et al.
Publicado: (2026)
PRISMA: Preference-Reinforced Self-Training Approach for Interpretable Emotionally Intelligent Negotiation Dialogues
por: Kajare, Prajwal Vijay, et al.
Publicado: (2026)
por: Kajare, Prajwal Vijay, et al.
Publicado: (2026)
SECURA: Sigmoid-Enhanced CUR Decomposition with Uninterrupted Retention and Low-Rank Adaptation in Large Language Models
por: Zhang, Yuxuan
Publicado: (2025)
por: Zhang, Yuxuan
Publicado: (2025)
OntoLogX: Ontology-Guided Knowledge Graph Extraction from Cybersecurity Logs with Large Language Models
por: Cotti, Luca, et al.
Publicado: (2025)
por: Cotti, Luca, et al.
Publicado: (2025)
Survey Transfer Learning: Recycling Data with Silicon Responses
por: Amini, Ali
Publicado: (2025)
por: Amini, Ali
Publicado: (2025)
TwinVoice: A Multi-dimensional Benchmark Towards Digital Twins via LLM Persona Simulation
por: Du, Bangde, et al.
Publicado: (2025)
por: Du, Bangde, et al.
Publicado: (2025)
PersonalLLM: Tailoring LLMs to Individual Preferences
por: Zollo, Thomas P., et al.
Publicado: (2024)
por: Zollo, Thomas P., et al.
Publicado: (2024)
RMGAP: Benchmarking the Generalization of Reward Models across Diverse Preferences
por: Zhou, Yangyang, et al.
Publicado: (2026)
por: Zhou, Yangyang, et al.
Publicado: (2026)
Constitution or Collapse? Exploring Constitutional AI with Llama 3-8B
por: Zhang, Xue
Publicado: (2025)
por: Zhang, Xue
Publicado: (2025)
FastGRPO: Accelerating Policy Optimization via Concurrency-aware Speculative Decoding and Online Draft Learning
por: Zhang, Yizhou, et al.
Publicado: (2025)
por: Zhang, Yizhou, et al.
Publicado: (2025)
XDecomposer: Learning Prior-Free Set Decomposition for Multiphase X-ray Diffraction
por: Gao, Hanyu, et al.
Publicado: (2026)
por: Gao, Hanyu, et al.
Publicado: (2026)
Automated Bug Triaging using Instruction-Tuned Large Language Models
por: Kiashemshaki, Kiana, et al.
Publicado: (2025)
por: Kiashemshaki, Kiana, et al.
Publicado: (2025)
Merge-Bench: Resolve Merge Conflicts with Large Language Models
por: Schesch, Benedikt, et al.
Publicado: (2026)
por: Schesch, Benedikt, et al.
Publicado: (2026)
Eyla: Toward an Identity-Anchored LLM Architecture with Integrated Biological Priors -- Vision, Implementation Attempt, and Lessons from AI-Assisted Development
por: Aditto, Arif
Publicado: (2026)
por: Aditto, Arif
Publicado: (2026)
TIAR: Trajectory-Informed Advantage Reweighting for LLM Abstention Learning
por: Pan, Muyu, et al.
Publicado: (2026)
por: Pan, Muyu, et al.
Publicado: (2026)
Evo-DKD: Dual-Knowledge Decoding for Autonomous Ontology Evolution in Large Language Models
por: Raman, Vishal, et al.
Publicado: (2025)
por: Raman, Vishal, et al.
Publicado: (2025)
ACE: Exploring Activation Cosine Similarity and Variance for Accurate and Calibration-Efficient LLM Pruning
por: Mi, Zhendong, et al.
Publicado: (2025)
por: Mi, Zhendong, et al.
Publicado: (2025)
Layer-Aware Embedding Fusion for LLMs in Text Classifications
por: Gwak, Jiho, et al.
Publicado: (2025)
por: Gwak, Jiho, et al.
Publicado: (2025)
On the Influence of Discourse Relations in Persuasive Texts
por: Turk, Nawar, et al.
Publicado: (2025)
por: Turk, Nawar, et al.
Publicado: (2025)
Latent Instruction Representation Alignment: defending against jailbreaks, backdoors and undesired knowledge in LLMs
por: Easley, Eric, et al.
Publicado: (2026)
por: Easley, Eric, et al.
Publicado: (2026)
Calibrated Confidence Estimation for Tabular Question Answering
por: Voss, Lukas
Publicado: (2026)
por: Voss, Lukas
Publicado: (2026)
Automated CAD Modeling Sequence Generation from Text Descriptions via Transformer-Based Large Language Models
por: Liao, Jianxing, et al.
Publicado: (2025)
por: Liao, Jianxing, et al.
Publicado: (2025)
Scalable GPU-Accelerated Euler Characteristic Curves: Optimization and Differentiable Learning for PyTorch
por: Saxena, Udit
Publicado: (2025)
por: Saxena, Udit
Publicado: (2025)
JURY-RL: Votes Propose, Proofs Dispose for Label-Free RLVR
por: Chen, Xinjie, et al.
Publicado: (2026)
por: Chen, Xinjie, et al.
Publicado: (2026)
Descriptive Collision in Sparse Autoencoder Auto-Interpretability: When One Explanation Describes Many Features
por: McCann, Jordan F.
Publicado: (2026)
por: McCann, Jordan F.
Publicado: (2026)
Super Apriel: One Checkpoint, Many Speeds
por: Labs, SLAM, et al.
Publicado: (2026)
por: Labs, SLAM, et al.
Publicado: (2026)
OSCToM: RL-Guided Adversarial Generation for High-Order Theory of Mind
por: Srishty, Sharmin Sultana, et al.
Publicado: (2026)
por: Srishty, Sharmin Sultana, et al.
Publicado: (2026)
Ejemplares similares
-
Bridging SFT and DPO for Diffusion Model Alignment with Self-Sampling Preference Optimization
por: Zhang, Daoan, et al.
Publicado: (2024) -
Reinforcement Learning for Scalable and Trustworthy Intelligent Systems
por: Lan, Guangchen
Publicado: (2026) -
Contextual Integrity in LLMs via Reasoning and Reinforcement Learning
por: Lan, Guangchen, et al.
Publicado: (2025) -
HIP Network: Historical Information Passing Network for Extrapolation Reasoning on Temporal Knowledge Graph
por: He, Yongquan, et al.
Publicado: (2024) -
Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy
por: Lan, Guangchen, et al.
Publicado: (2026)