A Practical Analysis of Human Alignment with *PO
Fuente:
arXiv
Guardado en:
| Autores principales: | Ahrabian, Kian, Lin, Xihui, Patra, Barun, Chaudhary, Vishrav, Benhaim, Alon, Pujara, Jay, Song, Xia |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
On The Adaptation of Unlimiformer for Decoder-Only Transformers
por: Ahrabian, Kian, et al.
Publicado: (2024)
por: Ahrabian, Kian, et al.
Publicado: (2024)
Scaling Optimal LR Across Token Horizons
por: Bjorck, Johan, et al.
Publicado: (2024)
por: Bjorck, Johan, et al.
Publicado: (2024)
A Systematic Analysis of Base Model Choice for Reward Modeling
por: Ahrabian, Kian, et al.
Publicado: (2025)
por: Ahrabian, Kian, et al.
Publicado: (2025)
Toward Better Temporal Structures for Geopolitical Events Forecasting
por: Ahrabian, Kian, et al.
Publicado: (2026)
por: Ahrabian, Kian, et al.
Publicado: (2026)
Scaling Laws for Multilingual Language Models
por: He, Yifei, et al.
Publicado: (2024)
por: He, Yifei, et al.
Publicado: (2024)
S2-Attention: Hardware-Aware Context Sharding Among Attention Heads
por: Lin, Xihui, et al.
Publicado: (2024)
por: Lin, Xihui, et al.
Publicado: (2024)
A Glitch in the Matrix? Locating and Detecting Language Model Grounding with Fakepedia
por: Monea, Giovanni, et al.
Publicado: (2023)
por: Monea, Giovanni, et al.
Publicado: (2023)
POROver: Improving Safety and Reducing Overrefusal in Large Language Models with Overgeneration and Preference Optimization
por: Karaman, Batuhan K., et al.
Publicado: (2024)
por: Karaman, Batuhan K., et al.
Publicado: (2024)
The Curious Case of Nonverbal Abstract Reasoning with Multi-Modal Large Language Models
por: Ahrabian, Kian, et al.
Publicado: (2024)
por: Ahrabian, Kian, et al.
Publicado: (2024)
ComPO: Preference Alignment via Comparison Oracles
por: Chen, Peter, et al.
Publicado: (2025)
por: Chen, Peter, et al.
Publicado: (2025)
Preference Ranking Optimization for Human Alignment
por: Song, Feifan, et al.
Publicado: (2023)
por: Song, Feifan, et al.
Publicado: (2023)
Alignment Tuning for Large Language Models: A Data-Centric Lens on Alignment Data Pipelines
por: Song, Hwanjun
Publicado: (2026)
por: Song, Hwanjun
Publicado: (2026)
Scaling Data Diversity for Fine-Tuning Language Models in Human Alignment
por: Song, Feifan, et al.
Publicado: (2024)
por: Song, Feifan, et al.
Publicado: (2024)
BFS-PO: Best-First Search for Large Reasoning Models
por: Parascandolo, Fiorenzo, et al.
Publicado: (2026)
por: Parascandolo, Fiorenzo, et al.
Publicado: (2026)
SparsePO: Controlling Preference Alignment of LLMs via Sparse Token Masks
por: Christopoulou, Fenia, et al.
Publicado: (2024)
por: Christopoulou, Fenia, et al.
Publicado: (2024)
Can LLMs Truly Embody Human Personality? Analyzing AI and Human Behavior Alignment in Dispute Resolution
por: Kwon, Deuksin, et al.
Publicado: (2026)
por: Kwon, Deuksin, et al.
Publicado: (2026)
Simulated Reasoning is Reasoning
por: Kempt, Hendrik, et al.
Publicado: (2026)
por: Kempt, Hendrik, et al.
Publicado: (2026)
Northeastern Uni at Multilingual Counterspeech Generation: Enhancing Counter Speech Generation with LLM Alignment through Direct Preference Optimization
por: Wadhwa, Sahil, et al.
Publicado: (2024)
por: Wadhwa, Sahil, et al.
Publicado: (2024)
SAHOO: Safeguarded Alignment for High-Order Optimization Objectives in Recursive Self-Improvement
por: Sahoo, Subramanyam, et al.
Publicado: (2026)
por: Sahoo, Subramanyam, et al.
Publicado: (2026)
Self-Discover: Large Language Models Self-Compose Reasoning Structures
por: Zhou, Pei, et al.
Publicado: (2024)
por: Zhou, Pei, et al.
Publicado: (2024)
Evaluating Human Alignment and Model Faithfulness of LLM Rationale
por: Fayyaz, Mohsen, et al.
Publicado: (2024)
por: Fayyaz, Mohsen, et al.
Publicado: (2024)
Alignment-Constrained Dynamic Pruning for LLMs: Identifying and Preserving Alignment-Critical Circuits
por: Patel, Dev, et al.
Publicado: (2025)
por: Patel, Dev, et al.
Publicado: (2025)
Uncovering Factor Level Preferences to Improve Human-Model Alignment
por: Oh, Juhyun, et al.
Publicado: (2024)
por: Oh, Juhyun, et al.
Publicado: (2024)
Human-Instruction-Free LLM Self-Alignment with Limited Samples
por: Guo, Hongyi, et al.
Publicado: (2024)
por: Guo, Hongyi, et al.
Publicado: (2024)
SDA: Steering-Driven Distribution Alignment for Open LLMs without Fine-Tuning
por: Xia, Wei, et al.
Publicado: (2025)
por: Xia, Wei, et al.
Publicado: (2025)
Same Task, More Tokens: the Impact of Input Length on the Reasoning Performance of Large Language Models
por: Levy, Mosh, et al.
Publicado: (2024)
por: Levy, Mosh, et al.
Publicado: (2024)
Faithfulness Serum: Mitigating the Faithfulness Gap in Textual Explanations of LLM Decisions via Attribution Guidance
por: Alon, Bar, et al.
Publicado: (2026)
por: Alon, Bar, et al.
Publicado: (2026)
Psychometric Alignment: Capturing Human Knowledge Distributions via Language Models
por: He-Yueya, Joy, et al.
Publicado: (2024)
por: He-Yueya, Joy, et al.
Publicado: (2024)
Implicature in Interaction: Understanding Implicature Improves Alignment in Human-LLM Interaction
por: Hota, Asutosh, et al.
Publicado: (2025)
por: Hota, Asutosh, et al.
Publicado: (2025)
Human-Alignment and Calibration of Inference-Time Uncertainty in Large Language Models
por: Moore, Kyle, et al.
Publicado: (2025)
por: Moore, Kyle, et al.
Publicado: (2025)
HumanLM: Simulating Users with State Alignment Beats Response Imitation
por: Wu, Shirley, et al.
Publicado: (2026)
por: Wu, Shirley, et al.
Publicado: (2026)
Surprising Resilience of Science During a Global Pandemic: A Large-Scale Descriptive Analysis
por: Ahrabian, Kian, et al.
Publicado: (2024)
por: Ahrabian, Kian, et al.
Publicado: (2024)
Discourse vs emissions: Analysis of corporate narratives, symbolic practices, and mimicry through LLMs
por: Hassani, Bertrand Kian, et al.
Publicado: (2025)
por: Hassani, Bertrand Kian, et al.
Publicado: (2025)
Dropouts in Confidence: Moral Uncertainty in Human-LLM Alignment
por: Kwon, Jea, et al.
Publicado: (2025)
por: Kwon, Jea, et al.
Publicado: (2025)
Surprising Resilience of Scientific Publication during a Global Pandemic: A Large‐Scale Bibliometric Analysis
por: Casandra Rusti, et al.
Publicado: (2025)
por: Casandra Rusti, et al.
Publicado: (2025)
Negating Negatives: Alignment with Human Negative Samples via Distributional Dispreference Optimization
por: Duan, Shitong, et al.
Publicado: (2024)
por: Duan, Shitong, et al.
Publicado: (2024)
The Knowledge Alignment Problem: Bridging Human and External Knowledge for Large Language Models
por: Zhang, Shuo, et al.
Publicado: (2023)
por: Zhang, Shuo, et al.
Publicado: (2023)
Psycholinguistic Word Features: a New Approach for the Evaluation of LLMs Alignment with Humans
por: Conde, Javier, et al.
Publicado: (2025)
por: Conde, Javier, et al.
Publicado: (2025)
Diverse Human Value Alignment for Large Language Models via Ethical Reasoning
por: Wang, Jiahao, et al.
Publicado: (2025)
por: Wang, Jiahao, et al.
Publicado: (2025)
Elo-Evolve: A Co-evolutionary Framework for Language Model Alignment
por: Zhao, Jing, et al.
Publicado: (2026)
por: Zhao, Jing, et al.
Publicado: (2026)
Ejemplares similares
-
On The Adaptation of Unlimiformer for Decoder-Only Transformers
por: Ahrabian, Kian, et al.
Publicado: (2024) -
Scaling Optimal LR Across Token Horizons
por: Bjorck, Johan, et al.
Publicado: (2024) -
A Systematic Analysis of Base Model Choice for Reward Modeling
por: Ahrabian, Kian, et al.
Publicado: (2025) -
Toward Better Temporal Structures for Geopolitical Events Forecasting
por: Ahrabian, Kian, et al.
Publicado: (2026) -
Scaling Laws for Multilingual Language Models
por: He, Yifei, et al.
Publicado: (2024)