HAL: Inducing Human-likeness in LLMs with Alignment
Fuente:
arXiv
Salvato in:
| Autori principali: | Hasan, Masum, Zhao, Junjie, Hoque, Ehsan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
LLMs as Zero-shot Graph Learners: Alignment of GNN Representations with LLM Token Embeddings
di: Wang, Duo, et al.
Pubblicazione: (2024)
di: Wang, Duo, et al.
Pubblicazione: (2024)
Enhancing Hyperspace Analogue to Language (HAL) Representations via Attention-Based Pooling for Text Classification
di: Sakour, Ali, et al.
Pubblicazione: (2026)
di: Sakour, Ali, et al.
Pubblicazione: (2026)
Sample-Efficient Alignment for LLMs
di: Liu, Zichen, et al.
Pubblicazione: (2024)
di: Liu, Zichen, et al.
Pubblicazione: (2024)
Aligners: Decoupling LLMs and Alignment
di: Ngweta, Lilian, et al.
Pubblicazione: (2024)
di: Ngweta, Lilian, et al.
Pubblicazione: (2024)
Enhancing Time Series Forecasting via Multi-Level Text Alignment with LLMs
di: Zhao, Taibiao, et al.
Pubblicazione: (2025)
di: Zhao, Taibiao, et al.
Pubblicazione: (2025)
A Nurse is Blue and Elephant is Rugby: Cross Domain Alignment in Large Language Models Reveal Human-like Patterns
di: Yehudai, Asaf, et al.
Pubblicazione: (2024)
di: Yehudai, Asaf, et al.
Pubblicazione: (2024)
Bridging the Creativity Understanding Gap: Small-Scale Human Alignment Enables Expert-Level Humor Ranking in LLMs
di: Zhou, Kuan Lok, et al.
Pubblicazione: (2025)
di: Zhou, Kuan Lok, et al.
Pubblicazione: (2025)
Alignment-Constrained Dynamic Pruning for LLMs: Identifying and Preserving Alignment-Critical Circuits
di: Patel, Dev, et al.
Pubblicazione: (2025)
di: Patel, Dev, et al.
Pubblicazione: (2025)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
di: Liu, Yixin, et al.
Pubblicazione: (2025)
di: Liu, Yixin, et al.
Pubblicazione: (2025)
A Comprehensive Evaluation framework of Alignment Techniques for LLMs
di: Azmat, Muneeza, et al.
Pubblicazione: (2025)
di: Azmat, Muneeza, et al.
Pubblicazione: (2025)
Towards Scalable Automated Alignment of LLMs: A Survey
di: Cao, Boxi, et al.
Pubblicazione: (2024)
di: Cao, Boxi, et al.
Pubblicazione: (2024)
RLTHF: Targeted Human Feedback for LLM Alignment
di: Xu, Yifei, et al.
Pubblicazione: (2025)
di: Xu, Yifei, et al.
Pubblicazione: (2025)
Arithmetic Control of LLMs for Diverse User Preferences: Directional Preference Alignment with Multi-Objective Rewards
di: Wang, Haoxiang, et al.
Pubblicazione: (2024)
di: Wang, Haoxiang, et al.
Pubblicazione: (2024)
Nudging: Inference-time Alignment of LLMs via Guided Decoding
di: Fei, Yu, et al.
Pubblicazione: (2024)
di: Fei, Yu, et al.
Pubblicazione: (2024)
Learning from Reference Answers: Versatile Language Model Alignment without Binary Human Preference Data
di: Zhao, Shuai, et al.
Pubblicazione: (2025)
di: Zhao, Shuai, et al.
Pubblicazione: (2025)
Towards Unified Alignment Between Agents, Humans, and Environment
di: Yang, Zonghan, et al.
Pubblicazione: (2024)
di: Yang, Zonghan, et al.
Pubblicazione: (2024)
Detecting Hallucinations in SpeechLLMs at Inference Time Using Attention Maps
di: Waldendorf, Jonas, et al.
Pubblicazione: (2026)
di: Waldendorf, Jonas, et al.
Pubblicazione: (2026)
Skin-in-the-Game: Decision Making via Multi-Stakeholder Alignment in LLMs
di: Sel, Bilgehan, et al.
Pubblicazione: (2024)
di: Sel, Bilgehan, et al.
Pubblicazione: (2024)
Easy-to-Hard Generalization: Scalable Alignment Beyond Human Supervision
di: Sun, Zhiqing, et al.
Pubblicazione: (2024)
di: Sun, Zhiqing, et al.
Pubblicazione: (2024)
The Alignment Tax: Response Homogenization in Aligned LLMs and Its Implications for Uncertainty Estimation
di: Liu, Mingyi
Pubblicazione: (2026)
di: Liu, Mingyi
Pubblicazione: (2026)
SparsePO: Controlling Preference Alignment of LLMs via Sparse Token Masks
di: Christopoulou, Fenia, et al.
Pubblicazione: (2024)
di: Christopoulou, Fenia, et al.
Pubblicazione: (2024)
Direct Alignment of Draft Model for Speculative Decoding with Chat-Fine-Tuned LLMs
di: Goel, Raghavv, et al.
Pubblicazione: (2024)
di: Goel, Raghavv, et al.
Pubblicazione: (2024)
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection
di: Shairah, Harethah Abu, et al.
Pubblicazione: (2025)
di: Shairah, Harethah Abu, et al.
Pubblicazione: (2025)
UniMaia: Steering Chess Policies with Language for Human-like Play
di: Siu, Sherman, et al.
Pubblicazione: (2026)
di: Siu, Sherman, et al.
Pubblicazione: (2026)
C-3PO: Compact Plug-and-Play Proxy Optimization to Achieve Human-like Retrieval-Augmented Generation
di: Chen, Guoxin, et al.
Pubblicazione: (2025)
di: Chen, Guoxin, et al.
Pubblicazione: (2025)
Beyond Labels: Aligning Large Language Models with Human-like Reasoning
di: Kabir, Muhammad Rafsan, et al.
Pubblicazione: (2024)
di: Kabir, Muhammad Rafsan, et al.
Pubblicazione: (2024)
Human Inspired Progressive Alignment and Comparative Learning for Grounded Word Acquisition
di: Bao, Yuwei, et al.
Pubblicazione: (2023)
di: Bao, Yuwei, et al.
Pubblicazione: (2023)
Re-evaluating Automatic LLM System Ranking for Alignment with Human Preference
di: Gao, Mingqi, et al.
Pubblicazione: (2024)
di: Gao, Mingqi, et al.
Pubblicazione: (2024)
Frictive Policy Optimization for LLMs: Epistemic Intervention, Risk-Sensitive Control, and Reflective Alignment
di: Pustejovsky, James, et al.
Pubblicazione: (2026)
di: Pustejovsky, James, et al.
Pubblicazione: (2026)
Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs
di: Giordani, Jeremiah
Pubblicazione: (2025)
di: Giordani, Jeremiah
Pubblicazione: (2025)
Picky LLMs and Unreliable RMs: An Empirical Study on Safety Alignment after Instruction Tuning
di: Li, Guanlin, et al.
Pubblicazione: (2025)
di: Li, Guanlin, et al.
Pubblicazione: (2025)
NeuronTune: Fine-Grained Neuron Modulation for Balanced Safety-Utility Alignment in LLMs
di: Pan, Birong, et al.
Pubblicazione: (2025)
di: Pan, Birong, et al.
Pubblicazione: (2025)
Guided Self-Evolving LLMs with Minimal Human Supervision
di: Yu, Wenhao, et al.
Pubblicazione: (2025)
di: Yu, Wenhao, et al.
Pubblicazione: (2025)
Model Merging and Safety Alignment: One Bad Model Spoils the Bunch
di: Hammoud, Hasan Abed Al Kader, et al.
Pubblicazione: (2024)
di: Hammoud, Hasan Abed Al Kader, et al.
Pubblicazione: (2024)
LLMs on a Budget? Say HOLA
di: Siddiqui, Zohaib Hasan, et al.
Pubblicazione: (2025)
di: Siddiqui, Zohaib Hasan, et al.
Pubblicazione: (2025)
Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators
di: Liu, Yinhong, et al.
Pubblicazione: (2024)
di: Liu, Yinhong, et al.
Pubblicazione: (2024)
Are the Values of LLMs Structurally Aligned with Humans? A Causal Perspective
di: Kang, Yipeng, et al.
Pubblicazione: (2024)
di: Kang, Yipeng, et al.
Pubblicazione: (2024)
How Likely Do LLMs with CoT Mimic Human Reasoning?
di: Bao, Guangsheng, et al.
Pubblicazione: (2024)
di: Bao, Guangsheng, et al.
Pubblicazione: (2024)
One STEP at a time: Language Agents are Stepwise Planners
di: Nguyen, Minh, et al.
Pubblicazione: (2024)
di: Nguyen, Minh, et al.
Pubblicazione: (2024)
When Numbers Tell Half the Story: Human-Metric Alignment in Topic Model Evaluation
di: Prouteau, Thibault, et al.
Pubblicazione: (2026)
di: Prouteau, Thibault, et al.
Pubblicazione: (2026)
Documenti analoghi
-
LLMs as Zero-shot Graph Learners: Alignment of GNN Representations with LLM Token Embeddings
di: Wang, Duo, et al.
Pubblicazione: (2024) -
Enhancing Hyperspace Analogue to Language (HAL) Representations via Attention-Based Pooling for Text Classification
di: Sakour, Ali, et al.
Pubblicazione: (2026) -
Sample-Efficient Alignment for LLMs
di: Liu, Zichen, et al.
Pubblicazione: (2024) -
Aligners: Decoupling LLMs and Alignment
di: Ngweta, Lilian, et al.
Pubblicazione: (2024) -
Enhancing Time Series Forecasting via Multi-Level Text Alignment with LLMs
di: Zhao, Taibiao, et al.
Pubblicazione: (2025)