Aligners: Decoupling LLMs and Alignment
Fuente:
arXiv
Salvato in:
| Autori principali: | Ngweta, Lilian, Agarwal, Mayank, Maity, Subha, Gittens, Alex, Sun, Yuekai, Yurochkin, Mikhail |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Weak Supervision Performance Evaluation via Partial Identification
di: Polo, Felipe Maia, et al.
Pubblicazione: (2023)
di: Polo, Felipe Maia, et al.
Pubblicazione: (2023)
tinyBenchmarks: evaluating LLMs with fewer examples
di: Polo, Felipe Maia, et al.
Pubblicazione: (2024)
di: Polo, Felipe Maia, et al.
Pubblicazione: (2024)
Bridging Human and LLM Judgments: Understanding and Narrowing the Gap
di: Polo, Felipe Maia, et al.
Pubblicazione: (2025)
di: Polo, Felipe Maia, et al.
Pubblicazione: (2025)
Efficient multi-prompt evaluation of LLMs
di: Polo, Felipe Maia, et al.
Pubblicazione: (2024)
di: Polo, Felipe Maia, et al.
Pubblicazione: (2024)
Aligner: Efficient Alignment by Learning to Correct
di: Ji, Jiaming, et al.
Pubblicazione: (2024)
di: Ji, Jiaming, et al.
Pubblicazione: (2024)
NeMo-Aligner: Scalable Toolkit for Efficient Model Alignment
di: Shen, Gerald, et al.
Pubblicazione: (2024)
di: Shen, Gerald, et al.
Pubblicazione: (2024)
Prompt Exploration with Prompt Regression
di: Feffer, Michael, et al.
Pubblicazione: (2024)
di: Feffer, Michael, et al.
Pubblicazione: (2024)
Out-of-Distribution Detection using Synthetic Data Generation
di: Abbas, Momin, et al.
Pubblicazione: (2025)
di: Abbas, Momin, et al.
Pubblicazione: (2025)
Stream Aligner: Efficient Sentence-Level Alignment via Distribution Induction
di: Lou, Hantao, et al.
Pubblicazione: (2025)
di: Lou, Hantao, et al.
Pubblicazione: (2025)
Limitations of refinement methods for weak to strong generalization
di: Somerstep, Seamus, et al.
Pubblicazione: (2025)
di: Somerstep, Seamus, et al.
Pubblicazione: (2025)
Sloth: scaling laws for LLM skills to predict multi-benchmark performance across families
di: Polo, Felipe Maia, et al.
Pubblicazione: (2024)
di: Polo, Felipe Maia, et al.
Pubblicazione: (2024)
Sample-Efficient Alignment for LLMs
di: Liu, Zichen, et al.
Pubblicazione: (2024)
di: Liu, Zichen, et al.
Pubblicazione: (2024)
A Comment On "The Illusion of Thinking": Reframing the Reasoning Cliff as an Agentic Gap
di: Khan, Sheraz, et al.
Pubblicazione: (2025)
di: Khan, Sheraz, et al.
Pubblicazione: (2025)
Proxy-RLHF: Decoupling Generation and Alignment in Large Language Model with Proxy
di: Zhu, Yu, et al.
Pubblicazione: (2024)
di: Zhu, Yu, et al.
Pubblicazione: (2024)
Languages are Modalities: Cross-Lingual Alignment via Encoder Injection
di: Agarwal, Rajan, et al.
Pubblicazione: (2025)
di: Agarwal, Rajan, et al.
Pubblicazione: (2025)
Enhancing Time Series Forecasting via Multi-Level Text Alignment with LLMs
di: Zhao, Taibiao, et al.
Pubblicazione: (2025)
di: Zhao, Taibiao, et al.
Pubblicazione: (2025)
Towards Scalable Automated Alignment of LLMs: A Survey
di: Cao, Boxi, et al.
Pubblicazione: (2024)
di: Cao, Boxi, et al.
Pubblicazione: (2024)
Automated Meta Prompt Engineering for Alignment with the Theory of Mind
di: Baughman, Aaron, et al.
Pubblicazione: (2025)
di: Baughman, Aaron, et al.
Pubblicazione: (2025)
Alignment-Constrained Dynamic Pruning for LLMs: Identifying and Preserving Alignment-Critical Circuits
di: Patel, Dev, et al.
Pubblicazione: (2025)
di: Patel, Dev, et al.
Pubblicazione: (2025)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
di: Liu, Yixin, et al.
Pubblicazione: (2025)
di: Liu, Yixin, et al.
Pubblicazione: (2025)
HAL: Inducing Human-likeness in LLMs with Alignment
di: Hasan, Masum, et al.
Pubblicazione: (2026)
di: Hasan, Masum, et al.
Pubblicazione: (2026)
Aligner-Guided Training Paradigm: Advancing Text-to-Speech Models with Aligner Guided Duration
di: Lou, Haowei, et al.
Pubblicazione: (2024)
di: Lou, Haowei, et al.
Pubblicazione: (2024)
Meta-Aligner: Bidirectional Preference-Policy Optimization for Multi-Objective LLMs Alignment
di: Xu, Wenzhe, et al.
Pubblicazione: (2026)
di: Xu, Wenzhe, et al.
Pubblicazione: (2026)
A Comprehensive Evaluation framework of Alignment Techniques for LLMs
di: Azmat, Muneeza, et al.
Pubblicazione: (2025)
di: Azmat, Muneeza, et al.
Pubblicazione: (2025)
Where Does Toxicity Live? Mechanistic Localization and Targeted Suppression in Language Models
di: Beniwal, Himanshu, et al.
Pubblicazione: (2026)
di: Beniwal, Himanshu, et al.
Pubblicazione: (2026)
Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning
di: Beutel, Alex, et al.
Pubblicazione: (2024)
di: Beutel, Alex, et al.
Pubblicazione: (2024)
Nudging: Inference-time Alignment of LLMs via Guided Decoding
di: Fei, Yu, et al.
Pubblicazione: (2024)
di: Fei, Yu, et al.
Pubblicazione: (2024)
In-Context Explainers: Harnessing LLMs for Explaining Black Box Models
di: Kroeger, Nicholas, et al.
Pubblicazione: (2023)
di: Kroeger, Nicholas, et al.
Pubblicazione: (2023)
Compress then Serve: Serving Thousands of LoRA Adapters with Little Overhead
di: Brüel-Gabrielsson, Rickard, et al.
Pubblicazione: (2024)
di: Brüel-Gabrielsson, Rickard, et al.
Pubblicazione: (2024)
Selective Prompting Tuning for Personalized Conversations with LLMs
di: Huang, Qiushi, et al.
Pubblicazione: (2024)
di: Huang, Qiushi, et al.
Pubblicazione: (2024)
Skin-in-the-Game: Decision Making via Multi-Stakeholder Alignment in LLMs
di: Sel, Bilgehan, et al.
Pubblicazione: (2024)
di: Sel, Bilgehan, et al.
Pubblicazione: (2024)
Fragile Thoughts: How Large Language Models Handle Chain-of-Thought Perturbations
di: Aravindan, Ashwath Vaithinathan, et al.
Pubblicazione: (2026)
di: Aravindan, Ashwath Vaithinathan, et al.
Pubblicazione: (2026)
Quokka: Accelerating Program Verification with LLMs via Invariant Synthesis
di: Wei, Anjiang, et al.
Pubblicazione: (2025)
di: Wei, Anjiang, et al.
Pubblicazione: (2025)
SparsePO: Controlling Preference Alignment of LLMs via Sparse Token Masks
di: Christopoulou, Fenia, et al.
Pubblicazione: (2024)
di: Christopoulou, Fenia, et al.
Pubblicazione: (2024)
Direct Alignment of Draft Model for Speculative Decoding with Chat-Fine-Tuned LLMs
di: Goel, Raghavv, et al.
Pubblicazione: (2024)
di: Goel, Raghavv, et al.
Pubblicazione: (2024)
The Alignment Tax: Response Homogenization in Aligned LLMs and Its Implications for Uncertainty Estimation
di: Liu, Mingyi
Pubblicazione: (2026)
di: Liu, Mingyi
Pubblicazione: (2026)
Deep Knowledge-Infusion For Explainable Depression Detection
di: Dalal, Sumit, et al.
Pubblicazione: (2024)
di: Dalal, Sumit, et al.
Pubblicazione: (2024)
Error Taxonomy-Guided Prompt Optimization
di: Singh, Mayank, et al.
Pubblicazione: (2026)
di: Singh, Mayank, et al.
Pubblicazione: (2026)
LLMs as Zero-shot Graph Learners: Alignment of GNN Representations with LLM Token Embeddings
di: Wang, Duo, et al.
Pubblicazione: (2024)
di: Wang, Duo, et al.
Pubblicazione: (2024)
Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs
di: Giordani, Jeremiah
Pubblicazione: (2025)
di: Giordani, Jeremiah
Pubblicazione: (2025)
Documenti analoghi
-
Weak Supervision Performance Evaluation via Partial Identification
di: Polo, Felipe Maia, et al.
Pubblicazione: (2023) -
tinyBenchmarks: evaluating LLMs with fewer examples
di: Polo, Felipe Maia, et al.
Pubblicazione: (2024) -
Bridging Human and LLM Judgments: Understanding and Narrowing the Gap
di: Polo, Felipe Maia, et al.
Pubblicazione: (2025) -
Efficient multi-prompt evaluation of LLMs
di: Polo, Felipe Maia, et al.
Pubblicazione: (2024) -
Aligner: Efficient Alignment by Learning to Correct
di: Ji, Jiaming, et al.
Pubblicazione: (2024)