Is Free Self-Alignment Possible?
Fuente:
arXiv
Guardado en:
| Autores principales: | Adila, Dyah, Shin, Changho, Zhang, Yijing, Sala, Frederic |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Personalize Your LLM: Fake it then Align it
por: Zhang, Yijing, et al.
Publicado: (2025)
por: Zhang, Yijing, et al.
Publicado: (2025)
Zero-Shot Robustification of Zero-Shot Models
por: Adila, Dyah, et al.
Publicado: (2023)
por: Adila, Dyah, et al.
Publicado: (2023)
Multimodal Data Curation via Object Detection and Filter Ensembles
por: Huang, Tzu-Heng, et al.
Publicado: (2024)
por: Huang, Tzu-Heng, et al.
Publicado: (2024)
CrEst: Credibility Estimation for Contexts in LLMs via Weak Supervision
por: Adila, Dyah, et al.
Publicado: (2025)
por: Adila, Dyah, et al.
Publicado: (2025)
Weight Updates as Activation Shifts: A Principled Framework for Steering
por: Adila, Dyah, et al.
Publicado: (2026)
por: Adila, Dyah, et al.
Publicado: (2026)
Weak-to-Strong Generalization Through the Data-Centric Lens
por: Shin, Changho, et al.
Publicado: (2024)
por: Shin, Changho, et al.
Publicado: (2024)
SlopCodeBench: Benchmarking How Coding Agents Degrade Over Long-Horizon Iterative Tasks
por: Orlanski, Gabriel, et al.
Publicado: (2026)
por: Orlanski, Gabriel, et al.
Publicado: (2026)
Evaluating Language Model Context Windows: A "Working Memory" Test and Inference-time Correction
por: Dsouza, Amanda, et al.
Publicado: (2024)
por: Dsouza, Amanda, et al.
Publicado: (2024)
Discovering Bias in Latent Space: An Unsupervised Debiasing Approach
por: Adila, Dyah, et al.
Publicado: (2024)
por: Adila, Dyah, et al.
Publicado: (2024)
Reasoning Boosts Opinion Alignment in LLMs
por: Berdoz, Frédéric, et al.
Publicado: (2026)
por: Berdoz, Frédéric, et al.
Publicado: (2026)
Extensive Self-Contrast Enables Feedback-Free Language Model Alignment
por: Liu, Xiao, et al.
Publicado: (2024)
por: Liu, Xiao, et al.
Publicado: (2024)
SeMe: Training-Free Language Model Merging via Semantic Alignment
por: Gu, Jian, et al.
Publicado: (2025)
por: Gu, Jian, et al.
Publicado: (2025)
Grow, Don't Overwrite: Fine-tuning Without Forgetting
por: Adila, Dyah, et al.
Publicado: (2026)
por: Adila, Dyah, et al.
Publicado: (2026)
OTTER: Effortless Label Distribution Adaptation of Zero-shot Models
por: Shin, Changho, et al.
Publicado: (2024)
por: Shin, Changho, et al.
Publicado: (2024)
The ALCHEmist: Automated Labeling 500x CHEaper Than LLM Data Annotators
por: Huang, Tzu-Heng, et al.
Publicado: (2024)
por: Huang, Tzu-Heng, et al.
Publicado: (2024)
SelfCodeAlign: Self-Alignment for Code Generation
por: Wei, Yuxiang, et al.
Publicado: (2024)
por: Wei, Yuxiang, et al.
Publicado: (2024)
REFA: Reference Free Alignment for multi-preference optimization
por: Gupta, Taneesh, et al.
Publicado: (2024)
por: Gupta, Taneesh, et al.
Publicado: (2024)
A Theoretical Understanding of Self-Correction through In-context Alignment
por: Wang, Yifei, et al.
Publicado: (2024)
por: Wang, Yifei, et al.
Publicado: (2024)
Self-Improvement Towards Pareto Optimality: Mitigating Preference Conflicts in Multi-Objective Alignment
por: Li, Moxin, et al.
Publicado: (2025)
por: Li, Moxin, et al.
Publicado: (2025)
Self-Refining Language Model Anonymizers via Adversarial Distillation
por: Kim, Kyuyoung, et al.
Publicado: (2025)
por: Kim, Kyuyoung, et al.
Publicado: (2025)
SALMON: Self-Alignment with Instructable Reward Models
por: Sun, Zhiqing, et al.
Publicado: (2023)
por: Sun, Zhiqing, et al.
Publicado: (2023)
Self-Augmented Preference Optimization: Off-Policy Paradigms for Language Model Alignment
por: Yin, Yueqin, et al.
Publicado: (2024)
por: Yin, Yueqin, et al.
Publicado: (2024)
Translating Expert Intuition into Quantifiable Features: Encode Investigator Domain Knowledge via LLM for Enhanced Predictive Analytics
por: Jing, Phoebe, et al.
Publicado: (2024)
por: Jing, Phoebe, et al.
Publicado: (2024)
Alchemist: Towards the Design of Efficient Online Continual Learning System
por: Huang, Yuyang, et al.
Publicado: (2025)
por: Huang, Yuyang, et al.
Publicado: (2025)
TARDIS: Mitigating Temporal Misalignment via Representation Steering
por: Shin, Changho, et al.
Publicado: (2025)
por: Shin, Changho, et al.
Publicado: (2025)
LLM-Integrated Bayesian State Space Models for Multimodal Time-Series Forecasting
por: Cho, Sungjun, et al.
Publicado: (2025)
por: Cho, Sungjun, et al.
Publicado: (2025)
Multilingual Safety Alignment via Self-Distillation
por: Qin, Ruiyang, et al.
Publicado: (2026)
por: Qin, Ruiyang, et al.
Publicado: (2026)
Safety Alignment as Continual Learning: Mitigating the Alignment Tax via Orthogonal Gradient Projection
por: Sun, Guanglong, et al.
Publicado: (2026)
por: Sun, Guanglong, et al.
Publicado: (2026)
ReVISE: Learning to Refine at Test-Time via Intrinsic Self-Verification
por: Lee, Hyunseok, et al.
Publicado: (2025)
por: Lee, Hyunseok, et al.
Publicado: (2025)
Large Language Models are Superpositions of All Characters: Attaining Arbitrary Role-play via Self-Alignment
por: Lu, Keming, et al.
Publicado: (2024)
por: Lu, Keming, et al.
Publicado: (2024)
How to Leverage Demonstration Data in Alignment for Large Language Model? A Self-Imitation Learning Perspective
por: Xiao, Teng, et al.
Publicado: (2024)
por: Xiao, Teng, et al.
Publicado: (2024)
Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment
por: Kim, Dongyoung, et al.
Publicado: (2024)
por: Kim, Dongyoung, et al.
Publicado: (2024)
SelfCite: Self-Supervised Alignment for Context Attribution in Large Language Models
por: Chuang, Yung-Sung, et al.
Publicado: (2025)
por: Chuang, Yung-Sung, et al.
Publicado: (2025)
Self-Supervised Visual Preference Alignment
por: Zhu, Ke, et al.
Publicado: (2024)
por: Zhu, Ke, et al.
Publicado: (2024)
Baichuan Alignment Technical Report
por: Lin, Mingan, et al.
Publicado: (2024)
por: Lin, Mingan, et al.
Publicado: (2024)
Self-Play Preference Optimization for Language Model Alignment
por: Wu, Yue, et al.
Publicado: (2024)
por: Wu, Yue, et al.
Publicado: (2024)
Test-Time Scaling Makes Overtraining Compute-Optimal
por: Roberts, Nicholas, et al.
Publicado: (2026)
por: Roberts, Nicholas, et al.
Publicado: (2026)
MULTIVERSE: Exposing Large Language Model Alignment Problems in Diverse Worlds
por: Jin, Xiaolong, et al.
Publicado: (2024)
por: Jin, Xiaolong, et al.
Publicado: (2024)
LLMs Are Already Good Tutors: Training-Free Prompt Optimization for Pedagogical Math Tutoring
por: Lee, Unggi, et al.
Publicado: (2026)
por: Lee, Unggi, et al.
Publicado: (2026)
Energy-Based Reward Models for Robust Language Model Alignment
por: Lochab, Anamika, et al.
Publicado: (2025)
por: Lochab, Anamika, et al.
Publicado: (2025)
Ejemplares similares
-
Personalize Your LLM: Fake it then Align it
por: Zhang, Yijing, et al.
Publicado: (2025) -
Zero-Shot Robustification of Zero-Shot Models
por: Adila, Dyah, et al.
Publicado: (2023) -
Multimodal Data Curation via Object Detection and Filter Ensembles
por: Huang, Tzu-Heng, et al.
Publicado: (2024) -
CrEst: Credibility Estimation for Contexts in LLMs via Weak Supervision
por: Adila, Dyah, et al.
Publicado: (2025) -
Weight Updates as Activation Shifts: A Principled Framework for Steering
por: Adila, Dyah, et al.
Publicado: (2026)