The Alignment Game: A Theory of Long-Horizon Alignment Through Recursive Curation
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Falahati, Ali, Amiri, Mohammad Mohammadi, Larson, Kate, Golab, Lukasz |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Curated Synthetic Data Doesn't Have to Collapse: A Theoretical Study of Generative Retraining with Pluralistic Preferences
par: Falahati, Ali, et autres
Publié: (2026)
par: Falahati, Ali, et autres
Publié: (2026)
Disentangled Structural and Featural Representation for Task-Agnostic Graph Valuation
par: Falahati, Ali, et autres
Publié: (2024)
par: Falahati, Ali, et autres
Publié: (2024)
A Median Perspective on Unlabeled Data for Out-of-Distribution Detection
par: Abbas, Momin, et autres
Publié: (2025)
par: Abbas, Momin, et autres
Publié: (2025)
SentenceKV: Efficient LLM Inference via Sentence-Level Semantic KV Caching
par: Zhu, Yuxuan, et autres
Publié: (2025)
par: Zhu, Yuxuan, et autres
Publié: (2025)
Toward Efficient Influence Function: Dropout as a Compression Tool
par: Zhang, Yuchen, et autres
Publié: (2025)
par: Zhang, Yuchen, et autres
Publié: (2025)
Jackpot! Alignment as a Maximal Lottery
par: Maura-Rivero, Roberto-Rafael, et autres
Publié: (2025)
par: Maura-Rivero, Roberto-Rafael, et autres
Publié: (2025)
Optimal Singular Damage: Efficient LLM Inference in Low Storage Regimes
par: Alipour, Mohammadsajad, et autres
Publié: (2025)
par: Alipour, Mohammadsajad, et autres
Publié: (2025)
DriftXpress: Faster Drifting Models via Projected RKHS Fields
par: Falahati, Ali, et autres
Publié: (2026)
par: Falahati, Ali, et autres
Publié: (2026)
GASTON: Graph-Aware Social Transformer for Online Networks
par: Wloch, Olha, et autres
Publié: (2026)
par: Wloch, Olha, et autres
Publié: (2026)
When and How Human Curation Backfires: Preference Alignment under Multi-Model Self-Consuming Loop
par: Zhang, Yang, et autres
Publié: (2026)
par: Zhang, Yang, et autres
Publié: (2026)
Reward Shaping for Inference-Time Alignment: A Stackelberg Game Perspective
par: Wang, Haichuan, et autres
Publié: (2026)
par: Wang, Haichuan, et autres
Publié: (2026)
Data Advisor: Dynamic Data Curation for Safety Alignment of Large Language Models
par: Wang, Fei, et autres
Publié: (2024)
par: Wang, Fei, et autres
Publié: (2024)
Skin-in-the-Game: Decision Making via Multi-Stakeholder Alignment in LLMs
par: Sel, Bilgehan, et autres
Publié: (2024)
par: Sel, Bilgehan, et autres
Publié: (2024)
Towards a Learning Theory of Representation Alignment
par: Insulla, Francesco, et autres
Publié: (2025)
par: Insulla, Francesco, et autres
Publié: (2025)
Tokenized Bandit for LLM Decoding and Alignment
par: Shin, Suho, et autres
Publié: (2025)
par: Shin, Suho, et autres
Publié: (2025)
Information as Structural Alignment: A Dynamical Theory of Continual Learning
par: Negulescu, Radu
Publié: (2026)
par: Negulescu, Radu
Publié: (2026)
LLM Probability Concentration: How Alignment Shrinks the Generative Horizon
par: Yang, Chenghao, et autres
Publié: (2025)
par: Yang, Chenghao, et autres
Publié: (2025)
OFMU: Optimization-Driven Framework for Machine Unlearning
par: Asif, Sadia, et autres
Publié: (2025)
par: Asif, Sadia, et autres
Publié: (2025)
Towards Reversible Model Merging For Low-rank Weights
par: Alipour, Mohammadsajad, et autres
Publié: (2025)
par: Alipour, Mohammadsajad, et autres
Publié: (2025)
Manifold Approximation leads to Robust Kernel Alignment
par: Islam, Mohammad Tariqul, et autres
Publié: (2025)
par: Islam, Mohammad Tariqul, et autres
Publié: (2025)
SAHOO: Safeguarded Alignment for High-Order Optimization Objectives in Recursive Self-Improvement
par: Sahoo, Subramanyam, et autres
Publié: (2026)
par: Sahoo, Subramanyam, et autres
Publié: (2026)
Direct Alignment with Heterogeneous Preferences
par: Shirali, Ali, et autres
Publié: (2025)
par: Shirali, Ali, et autres
Publié: (2025)
Learning Hyperspectral Images with Curated Text Prompts for Efficient Multimodal Alignment
par: Chatterjee, Abhiroop, et autres
Publié: (2025)
par: Chatterjee, Abhiroop, et autres
Publié: (2025)
Pharmacist: Safety Alignment Data Curation for Large Language Models against Harmful Fine-tuning
par: Liu, Guozhi, et autres
Publié: (2025)
par: Liu, Guozhi, et autres
Publié: (2025)
The Sign Estimator: LLM Alignment in the Face of Choice Heterogeneity
par: Aouad, Ali, et autres
Publié: (2025)
par: Aouad, Ali, et autres
Publié: (2025)
RefusalGuard: Geometry-Preserving Fine-Tuning for Safety in LLMs
par: Asif, Sadia, et autres
Publié: (2026)
par: Asif, Sadia, et autres
Publié: (2026)
RDD: Retrieval-Based Demonstration Decomposer for Planner Alignment in Long-Horizon Tasks
par: Yan, Mingxuan, et autres
Publié: (2025)
par: Yan, Mingxuan, et autres
Publié: (2025)
Liquid Ensemble Selection for Continual Learning
par: Blair, Carter, et autres
Publié: (2024)
par: Blair, Carter, et autres
Publié: (2024)
Rethinking Inverse Reinforcement Learning: from Data Alignment to Task Alignment
par: Zhou, Weichao, et autres
Publié: (2024)
par: Zhou, Weichao, et autres
Publié: (2024)
Fair Dataset Distillation via Cross-Group Barycenter Alignment
par: Moslemi, Mohammad Hossein, et autres
Publié: (2026)
par: Moslemi, Mohammad Hossein, et autres
Publié: (2026)
Towards Generalisable Imitation Learning Through Conditioned Transition Estimation and Online Behaviour Alignment
par: Gavenski, Nathan, et autres
Publié: (2026)
par: Gavenski, Nathan, et autres
Publié: (2026)
Liquid Democracy for Low-Cost Ensemble Pruning
par: Armstrong, Ben, et autres
Publié: (2024)
par: Armstrong, Ben, et autres
Publié: (2024)
Model Alignment Search
par: Grant, Satchel
Publié: (2025)
par: Grant, Satchel
Publié: (2025)
ECLIPTICA -- A Framework for Switchable LLM Alignment via CITA - Contrastive Instruction-Tuned Alignment
par: Wanaskar, Kapil, et autres
Publié: (2026)
par: Wanaskar, Kapil, et autres
Publié: (2026)
Any-Depth Alignment: Unlocking Innate Safety Alignment of LLMs to Any-Depth
par: Zhang, Jiawei, et autres
Publié: (2025)
par: Zhang, Jiawei, et autres
Publié: (2025)
Conformal Feedback Alignment: Quantifying Answer-Level Reliability for Robust LLM Alignment
par: Chen, Tiejin, et autres
Publié: (2026)
par: Chen, Tiejin, et autres
Publié: (2026)
Automated Meta Prompt Engineering for Alignment with the Theory of Mind
par: Baughman, Aaron, et autres
Publié: (2025)
par: Baughman, Aaron, et autres
Publié: (2025)
Space Alignment Matters: The Missing Piece for Inducing Neural Collapse in Long-Tailed Learning
par: Wang, Jinping, et autres
Publié: (2025)
par: Wang, Jinping, et autres
Publié: (2025)
Four Things People Should Know About Migraines
par: Parsa, Mohammad S., et autres
Publié: (2025)
par: Parsa, Mohammad S., et autres
Publié: (2025)
Federated Unsupervised Domain Generalization using Global and Local Alignment of Gradients
par: Pourpanah, Farhad, et autres
Publié: (2024)
par: Pourpanah, Farhad, et autres
Publié: (2024)
Documents similaires
-
Curated Synthetic Data Doesn't Have to Collapse: A Theoretical Study of Generative Retraining with Pluralistic Preferences
par: Falahati, Ali, et autres
Publié: (2026) -
Disentangled Structural and Featural Representation for Task-Agnostic Graph Valuation
par: Falahati, Ali, et autres
Publié: (2024) -
A Median Perspective on Unlabeled Data for Out-of-Distribution Detection
par: Abbas, Momin, et autres
Publié: (2025) -
SentenceKV: Efficient LLM Inference via Sentence-Level Semantic KV Caching
par: Zhu, Yuxuan, et autres
Publié: (2025) -
Toward Efficient Influence Function: Dropout as a Compression Tool
par: Zhang, Yuchen, et autres
Publié: (2025)