Stackelberg Self-Annotation: A Robust Approach to Data-Efficient LLM Alignment
Fuente:
arXiv
Guardado en:
| Autores principales: | Chu, Xu, Zhang, Zhixin, Jia, Tianyu, Jin, Yujie |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Data Selection for LLM Alignment Using Fine-Grained Preferences
por: Zhang, Jia, et al.
Publicado: (2025)
por: Zhang, Jia, et al.
Publicado: (2025)
Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment
por: Kim, Dongyoung, et al.
Publicado: (2024)
por: Kim, Dongyoung, et al.
Publicado: (2024)
Reward Shaping for Inference-Time Alignment: A Stackelberg Game Perspective
por: Wang, Haichuan, et al.
Publicado: (2026)
por: Wang, Haichuan, et al.
Publicado: (2026)
Convergence Acceleration in Wireless Federated Learning: A Stackelberg Game Approach
por: Wang, Kaidi, et al.
Publicado: (2022)
por: Wang, Kaidi, et al.
Publicado: (2022)
PluralLLM: Pluralistic Alignment in LLMs via Federated Learning
por: Srewa, Mahmoud, et al.
Publicado: (2025)
por: Srewa, Mahmoud, et al.
Publicado: (2025)
Robust LLM Alignment via Distributionally Robust Direct Preference Optimization
por: Xu, Zaiyan, et al.
Publicado: (2025)
por: Xu, Zaiyan, et al.
Publicado: (2025)
DrugSAGE:Self-evolving Agent Experience for Efficient State-of-the-Art Drug Discovery
por: Zhang, Yikun, et al.
Publicado: (2026)
por: Zhang, Yikun, et al.
Publicado: (2026)
Conformal Thresholded Intervals for Efficient Regression
por: Luo, Rui, et al.
Publicado: (2024)
por: Luo, Rui, et al.
Publicado: (2024)
Adversarial Preference Learning for Robust LLM Alignment
por: Wang, Yuanfu, et al.
Publicado: (2025)
por: Wang, Yuanfu, et al.
Publicado: (2025)
SMART: Towards Pre-trained Missing-Aware Model for Patient Health Status Prediction
por: Yu, Zhihao, et al.
Publicado: (2024)
por: Yu, Zhihao, et al.
Publicado: (2024)
Robust Reward Alignment via Hypothesis Space Batch Cutting
por: Xie, Zhixian, et al.
Publicado: (2025)
por: Xie, Zhixian, et al.
Publicado: (2025)
Unsupervised Conformal Inference: Bootstrapping and Alignment to Control LLM Uncertainty
por: Pang, Lingyou, et al.
Publicado: (2025)
por: Pang, Lingyou, et al.
Publicado: (2025)
Data-Centric Human Preference with Rationales for Direct Preference Alignment
por: Just, Hoang Anh, et al.
Publicado: (2024)
por: Just, Hoang Anh, et al.
Publicado: (2024)
D3: Diversity, Difficulty, and Dependability-Aware Data Selection for Sample-Efficient LLM Instruction Tuning
por: Zhang, Jia, et al.
Publicado: (2025)
por: Zhang, Jia, et al.
Publicado: (2025)
IntelliCare: Improving Healthcare Analysis with Variance-Controlled Patient-Level Knowledge from Large Language Models
por: Yu, Zhihao, et al.
Publicado: (2024)
por: Yu, Zhihao, et al.
Publicado: (2024)
From Parameters to Data: A Task-Parameter-Guided Fine-Tuning Pipeline for Efficient LLM Alignment
por: Chen, Hao, et al.
Publicado: (2026)
por: Chen, Hao, et al.
Publicado: (2026)
Volume-Sorted Prediction Set: Efficient Conformal Prediction for Multi-Target Regression
por: Luo, Rui, et al.
Publicado: (2025)
por: Luo, Rui, et al.
Publicado: (2025)
SCAN: Self-Denoising Monte Carlo Annotation for Robust Process Reward Learning
por: Ding, Yuyang, et al.
Publicado: (2025)
por: Ding, Yuyang, et al.
Publicado: (2025)
ALMA: Alignment with Minimal Annotation
por: Yasunaga, Michihiro, et al.
Publicado: (2024)
por: Yasunaga, Michihiro, et al.
Publicado: (2024)
Conformal Feedback Alignment: Quantifying Answer-Level Reliability for Robust LLM Alignment
por: Chen, Tiejin, et al.
Publicado: (2026)
por: Chen, Tiejin, et al.
Publicado: (2026)
Adaptive Ensembles of Fine-Tuned Transformers for LLM-Generated Text Detection
por: Lai, Zhixin, et al.
Publicado: (2024)
por: Lai, Zhixin, et al.
Publicado: (2024)
Online Distributionally Robust LLM Alignment via Regression to Relative Reward
por: Sahu, Sharan, et al.
Publicado: (2025)
por: Sahu, Sharan, et al.
Publicado: (2025)
Revisiting Robustness for LLM Safety Alignment via Selective Geometry Control
por: Yang, Yonghui, et al.
Publicado: (2026)
por: Yang, Yonghui, et al.
Publicado: (2026)
Multimodal-Guided Dynamic Dataset Pruning for Robust and Efficient Data-Centric Learning
por: Yang, Suorong, et al.
Publicado: (2025)
por: Yang, Suorong, et al.
Publicado: (2025)
Rubric-Conditioned LLM Grading: Alignment, Uncertainty, and Robustness
por: Deng, Haotian, et al.
Publicado: (2025)
por: Deng, Haotian, et al.
Publicado: (2025)
Contextual Optimization under Covariate Shift: A Robust Approach by Intersecting Wasserstein Balls
por: Wang, Tianyu, et al.
Publicado: (2024)
por: Wang, Tianyu, et al.
Publicado: (2024)
Reducing the Safety Tax in LLM Safety Alignment with On-Policy Self-Distillation
por: Fu, Yu, et al.
Publicado: (2026)
por: Fu, Yu, et al.
Publicado: (2026)
Prior-Informed Zeroth-Order Optimization with Adaptive Direction Alignment for Memory-Efficient LLM Fine-Tuning
por: Jin, Feihu, et al.
Publicado: (2026)
por: Jin, Feihu, et al.
Publicado: (2026)
LLM Chain Ensembles for Scalable and Accurate Data Annotation
por: Farr, David, et al.
Publicado: (2024)
por: Farr, David, et al.
Publicado: (2024)
Rethinking Neural Network Learning Rates: A Stackelberg Perspective
por: Zeng, Sihan, et al.
Publicado: (2026)
por: Zeng, Sihan, et al.
Publicado: (2026)
TriPlay-RL: Tri-Role Self-Play Reinforcement Learning for LLM Safety Alignment
por: Tan, Zhewen, et al.
Publicado: (2026)
por: Tan, Zhewen, et al.
Publicado: (2026)
Meta Stackelberg Game: Robust Federated Learning against Adaptive and Mixed Poisoning Attacks
por: Li, Tao, et al.
Publicado: (2024)
por: Li, Tao, et al.
Publicado: (2024)
CluCERT: Certifying LLM Robustness via Clustering-Guided Denoising Smoothing
por: Wang, Zixia, et al.
Publicado: (2025)
por: Wang, Zixia, et al.
Publicado: (2025)
AnnotatedTables: A Large Tabular Dataset with Language Model Annotations
por: Hu, Yaojie, et al.
Publicado: (2024)
por: Hu, Yaojie, et al.
Publicado: (2024)
Weakly Supervised Anomaly Detection via Knowledge-Data Alignment
por: Zhao, Haihong, et al.
Publicado: (2024)
por: Zhao, Haihong, et al.
Publicado: (2024)
Gradient Compression May Hurt Generalization: A Remedy by Synthetic Data Guided Sharpness Aware Minimization
por: Gu, Yujie, et al.
Publicado: (2026)
por: Gu, Yujie, et al.
Publicado: (2026)
RobustExplain: Evaluating Robustness of LLM-Based Explanation Agents for Recommendation
por: Zhang, Guilin, et al.
Publicado: (2026)
por: Zhang, Guilin, et al.
Publicado: (2026)
Advantage Alignment Algorithms
por: Duque, Juan Agustin, et al.
Publicado: (2024)
por: Duque, Juan Agustin, et al.
Publicado: (2024)
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment
por: Trivedi, Prashant, et al.
Publicado: (2025)
por: Trivedi, Prashant, et al.
Publicado: (2025)
Efficient Text-Attributed Graph Learning through Selective Annotation and Graph Alignment
por: Xie, Huanyi, et al.
Publicado: (2025)
por: Xie, Huanyi, et al.
Publicado: (2025)
Ejemplares similares
-
Data Selection for LLM Alignment Using Fine-Grained Preferences
por: Zhang, Jia, et al.
Publicado: (2025) -
Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment
por: Kim, Dongyoung, et al.
Publicado: (2024) -
Reward Shaping for Inference-Time Alignment: A Stackelberg Game Perspective
por: Wang, Haichuan, et al.
Publicado: (2026) -
Convergence Acceleration in Wireless Federated Learning: A Stackelberg Game Approach
por: Wang, Kaidi, et al.
Publicado: (2022) -
PluralLLM: Pluralistic Alignment in LLMs via Federated Learning
por: Srewa, Mahmoud, et al.
Publicado: (2025)