Stackelberg Self-Annotation: A Robust Approach to Data-Efficient LLM Alignment
Fuente:
arXiv
Salvato in:
| Autori principali: | Chu, Xu, Zhang, Zhixin, Jia, Tianyu, Jin, Yujie |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Data Selection for LLM Alignment Using Fine-Grained Preferences
di: Zhang, Jia, et al.
Pubblicazione: (2025)
di: Zhang, Jia, et al.
Pubblicazione: (2025)
Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment
di: Kim, Dongyoung, et al.
Pubblicazione: (2024)
di: Kim, Dongyoung, et al.
Pubblicazione: (2024)
Reward Shaping for Inference-Time Alignment: A Stackelberg Game Perspective
di: Wang, Haichuan, et al.
Pubblicazione: (2026)
di: Wang, Haichuan, et al.
Pubblicazione: (2026)
Convergence Acceleration in Wireless Federated Learning: A Stackelberg Game Approach
di: Wang, Kaidi, et al.
Pubblicazione: (2022)
di: Wang, Kaidi, et al.
Pubblicazione: (2022)
PluralLLM: Pluralistic Alignment in LLMs via Federated Learning
di: Srewa, Mahmoud, et al.
Pubblicazione: (2025)
di: Srewa, Mahmoud, et al.
Pubblicazione: (2025)
Robust LLM Alignment via Distributionally Robust Direct Preference Optimization
di: Xu, Zaiyan, et al.
Pubblicazione: (2025)
di: Xu, Zaiyan, et al.
Pubblicazione: (2025)
DrugSAGE:Self-evolving Agent Experience for Efficient State-of-the-Art Drug Discovery
di: Zhang, Yikun, et al.
Pubblicazione: (2026)
di: Zhang, Yikun, et al.
Pubblicazione: (2026)
Conformal Thresholded Intervals for Efficient Regression
di: Luo, Rui, et al.
Pubblicazione: (2024)
di: Luo, Rui, et al.
Pubblicazione: (2024)
Adversarial Preference Learning for Robust LLM Alignment
di: Wang, Yuanfu, et al.
Pubblicazione: (2025)
di: Wang, Yuanfu, et al.
Pubblicazione: (2025)
SMART: Towards Pre-trained Missing-Aware Model for Patient Health Status Prediction
di: Yu, Zhihao, et al.
Pubblicazione: (2024)
di: Yu, Zhihao, et al.
Pubblicazione: (2024)
Robust Reward Alignment via Hypothesis Space Batch Cutting
di: Xie, Zhixian, et al.
Pubblicazione: (2025)
di: Xie, Zhixian, et al.
Pubblicazione: (2025)
Unsupervised Conformal Inference: Bootstrapping and Alignment to Control LLM Uncertainty
di: Pang, Lingyou, et al.
Pubblicazione: (2025)
di: Pang, Lingyou, et al.
Pubblicazione: (2025)
Data-Centric Human Preference with Rationales for Direct Preference Alignment
di: Just, Hoang Anh, et al.
Pubblicazione: (2024)
di: Just, Hoang Anh, et al.
Pubblicazione: (2024)
D3: Diversity, Difficulty, and Dependability-Aware Data Selection for Sample-Efficient LLM Instruction Tuning
di: Zhang, Jia, et al.
Pubblicazione: (2025)
di: Zhang, Jia, et al.
Pubblicazione: (2025)
IntelliCare: Improving Healthcare Analysis with Variance-Controlled Patient-Level Knowledge from Large Language Models
di: Yu, Zhihao, et al.
Pubblicazione: (2024)
di: Yu, Zhihao, et al.
Pubblicazione: (2024)
From Parameters to Data: A Task-Parameter-Guided Fine-Tuning Pipeline for Efficient LLM Alignment
di: Chen, Hao, et al.
Pubblicazione: (2026)
di: Chen, Hao, et al.
Pubblicazione: (2026)
Volume-Sorted Prediction Set: Efficient Conformal Prediction for Multi-Target Regression
di: Luo, Rui, et al.
Pubblicazione: (2025)
di: Luo, Rui, et al.
Pubblicazione: (2025)
SCAN: Self-Denoising Monte Carlo Annotation for Robust Process Reward Learning
di: Ding, Yuyang, et al.
Pubblicazione: (2025)
di: Ding, Yuyang, et al.
Pubblicazione: (2025)
ALMA: Alignment with Minimal Annotation
di: Yasunaga, Michihiro, et al.
Pubblicazione: (2024)
di: Yasunaga, Michihiro, et al.
Pubblicazione: (2024)
Conformal Feedback Alignment: Quantifying Answer-Level Reliability for Robust LLM Alignment
di: Chen, Tiejin, et al.
Pubblicazione: (2026)
di: Chen, Tiejin, et al.
Pubblicazione: (2026)
Adaptive Ensembles of Fine-Tuned Transformers for LLM-Generated Text Detection
di: Lai, Zhixin, et al.
Pubblicazione: (2024)
di: Lai, Zhixin, et al.
Pubblicazione: (2024)
Online Distributionally Robust LLM Alignment via Regression to Relative Reward
di: Sahu, Sharan, et al.
Pubblicazione: (2025)
di: Sahu, Sharan, et al.
Pubblicazione: (2025)
Revisiting Robustness for LLM Safety Alignment via Selective Geometry Control
di: Yang, Yonghui, et al.
Pubblicazione: (2026)
di: Yang, Yonghui, et al.
Pubblicazione: (2026)
Multimodal-Guided Dynamic Dataset Pruning for Robust and Efficient Data-Centric Learning
di: Yang, Suorong, et al.
Pubblicazione: (2025)
di: Yang, Suorong, et al.
Pubblicazione: (2025)
Rubric-Conditioned LLM Grading: Alignment, Uncertainty, and Robustness
di: Deng, Haotian, et al.
Pubblicazione: (2025)
di: Deng, Haotian, et al.
Pubblicazione: (2025)
Contextual Optimization under Covariate Shift: A Robust Approach by Intersecting Wasserstein Balls
di: Wang, Tianyu, et al.
Pubblicazione: (2024)
di: Wang, Tianyu, et al.
Pubblicazione: (2024)
Reducing the Safety Tax in LLM Safety Alignment with On-Policy Self-Distillation
di: Fu, Yu, et al.
Pubblicazione: (2026)
di: Fu, Yu, et al.
Pubblicazione: (2026)
Prior-Informed Zeroth-Order Optimization with Adaptive Direction Alignment for Memory-Efficient LLM Fine-Tuning
di: Jin, Feihu, et al.
Pubblicazione: (2026)
di: Jin, Feihu, et al.
Pubblicazione: (2026)
LLM Chain Ensembles for Scalable and Accurate Data Annotation
di: Farr, David, et al.
Pubblicazione: (2024)
di: Farr, David, et al.
Pubblicazione: (2024)
Rethinking Neural Network Learning Rates: A Stackelberg Perspective
di: Zeng, Sihan, et al.
Pubblicazione: (2026)
di: Zeng, Sihan, et al.
Pubblicazione: (2026)
TriPlay-RL: Tri-Role Self-Play Reinforcement Learning for LLM Safety Alignment
di: Tan, Zhewen, et al.
Pubblicazione: (2026)
di: Tan, Zhewen, et al.
Pubblicazione: (2026)
Meta Stackelberg Game: Robust Federated Learning against Adaptive and Mixed Poisoning Attacks
di: Li, Tao, et al.
Pubblicazione: (2024)
di: Li, Tao, et al.
Pubblicazione: (2024)
CluCERT: Certifying LLM Robustness via Clustering-Guided Denoising Smoothing
di: Wang, Zixia, et al.
Pubblicazione: (2025)
di: Wang, Zixia, et al.
Pubblicazione: (2025)
AnnotatedTables: A Large Tabular Dataset with Language Model Annotations
di: Hu, Yaojie, et al.
Pubblicazione: (2024)
di: Hu, Yaojie, et al.
Pubblicazione: (2024)
Weakly Supervised Anomaly Detection via Knowledge-Data Alignment
di: Zhao, Haihong, et al.
Pubblicazione: (2024)
di: Zhao, Haihong, et al.
Pubblicazione: (2024)
Gradient Compression May Hurt Generalization: A Remedy by Synthetic Data Guided Sharpness Aware Minimization
di: Gu, Yujie, et al.
Pubblicazione: (2026)
di: Gu, Yujie, et al.
Pubblicazione: (2026)
RobustExplain: Evaluating Robustness of LLM-Based Explanation Agents for Recommendation
di: Zhang, Guilin, et al.
Pubblicazione: (2026)
di: Zhang, Guilin, et al.
Pubblicazione: (2026)
Advantage Alignment Algorithms
di: Duque, Juan Agustin, et al.
Pubblicazione: (2024)
di: Duque, Juan Agustin, et al.
Pubblicazione: (2024)
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment
di: Trivedi, Prashant, et al.
Pubblicazione: (2025)
di: Trivedi, Prashant, et al.
Pubblicazione: (2025)
Efficient Text-Attributed Graph Learning through Selective Annotation and Graph Alignment
di: Xie, Huanyi, et al.
Pubblicazione: (2025)
di: Xie, Huanyi, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Data Selection for LLM Alignment Using Fine-Grained Preferences
di: Zhang, Jia, et al.
Pubblicazione: (2025) -
Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment
di: Kim, Dongyoung, et al.
Pubblicazione: (2024) -
Reward Shaping for Inference-Time Alignment: A Stackelberg Game Perspective
di: Wang, Haichuan, et al.
Pubblicazione: (2026) -
Convergence Acceleration in Wireless Federated Learning: A Stackelberg Game Approach
di: Wang, Kaidi, et al.
Pubblicazione: (2022) -
PluralLLM: Pluralistic Alignment in LLMs via Federated Learning
di: Srewa, Mahmoud, et al.
Pubblicazione: (2025)