InfAlign: Inference-aware language model alignment
Fuente:
arXiv
Salvato in:
| Autori principali: | Balashankar, Ananth, Sun, Ziteng, Berant, Jonathan, Eisenstein, Jacob, Collins, Michael, Hutter, Adrian, Lee, Jong, Nagpal, Chirag, Prost, Flavien, Sinha, Aradhana, Suresh, Ananda Theertha, Beirami, Ahmad |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Theoretical guarantees on the best-of-n alignment policy
di: Beirami, Ahmad, et al.
Pubblicazione: (2024)
di: Beirami, Ahmad, et al.
Pubblicazione: (2024)
Inducing Group Fairness in Prompt-Based Language Model Decisions
di: Atwood, James, et al.
Pubblicazione: (2024)
di: Atwood, James, et al.
Pubblicazione: (2024)
Break it, Imitate it, Fix it: Robustness by Generating Human-Like Attacks
di: Sinha, Aradhana, et al.
Pubblicazione: (2023)
di: Sinha, Aradhana, et al.
Pubblicazione: (2023)
Reuse Your Rewards: Reward Model Transfer for Zero-Shot Cross-Lingual Alignment
di: Wu, Zhaofeng, et al.
Pubblicazione: (2024)
di: Wu, Zhaofeng, et al.
Pubblicazione: (2024)
Asymptotics of Language Model Alignment
di: Yang, Joy Qiping, et al.
Pubblicazione: (2024)
di: Yang, Joy Qiping, et al.
Pubblicazione: (2024)
Robust Preference Optimization through Reward Model Distillation
di: Fisch, Adam, et al.
Pubblicazione: (2024)
di: Fisch, Adam, et al.
Pubblicazione: (2024)
Transforming and Combining Rewards for Aligning Large Language Models
di: Wang, Zihao, et al.
Pubblicazione: (2024)
di: Wang, Zihao, et al.
Pubblicazione: (2024)
SpecTr: Fast Speculative Decoding via Optimal Transport
di: Sun, Ziteng, et al.
Pubblicazione: (2023)
di: Sun, Ziteng, et al.
Pubblicazione: (2023)
The importance of feature preprocessing for differentially private linear optimization
di: Sun, Ziteng, et al.
Pubblicazione: (2023)
di: Sun, Ziteng, et al.
Pubblicazione: (2023)
Automated Adversarial Discovery for Safety Classifiers
di: Lal, Yash Kumar, et al.
Pubblicazione: (2024)
di: Lal, Yash Kumar, et al.
Pubblicazione: (2024)
Subset-Based Instance Optimality in Private Estimation
di: Dick, Travis, et al.
Pubblicazione: (2023)
di: Dick, Travis, et al.
Pubblicazione: (2023)
Block Verification Accelerates Speculative Decoding
di: Sun, Ziteng, et al.
Pubblicazione: (2024)
di: Sun, Ziteng, et al.
Pubblicazione: (2024)
Multi-Mixer Models: Flexible Sequence Modeling with Shared Representations
di: Li, Kevin Y., et al.
Pubblicazione: (2026)
di: Li, Kevin Y., et al.
Pubblicazione: (2026)
Private federated discovery of out-of-vocabulary words for Gboard
di: Sun, Ziteng, et al.
Pubblicazione: (2024)
di: Sun, Ziteng, et al.
Pubblicazione: (2024)
Multi-Group Fairness Evaluation via Conditional Value-at-Risk Testing
di: Paes, Lucas Monteiro, et al.
Pubblicazione: (2023)
di: Paes, Lucas Monteiro, et al.
Pubblicazione: (2023)
CoDistill-GRPO: A Co-Distillation Recipe for Efficient Group Relative Policy Optimization
di: Kwon, Soo Min, et al.
Pubblicazione: (2026)
di: Kwon, Soo Min, et al.
Pubblicazione: (2026)
Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
di: Eisenstein, Jacob, et al.
Pubblicazione: (2023)
di: Eisenstein, Jacob, et al.
Pubblicazione: (2023)
Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning
di: Setlur, Amrith, et al.
Pubblicazione: (2024)
di: Setlur, Amrith, et al.
Pubblicazione: (2024)
Coupling without Communication and Drafter-Invariant Speculative Decoding
di: Daliri, Majid, et al.
Pubblicazione: (2024)
di: Daliri, Majid, et al.
Pubblicazione: (2024)
On Robust Hypothesis Testing with respect to the Hellinger Distance
di: Modak, Eeshan, et al.
Pubblicazione: (2025)
di: Modak, Eeshan, et al.
Pubblicazione: (2025)
Mean estimation in the add-remove model of differential privacy
di: Kulesza, Alex, et al.
Pubblicazione: (2023)
di: Kulesza, Alex, et al.
Pubblicazione: (2023)
FRAPPE: A Group Fairness Framework for Post-Processing Everything
di: Tifrea, Alexandru, et al.
Pubblicazione: (2023)
di: Tifrea, Alexandru, et al.
Pubblicazione: (2023)
Preference Models assume Proportional Hazards of Utilities
di: Nagpal, Chirag
Pubblicazione: (2025)
di: Nagpal, Chirag
Pubblicazione: (2025)
CafeQ: Calibration-free Quantization via Learned Transformations and Adaptive Rounding
di: Sun, Ziteng, et al.
Pubblicazione: (2025)
di: Sun, Ziteng, et al.
Pubblicazione: (2025)
Rate of Model Collapse in Recursive Training
di: Suresh, Ananda Theertha, et al.
Pubblicazione: (2024)
di: Suresh, Ananda Theertha, et al.
Pubblicazione: (2024)
BEFT: Bias-Efficient Fine-Tuning of Language Models in Low-Data Regimes
di: Huang, Baichuan, et al.
Pubblicazione: (2025)
di: Huang, Baichuan, et al.
Pubblicazione: (2025)
FieldFormer: Locality-Aware Transformers for Spatio-Temporal Modeling on Sparse Sensor Networks
di: Bhardwaj, Ankit, et al.
Pubblicazione: (2025)
di: Bhardwaj, Ankit, et al.
Pubblicazione: (2025)
Systematic Scaling Analysis of Jailbreak Attacks in Large Language Models
di: Wang, Xiangwen, et al.
Pubblicazione: (2026)
di: Wang, Xiangwen, et al.
Pubblicazione: (2026)
MT-PingEval: Evaluating Multi-Turn Collaboration with Private Information Games
di: Eisenstein, Jacob, et al.
Pubblicazione: (2026)
di: Eisenstein, Jacob, et al.
Pubblicazione: (2026)
Efficient Language Model Architectures for Differentially Private Federated Learning
di: Ro, Jae Hun, et al.
Pubblicazione: (2024)
di: Ro, Jae Hun, et al.
Pubblicazione: (2024)
Cost-Optimal Active AI Model Evaluation
di: Angelopoulos, Anastasios N., et al.
Pubblicazione: (2025)
di: Angelopoulos, Anastasios N., et al.
Pubblicazione: (2025)
Effect of Pre‐ and Post‐Milling Processing Techniques on the Physico‐Chemical, Functional, and Pasting Properties of Sorghum
di: Theertha DP, et al.
Pubblicazione: (2025)
di: Theertha DP, et al.
Pubblicazione: (2025)
Exploring and Improving Drafts in Blockwise Parallel Decoding
di: Kim, Taehyeon, et al.
Pubblicazione: (2024)
di: Kim, Taehyeon, et al.
Pubblicazione: (2024)
Hierarchical Retrieval: The Geometry and a Pretrain-Finetune Recipe
di: You, Chong, et al.
Pubblicazione: (2025)
di: You, Chong, et al.
Pubblicazione: (2025)
Plantain: Plan-Answer Interleaved Reasoning
di: Liang, Anthony, et al.
Pubblicazione: (2025)
di: Liang, Anthony, et al.
Pubblicazione: (2025)
ALTA: Compiler-Based Analysis of Transformers
di: Shaw, Peter, et al.
Pubblicazione: (2024)
di: Shaw, Peter, et al.
Pubblicazione: (2024)
Breaking the Illusion: Consensus-Based Generative Mitigation of Adversarial Illusions in Multi-Modal Embeddings
di: Akbarian, Fatemeh, et al.
Pubblicazione: (2025)
di: Akbarian, Fatemeh, et al.
Pubblicazione: (2025)
Robust LLM Performance Certification via Constrained Maximum Likelihood Estimation
di: Shen, Minghe, et al.
Pubblicazione: (2026)
di: Shen, Minghe, et al.
Pubblicazione: (2026)
Exploring Social Business Pathways: Green Map System as a Case in Point
di: Balashankar Mulloth
Pubblicazione: (2021)
di: Balashankar Mulloth
Pubblicazione: (2021)
Efficient and Asymptotically Unbiased Constrained Decoding for Large Language Models
di: Ye, Haotian, et al.
Pubblicazione: (2025)
di: Ye, Haotian, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Theoretical guarantees on the best-of-n alignment policy
di: Beirami, Ahmad, et al.
Pubblicazione: (2024) -
Inducing Group Fairness in Prompt-Based Language Model Decisions
di: Atwood, James, et al.
Pubblicazione: (2024) -
Break it, Imitate it, Fix it: Robustness by Generating Human-Like Attacks
di: Sinha, Aradhana, et al.
Pubblicazione: (2023) -
Reuse Your Rewards: Reward Model Transfer for Zero-Shot Cross-Lingual Alignment
di: Wu, Zhaofeng, et al.
Pubblicazione: (2024) -
Asymptotics of Language Model Alignment
di: Yang, Joy Qiping, et al.
Pubblicazione: (2024)