InfoPO: On Mutual Information Maximization for Large Language Model Alignment
Fuente:
arXiv
Guardado en:
| Autores principales: | Xiao, Teng, Ge, Zhen, Sanghavi, Sujay, Wang, Tian, Katz-Samuels, Julian, Versage, Marc, Cui, Qingjun, Chilimbi, Trishul |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
InfoPO: Information-Driven Policy Optimization for User-Centric Agents
por: Kong, Fanqi, et al.
Publicado: (2026)
por: Kong, Fanqi, et al.
Publicado: (2026)
Evolutionary Contrastive Distillation for Language Model Alignment
por: Katz-Samuels, Julian, et al.
Publicado: (2024)
por: Katz-Samuels, Julian, et al.
Publicado: (2024)
Exploring Reasoning-Infused Text Embedding with Large Language Models for Zero-Shot Dense Retrieval
por: Liu, Yuxiang, et al.
Publicado: (2025)
por: Liu, Yuxiang, et al.
Publicado: (2025)
SynerGen: Contextualized Generative Recommender for Unified Search and Recommendation
por: Gao, Vianne R., et al.
Publicado: (2025)
por: Gao, Vianne R., et al.
Publicado: (2025)
AutoMixAlign: Adaptive Data Mixing for Multi-Task Preference Optimization in LLMs
por: Corrado, Nicholas E., et al.
Publicado: (2025)
por: Corrado, Nicholas E., et al.
Publicado: (2025)
VidLA: Video-Language Alignment at Scale
por: Rizve, Mamshad Nayeem, et al.
Publicado: (2024)
por: Rizve, Mamshad Nayeem, et al.
Publicado: (2024)
Entropy Aware Reward Guidance for Diffusion Language Model Alignment
por: Tejaswi, Atula, et al.
Publicado: (2026)
por: Tejaswi, Atula, et al.
Publicado: (2026)
Context-Free Synthetic Data Mitigates Forgetting
por: Bansal, Parikshit, et al.
Publicado: (2025)
por: Bansal, Parikshit, et al.
Publicado: (2025)
Enabling Approximate Joint Sampling in Diffusion LMs
por: Bansal, Parikshit, et al.
Publicado: (2025)
por: Bansal, Parikshit, et al.
Publicado: (2025)
CoLLM: A Large Language Model for Composed Image Retrieval
por: Huynh, Chuong, et al.
Publicado: (2025)
por: Huynh, Chuong, et al.
Publicado: (2025)
InfoCTM: A Mutual Information Maximization Perspective of Cross-Lingual Topic Modeling
por: Wu, Xiaobao, et al.
Publicado: (2023)
por: Wu, Xiaobao, et al.
Publicado: (2023)
DreamBlend: Advancing Personalized Fine-tuning of Text-to-Image Diffusion Models
por: Ram, Shwetha, et al.
Publicado: (2024)
por: Ram, Shwetha, et al.
Publicado: (2024)
Learning Mixtures of Experts with EM: A Mirror Descent Perspective
por: Fruytier, Quentin, et al.
Publicado: (2024)
por: Fruytier, Quentin, et al.
Publicado: (2024)
Geometric Median (GM) Matching for Robust Data Pruning
por: Acharya, Anish, et al.
Publicado: (2024)
por: Acharya, Anish, et al.
Publicado: (2024)
Understanding Self-Supervised Learning via Gaussian Mixture Models
por: Bansal, Parikshit, et al.
Publicado: (2024)
por: Bansal, Parikshit, et al.
Publicado: (2024)
HiSpec: Hierarchical Speculative Decoding for LLMs
por: Kumar, Avinash, et al.
Publicado: (2025)
por: Kumar, Avinash, et al.
Publicado: (2025)
Test-Time Speculation
por: Kumar, Avinash, et al.
Publicado: (2026)
por: Kumar, Avinash, et al.
Publicado: (2026)
InfoRank: Unbiased Learning-to-Rank via Conditional Mutual Information Minimization
por: Jin, Jiarui, et al.
Publicado: (2024)
por: Jin, Jiarui, et al.
Publicado: (2024)
InfoBridge: Mutual Information estimation via Bridge Matching
por: Kholkin, Sergei, et al.
Publicado: (2025)
por: Kholkin, Sergei, et al.
Publicado: (2025)
InfoNorm: Mutual Information Shaping of Normals for Sparse-View Reconstruction
por: Wang, Xulong, et al.
Publicado: (2024)
por: Wang, Xulong, et al.
Publicado: (2024)
Towards Quantifying the Preconditioning Effect of Adam
por: Das, Rudrajit, et al.
Publicado: (2024)
por: Das, Rudrajit, et al.
Publicado: (2024)
RARe: Retrieval Augmented Retrieval with In-Context Examples
por: Tejaswi, Atula, et al.
Publicado: (2024)
por: Tejaswi, Atula, et al.
Publicado: (2024)
Blocking Bandits
por: Basu, Soumya, et al.
Publicado: (2019)
por: Basu, Soumya, et al.
Publicado: (2019)
InfoNet: Neural Estimation of Mutual Information without Test-Time Optimization
por: Hu, Zhengyang, et al.
Publicado: (2024)
por: Hu, Zhengyang, et al.
Publicado: (2024)
Robust Multi-Task Learning with Excess Risks
por: He, Yifei, et al.
Publicado: (2024)
por: He, Yifei, et al.
Publicado: (2024)
Geometric Median Matching for Robust k-Subset Selection from Noisy Data
por: Acharya, Anish, et al.
Publicado: (2025)
por: Acharya, Anish, et al.
Publicado: (2025)
HYPO: Hyperspherical Out-of-Distribution Generalization
por: Bai, Haoyue, et al.
Publicado: (2024)
por: Bai, Haoyue, et al.
Publicado: (2024)
InfoMerge: Information-aware Token Compression for Efficient Video Large Language Models
por: Liu, Xinxin, et al.
Publicado: (2026)
por: Liu, Xinxin, et al.
Publicado: (2026)
InfoFlood: Jailbreaking Large Language Models with Information Overload
por: Yadav, Advait, et al.
Publicado: (2025)
por: Yadav, Advait, et al.
Publicado: (2025)
VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization
por: Li, Mingxiao, et al.
Publicado: (2025)
por: Li, Mingxiao, et al.
Publicado: (2025)
Open Vocabulary Multi-Label Video Classification
por: Gupta, Rohit, et al.
Publicado: (2024)
por: Gupta, Rohit, et al.
Publicado: (2024)
X-Former: Unifying Contrastive and Reconstruction Learning for MLLMs
por: Swetha, Sirnam, et al.
Publicado: (2024)
por: Swetha, Sirnam, et al.
Publicado: (2024)
Parameter Estimation of Mutual Information Maximized Channels
por: Tavakoli, Hassan, et al.
Publicado: (2026)
por: Tavakoli, Hassan, et al.
Publicado: (2026)
Upweighting Easy Samples in Fine-Tuning Mitigates Forgetting
por: Sanyal, Sunny, et al.
Publicado: (2025)
por: Sanyal, Sunny, et al.
Publicado: (2025)
When Attention Collapses: How Degenerate Layers in LLMs Enable Smaller, Stronger Models
por: Sanyal, Sunny, et al.
Publicado: (2024)
por: Sanyal, Sunny, et al.
Publicado: (2024)
In-Context Learning with Transformers: Softmax Attention Adapts to Function Lipschitzness
por: Collins, Liam, et al.
Publicado: (2024)
por: Collins, Liam, et al.
Publicado: (2024)
Asymptotically-Optimal Gaussian Bandits with Side Observations
por: Atsidakou, Alexia, et al.
Publicado: (2025)
por: Atsidakou, Alexia, et al.
Publicado: (2025)
Finite-Time Logarithmic Bayes Regret Upper Bounds
por: Atsidakou, Alexia, et al.
Publicado: (2023)
por: Atsidakou, Alexia, et al.
Publicado: (2023)
Understanding the Training Speedup from Sampling with Approximate Losses
por: Das, Rudrajit, et al.
Publicado: (2024)
por: Das, Rudrajit, et al.
Publicado: (2024)
Adaptive and Optimal Second-order Optimistic Methods for Minimax Optimization
por: Jiang, Ruichen, et al.
Publicado: (2024)
por: Jiang, Ruichen, et al.
Publicado: (2024)
Ejemplares similares
-
InfoPO: Information-Driven Policy Optimization for User-Centric Agents
por: Kong, Fanqi, et al.
Publicado: (2026) -
Evolutionary Contrastive Distillation for Language Model Alignment
por: Katz-Samuels, Julian, et al.
Publicado: (2024) -
Exploring Reasoning-Infused Text Embedding with Large Language Models for Zero-Shot Dense Retrieval
por: Liu, Yuxiang, et al.
Publicado: (2025) -
SynerGen: Contextualized Generative Recommender for Unified Search and Recommendation
por: Gao, Vianne R., et al.
Publicado: (2025) -
AutoMixAlign: Adaptive Data Mixing for Multi-Task Preference Optimization in LLMs
por: Corrado, Nicholas E., et al.
Publicado: (2025)