QUATRO: Query-Adaptive Trust Region Policy Optimization for LLM Fine-tuning
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Doyeon, Lyou, Eunyi, Cho, Hyunsoo, Kim, Sookyung, Lee, Joonseok, Choi, Jaemoo |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Modality-Aware Representation Learning for Zero-shot Sketch-based Image Retrieval
by: Lyou, Eunyi, et al.
Published: (2024)
by: Lyou, Eunyi, et al.
Published: (2024)
Efficient Adjoint Matching for Fine-tuning Diffusion Models
by: Shin, Jeongwoo, et al.
Published: (2026)
by: Shin, Jeongwoo, et al.
Published: (2026)
A More Word-like Image Tokenization for MLLMs
by: Lee, Hyun, et al.
Published: (2026)
by: Lee, Hyun, et al.
Published: (2026)
Trust-Region Adaptive Policy Optimization
by: Su, Mingyu, et al.
Published: (2025)
by: Su, Mingyu, et al.
Published: (2025)
Coarse-to-Fine Compositional Diffusion for Long-Horizon Planning
by: Park, Byoungwoo, et al.
Published: (2026)
by: Park, Byoungwoo, et al.
Published: (2026)
Scalable Simulation-free Entropic Unbalanced Optimal Transport
by: Choi, Jaemoo, et al.
Published: (2024)
by: Choi, Jaemoo, et al.
Published: (2024)
Trust the uncertain teacher: distilling dark knowledge via calibrated uncertainty
by: Kim, Jeonghyun, et al.
Published: (2026)
by: Kim, Jeonghyun, et al.
Published: (2026)
Memento: Fine-tuning LLM Agents without Fine-tuning LLMs
by: Zhou, Huichi, et al.
Published: (2025)
by: Zhou, Huichi, et al.
Published: (2025)
Adaptive Divergence Regularized Policy Optimization for Fine-tuning Generative Models
by: Fan, Jiajun, et al.
Published: (2025)
by: Fan, Jiajun, et al.
Published: (2025)
Ratio-Variance Regularized Policy Optimization for Efficient LLM Fine-tuning
by: Luo, Yu, et al.
Published: (2026)
by: Luo, Yu, et al.
Published: (2026)
Overcoming Spurious Solutions in Semi-Dual Neural Optimal Transport: A Smoothing Approach for Learning the Optimal Transport Plan
by: Choi, Jaemoo, et al.
Published: (2025)
by: Choi, Jaemoo, et al.
Published: (2025)
MAGE: All-[MASK] Block Already Knows Where to Look in Diffusion LLM
by: Kwon, Omin, et al.
Published: (2026)
by: Kwon, Omin, et al.
Published: (2026)
Enhancing Reasoning for Diffusion LLMs via Distribution Matching Policy Optimization
by: Zhu, Yuchen, et al.
Published: (2025)
by: Zhu, Yuchen, et al.
Published: (2025)
Rethinking Safety in LLM Fine-tuning: An Optimization Perspective
by: Kim, Minseon, et al.
Published: (2025)
by: Kim, Minseon, et al.
Published: (2025)
Neural ODE Transformers: Analyzing Internal Dynamics and Adaptive Fine-tuning
by: Tong, Anh, et al.
Published: (2025)
by: Tong, Anh, et al.
Published: (2025)
Isometric Representation Learning for Disentangled Latent Space of Diffusion Models
by: Hahm, Jaehoon, et al.
Published: (2024)
by: Hahm, Jaehoon, et al.
Published: (2024)
ERPPO: Entropy Regularization-based Proximal Policy Optimization
by: Lee, Changha, et al.
Published: (2026)
by: Lee, Changha, et al.
Published: (2026)
DiSPA: Differential Substructure-Pathway Attention for Drug Response Prediction
by: Han, Yewon, et al.
Published: (2026)
by: Han, Yewon, et al.
Published: (2026)
Confidence-aware Reward Optimization for Fine-tuning Text-to-Image Models
by: Kim, Kyuyoung, et al.
Published: (2024)
by: Kim, Kyuyoung, et al.
Published: (2024)
MoSKA: Mixture of Shared KV Attention for Efficient Long-Sequence LLM Inference
by: Rhee, Myunghyun, et al.
Published: (2025)
by: Rhee, Myunghyun, et al.
Published: (2025)
Unveiling Imitation Learning: Exploring the Impact of Data Falsity to Large Language Model
by: Cho, Hyunsoo
Published: (2024)
by: Cho, Hyunsoo
Published: (2024)
Analyzing and Improving Optimal-Transport-based Adversarial Networks
by: Choi, Jaemoo, et al.
Published: (2023)
by: Choi, Jaemoo, et al.
Published: (2023)
Generative Modeling through the Semi-dual Formulation of Unbalanced Optimal Transport
by: Choi, Jaemoo, et al.
Published: (2023)
by: Choi, Jaemoo, et al.
Published: (2023)
Scalable Wasserstein Gradient Flow for Generative Modeling through Unbalanced Optimal Transport
by: Choi, Jaemoo, et al.
Published: (2024)
by: Choi, Jaemoo, et al.
Published: (2024)
Improving Neural Optimal Transport via Displacement Interpolation
by: Choi, Jaemoo, et al.
Published: (2024)
by: Choi, Jaemoo, et al.
Published: (2024)
FALQON: Accelerating LoRA Fine-tuning with Low-Bit Floating-Point Arithmetic
by: Choi, Kanghyun, et al.
Published: (2025)
by: Choi, Kanghyun, et al.
Published: (2025)
Graph Spectral Filtering with Chebyshev Interpolation for Recommendation
by: Kim, Chanwoo, et al.
Published: (2025)
by: Kim, Chanwoo, et al.
Published: (2025)
ETR: Outcome-Guided Elastic Trust Regions for Policy Optimization
by: Zhang, Shijie, et al.
Published: (2026)
by: Zhang, Shijie, et al.
Published: (2026)
Trust Region Q Adjoint Matching
by: Dong, Yonghoon, et al.
Published: (2026)
by: Dong, Yonghoon, et al.
Published: (2026)
Trust Regions Sell, But Who's Buying? Overlap Geometry as an Alternative Trust Region for Policy Optimization
by: Trivedi, Gaurish, et al.
Published: (2026)
by: Trivedi, Gaurish, et al.
Published: (2026)
TripleSumm: Adaptive Triple-Modality Fusion for Video Summarization
by: Kim, Sumin, et al.
Published: (2026)
by: Kim, Sumin, et al.
Published: (2026)
Matrix Low-Rank Trust Region Policy Optimization
by: Rozada, Sergio, et al.
Published: (2024)
by: Rozada, Sergio, et al.
Published: (2024)
Behavioral Mode Discovery for Fine-tuning Multimodal Generative Policies
by: Longhini, Alberta, et al.
Published: (2026)
by: Longhini, Alberta, et al.
Published: (2026)
SummDiff: Generative Modeling of Video Summarization with Diffusion
by: Kim, Kwanseok, et al.
Published: (2025)
by: Kim, Kwanseok, et al.
Published: (2025)
Query-Efficient Quantum Approximate Optimization via Graph-Conditioned Trust Regions
by: Huynh, Molena
Published: (2026)
by: Huynh, Molena
Published: (2026)
Equivariant Latent Alignment via Flow Matching under Group Symmetries
by: Kim, Sunghyun, et al.
Published: (2026)
by: Kim, Sunghyun, et al.
Published: (2026)
Let Multimodal Embedders Learn When to Augment Query via Adaptive Query Augmentation
by: Kim, Wongyu, et al.
Published: (2025)
by: Kim, Wongyu, et al.
Published: (2025)
Trust-Region Twisted Policy Improvement
by: de Vries, Joery A., et al.
Published: (2025)
by: de Vries, Joery A., et al.
Published: (2025)
Human Implicit Preference-Based Policy Fine-tuning for Multi-Agent Reinforcement Learning in USV Swarm
by: Kim, Hyeonjun, et al.
Published: (2025)
by: Kim, Hyeonjun, et al.
Published: (2025)
Generalized Schrödinger Bridge on Graphs
by: Theodoropoulos, Panagiotis, et al.
Published: (2026)
by: Theodoropoulos, Panagiotis, et al.
Published: (2026)
Similar Items
-
Modality-Aware Representation Learning for Zero-shot Sketch-based Image Retrieval
by: Lyou, Eunyi, et al.
Published: (2024) -
Efficient Adjoint Matching for Fine-tuning Diffusion Models
by: Shin, Jeongwoo, et al.
Published: (2026) -
A More Word-like Image Tokenization for MLLMs
by: Lee, Hyun, et al.
Published: (2026) -
Trust-Region Adaptive Policy Optimization
by: Su, Mingyu, et al.
Published: (2025) -
Coarse-to-Fine Compositional Diffusion for Long-Horizon Planning
by: Park, Byoungwoo, et al.
Published: (2026)