TIS-DPO: Token-level Importance Sampling for Direct Preference Optimization With Estimated Weights
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Aiwei, Bai, Haoping, Lu, Zhiyun, Sun, Yanchao, Kong, Xiang, Wang, Simon, Shan, Jiulong, Jose, Albin Madappally, Liu, Xiaojiang, Wen, Lijie, Yu, Philip S., Cao, Meng |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Direct Large Language Model Alignment Through Self-Rewarding Contrastive Prompt Distillation
by: Liu, Aiwei, et al.
Published: (2024)
by: Liu, Aiwei, et al.
Published: (2024)
A Semantic Invariant Robust Watermark for Large Language Models
by: Liu, Aiwei, et al.
Published: (2023)
by: Liu, Aiwei, et al.
Published: (2023)
An Unforgeable Publicly Verifiable Watermark for Large Language Models
by: Liu, Aiwei, et al.
Published: (2023)
by: Liu, Aiwei, et al.
Published: (2023)
ChatCite: LLM Agent with Human Workflow Guidance for Comparative Literature Summary
by: Li, Yutong, et al.
Published: (2024)
by: Li, Yutong, et al.
Published: (2024)
Can LLM Watermarks Robustly Prevent Unauthorized Knowledge Distillation?
by: Pan, Leyi, et al.
Published: (2025)
by: Pan, Leyi, et al.
Published: (2025)
On the Robustness of Document-Level Relation Extraction Models to Entity Name Variations
by: Meng, Shiao, et al.
Published: (2024)
by: Meng, Shiao, et al.
Published: (2024)
WaterSeeker: Pioneering Efficient Detection of Watermarked Segments in Large Documents
by: Pan, Leyi, et al.
Published: (2024)
by: Pan, Leyi, et al.
Published: (2024)
Can Watermarked LLMs be Identified by Users via Crafted Prompts?
by: Liu, Aiwei, et al.
Published: (2024)
by: Liu, Aiwei, et al.
Published: (2024)
Bridging SFT and DPO for Diffusion Model Alignment with Self-Sampling Preference Optimization
by: Zhang, Daoan, et al.
Published: (2024)
by: Zhang, Daoan, et al.
Published: (2024)
A Survey of Text Watermarking in the Era of Large Language Models
by: Liu, Aiwei, et al.
Published: (2023)
by: Liu, Aiwei, et al.
Published: (2023)
d-TreeRPO: Towards More Reliable Policy Optimization for Diffusion Language Models
by: Pan, Leyi, et al.
Published: (2025)
by: Pan, Leyi, et al.
Published: (2025)
MarkLLM: An Open-Source Toolkit for LLM Watermarking
by: Pan, Leyi, et al.
Published: (2024)
by: Pan, Leyi, et al.
Published: (2024)
Omni-SafetyBench: A Benchmark for Safety Evaluation of Audio-Visual Large Language Models
by: Pan, Leyi, et al.
Published: (2025)
by: Pan, Leyi, et al.
Published: (2025)
MarkDiffusion: An Open-Source Toolkit for Generative Watermarking of Latent Diffusion Models
by: Pan, Leyi, et al.
Published: (2025)
by: Pan, Leyi, et al.
Published: (2025)
Token-Importance Guided Direct Preference Optimization
by: Yang, Ning, et al.
Published: (2025)
by: Yang, Ning, et al.
Published: (2025)
WeDLM: Reconciling Diffusion Language Models with Standard Causal Attention for Fast Inference
by: Liu, Aiwei, et al.
Published: (2025)
by: Liu, Aiwei, et al.
Published: (2025)
FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models
by: Fu, Tianyu, et al.
Published: (2024)
by: Fu, Tianyu, et al.
Published: (2024)
A Class of Optimal Directed Graphs for Network Synchronization
by: Lu, Susie, et al.
Published: (2025)
by: Lu, Susie, et al.
Published: (2025)
An invariant-theoretic approach to three weight enumerators of self-dual quantum codes
by: Chen, Yin, et al.
Published: (2024)
by: Chen, Yin, et al.
Published: (2024)
On Modules Whose Pure Submodules Are Essential in Direct Summands
by: Gupta, Kaushal, et al.
Published: (2025)
by: Gupta, Kaushal, et al.
Published: (2025)
Which Pieces Does Unigram Tokenization Really Need?
by: Land, Sander, et al.
Published: (2025)
by: Land, Sander, et al.
Published: (2025)
Prescription of finite Dirichlet eigenvalues and area on surface with boundary
by: He, Xiang
Published: (2023)
by: He, Xiang
Published: (2023)
Modular invariants of a vector and a covector for some elementary abelian $p$-groups
by: Ren, Shan
Published: (2023)
by: Ren, Shan
Published: (2023)
Fishing for Magikarp: Automatically Detecting Under-trained Tokens in Large Language Models
by: Land, Sander, et al.
Published: (2024)
by: Land, Sander, et al.
Published: (2024)
ActiveDPO: Active Direct Preference Optimization for Sample-Efficient Alignment
by: Lin, Xiaoqiang, et al.
Published: (2025)
by: Lin, Xiaoqiang, et al.
Published: (2025)
$β$-DPO: Direct Preference Optimization with Dynamic $β$
by: Wu, Junkang, et al.
Published: (2024)
by: Wu, Junkang, et al.
Published: (2024)
Insights into Weighted Sum Sampling Approaches for Multi-Criteria Decision Making Problems
by: Williams, Aled, et al.
Published: (2024)
by: Williams, Aled, et al.
Published: (2024)
A Survey on Parallel Text Generation: From Parallel Decoding to Diffusion Language Models
by: Zhang, Lingzhe, et al.
Published: (2025)
by: Zhang, Lingzhe, et al.
Published: (2025)
Hydrography and water chemistry of the open Mediterranean Sea
by: Kremling, Klaus, et al.
Published: (1981)
by: Kremling, Klaus, et al.
Published: (1981)
Some four-dimensional orthogonal invariants
by: Ren, Shan, et al.
Published: (2025)
by: Ren, Shan, et al.
Published: (2025)
Modular matrix invariants under some transpose actions
by: Chen, Yin, et al.
Published: (2025)
by: Chen, Yin, et al.
Published: (2025)
Weighted Heights and GIT Heights
by: Shaska, Elira, et al.
Published: (2025)
by: Shaska, Elira, et al.
Published: (2025)
The growth rate on the volume of $\mathcal{M}_g^{<L(g)}$
by: Liu, Jinsong, et al.
Published: (2025)
by: Liu, Jinsong, et al.
Published: (2025)
Operadic Deformation Theory
by: Campos, Ricardo, et al.
Published: (2023)
by: Campos, Ricardo, et al.
Published: (2023)
AlphaDPO: Adaptive Reward Margin for Direct Preference Optimization
by: Wu, Junkang, et al.
Published: (2024)
by: Wu, Junkang, et al.
Published: (2024)
Vertex Weight Reconstruction in the Gel'fand's Inverse Problem on Connected Weighted Graphs
by: Li, Songshuo, et al.
Published: (2024)
by: Li, Songshuo, et al.
Published: (2024)
A Priori Log-Concavity Estimates for Dirichlet Eigenfunctions
by: Khan, Gabriel, et al.
Published: (2025)
by: Khan, Gabriel, et al.
Published: (2025)
Uncertainty Estimation and Quantification for LLMs: A Simple Supervised Approach
by: Liu, Linyu, et al.
Published: (2024)
by: Liu, Linyu, et al.
Published: (2024)
Tokenization Is More Than Compression
by: Schmidt, Craig W., et al.
Published: (2024)
by: Schmidt, Craig W., et al.
Published: (2024)
Reorientation Response of Magnetic Microspheres Attached to Gold Electrodes Under an Applied Magnetic Field
by: L. De Los Santos Valladares
Published: (2013)
by: L. De Los Santos Valladares
Published: (2013)
Similar Items
-
Direct Large Language Model Alignment Through Self-Rewarding Contrastive Prompt Distillation
by: Liu, Aiwei, et al.
Published: (2024) -
A Semantic Invariant Robust Watermark for Large Language Models
by: Liu, Aiwei, et al.
Published: (2023) -
An Unforgeable Publicly Verifiable Watermark for Large Language Models
by: Liu, Aiwei, et al.
Published: (2023) -
ChatCite: LLM Agent with Human Workflow Guidance for Comparative Literature Summary
by: Li, Yutong, et al.
Published: (2024) -
Can LLM Watermarks Robustly Prevent Unauthorized Knowledge Distillation?
by: Pan, Leyi, et al.
Published: (2025)