Logits-Based Finetuning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Jingyao, Yang, Senqiao, Wu, Sitong, Shi, Han, Zheng, Chuanyang, Xu, Hong, Jia, Jiaya |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
RoboCoder: Robotic Learning from Basic Skills to General Tasks with Large Language Models
von: Li, Jingyao, et al.
Veröffentlicht: (2024)
von: Li, Jingyao, et al.
Veröffentlicht: (2024)
QuickLLaMA: Query-aware Inference Acceleration for Large Language Models
von: Li, Jingyao, et al.
Veröffentlicht: (2024)
von: Li, Jingyao, et al.
Veröffentlicht: (2024)
VisionZip: Longer is Better but Not Necessary in Vision Language Models
von: Yang, Senqiao, et al.
Veröffentlicht: (2024)
von: Yang, Senqiao, et al.
Veröffentlicht: (2024)
MoTCoder: Elevating Large Language Models with Modular of Thought for Challenging Programming Tasks
von: Li, Jingyao, et al.
Veröffentlicht: (2023)
von: Li, Jingyao, et al.
Veröffentlicht: (2023)
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
von: Lai, Xin, et al.
Veröffentlicht: (2024)
von: Lai, Xin, et al.
Veröffentlicht: (2024)
VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning
von: Yang, Senqiao, et al.
Veröffentlicht: (2025)
von: Yang, Senqiao, et al.
Veröffentlicht: (2025)
DynamicBench: Evaluating Real-Time Report Generation in Large Language Models
von: Li, Jingyao, et al.
Veröffentlicht: (2025)
von: Li, Jingyao, et al.
Veröffentlicht: (2025)
iFormer: Integrating ConvNet and Transformer for Mobile Application
von: Zheng, Chuanyang
Veröffentlicht: (2025)
von: Zheng, Chuanyang
Veröffentlicht: (2025)
SmartSwitch: Advancing LLM Reasoning by Overcoming Underthinking via Promoting Deeper Thought Exploration
von: Zhang, Xichen, et al.
Veröffentlicht: (2025)
von: Zhang, Xichen, et al.
Veröffentlicht: (2025)
Scaf-GRPO: Scaffolded Group Relative Policy Optimization for Enhancing LLM Reasoning
von: Zhang, Xichen, et al.
Veröffentlicht: (2025)
von: Zhang, Xichen, et al.
Veröffentlicht: (2025)
UniPool: A Globally Shared Expert Pool for Mixture-of-Experts
von: Huang, Minbin, et al.
Veröffentlicht: (2026)
von: Huang, Minbin, et al.
Veröffentlicht: (2026)
LoCa: Logit Calibration for Knowledge Distillation
von: Yang, Runming, et al.
Veröffentlicht: (2024)
von: Yang, Runming, et al.
Veröffentlicht: (2024)
DALD: Improving Logits-based Detector without Logits from Black-box LLMs
von: Zeng, Cong, et al.
Veröffentlicht: (2024)
von: Zeng, Cong, et al.
Veröffentlicht: (2024)
Unified Language-driven Zero-shot Domain Adaptation
von: Yang, Senqiao, et al.
Veröffentlicht: (2024)
von: Yang, Senqiao, et al.
Veröffentlicht: (2024)
Spectral Logit Sculpting: Adaptive Low-Rank Logit Transformation for Controlled Text Generation
von: Li, Jin, et al.
Veröffentlicht: (2025)
von: Li, Jin, et al.
Veröffentlicht: (2025)
Orthogonal Finetuning for Direct Preference Optimization
von: Yang, Chenxu, et al.
Veröffentlicht: (2024)
von: Yang, Chenxu, et al.
Veröffentlicht: (2024)
Online Continual Learning via Logit Adjusted Softmax
von: Huang, Zhehao, et al.
Veröffentlicht: (2023)
von: Huang, Zhehao, et al.
Veröffentlicht: (2023)
Logits Poisoning Attack in Federated Distillation
von: Tang, Yuhan, et al.
Veröffentlicht: (2024)
von: Tang, Yuhan, et al.
Veröffentlicht: (2024)
Consistency Regularization for Domain Generalization with Logit Attribution Matching
von: Gao, Han, et al.
Veröffentlicht: (2023)
von: Gao, Han, et al.
Veröffentlicht: (2023)
Top-$nσ$: Not All Logits Are You Need
von: Tang, Chenxia, et al.
Veröffentlicht: (2024)
von: Tang, Chenxia, et al.
Veröffentlicht: (2024)
Logit Distillation on Manifolds: Mapping by Learning
von: Yang, Yiru, et al.
Veröffentlicht: (2026)
von: Yang, Yiru, et al.
Veröffentlicht: (2026)
Formalising the Logit Shift Induced by LoRA: A Technical Note
von: Shi, Xiang, et al.
Veröffentlicht: (2026)
von: Shi, Xiang, et al.
Veröffentlicht: (2026)
Progressive-Hint Prompting Improves Reasoning in Large Language Models
von: Zheng, Chuanyang, et al.
Veröffentlicht: (2023)
von: Zheng, Chuanyang, et al.
Veröffentlicht: (2023)
Enhancing Certified Robustness via Block Reflector Orthogonal Layers and Logit Annealing Loss
von: Lai, Bo-Han, et al.
Veröffentlicht: (2025)
von: Lai, Bo-Han, et al.
Veröffentlicht: (2025)
Efficient Reinforcement Finetuning via Adaptive Curriculum Learning
von: Shi, Taiwei, et al.
Veröffentlicht: (2025)
von: Shi, Taiwei, et al.
Veröffentlicht: (2025)
The Finetuner's Fallacy: When to Pretrain with Your Finetuning Data
von: Baek, Christina, et al.
Veröffentlicht: (2026)
von: Baek, Christina, et al.
Veröffentlicht: (2026)
Deconstructing Positional Information: From Attention Logits to Training Biases
von: Gu, Zihan, et al.
Veröffentlicht: (2025)
von: Gu, Zihan, et al.
Veröffentlicht: (2025)
SHRED: Retain-Set-Free Unlearning via Self-Distillation with Logit Demotion
von: Hu, Zizhao, et al.
Veröffentlicht: (2026)
von: Hu, Zizhao, et al.
Veröffentlicht: (2026)
The Implicit Bias of Logit Regularization
von: Beck, Alon, et al.
Veröffentlicht: (2026)
von: Beck, Alon, et al.
Veröffentlicht: (2026)
Reasoning-Finetuning Repurposes Latent Representations in Base Models
von: Ward, Jake, et al.
Veröffentlicht: (2025)
von: Ward, Jake, et al.
Veröffentlicht: (2025)
Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment
von: Li, Yuxi, et al.
Veröffentlicht: (2024)
von: Li, Yuxi, et al.
Veröffentlicht: (2024)
Logit Dynamics in Softmax Policy Gradient Methods
von: Li, Yingru
Veröffentlicht: (2025)
von: Li, Yingru
Veröffentlicht: (2025)
SSR-Zero: Simple Self-Rewarding Reinforcement Learning for Machine Translation
von: Yang, Wenjie, et al.
Veröffentlicht: (2025)
von: Yang, Wenjie, et al.
Veröffentlicht: (2025)
Machine Learning-Based Research on the Adaptability of Adolescents to Online Education
von: Wang, Mingwei, et al.
Veröffentlicht: (2024)
von: Wang, Mingwei, et al.
Veröffentlicht: (2024)
Learning Dynamics of VLM Finetuning
von: Zhang, Jusheng, et al.
Veröffentlicht: (2025)
von: Zhang, Jusheng, et al.
Veröffentlicht: (2025)
Towards Few-Shot Adaptation of Foundation Models via Multitask Finetuning
von: Xu, Zhuoyan, et al.
Veröffentlicht: (2024)
von: Xu, Zhuoyan, et al.
Veröffentlicht: (2024)
Representation Finetuning for Continual Learning
von: Luo, Haihua, et al.
Veröffentlicht: (2026)
von: Luo, Haihua, et al.
Veröffentlicht: (2026)
Stabilizing Policy Optimization via Logits Convexity
von: Chen, Hongzhan, et al.
Veröffentlicht: (2026)
von: Chen, Hongzhan, et al.
Veröffentlicht: (2026)
Provably Efficient Reinforcement Learning with Multinomial Logit Function Approximation
von: Li, Long-Fei, et al.
Veröffentlicht: (2024)
von: Li, Long-Fei, et al.
Veröffentlicht: (2024)
On the Out-of-Distribution Generalization of Self-Supervised Learning
von: Qiang, Wenwen, et al.
Veröffentlicht: (2025)
von: Qiang, Wenwen, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
RoboCoder: Robotic Learning from Basic Skills to General Tasks with Large Language Models
von: Li, Jingyao, et al.
Veröffentlicht: (2024) -
QuickLLaMA: Query-aware Inference Acceleration for Large Language Models
von: Li, Jingyao, et al.
Veröffentlicht: (2024) -
VisionZip: Longer is Better but Not Necessary in Vision Language Models
von: Yang, Senqiao, et al.
Veröffentlicht: (2024) -
MoTCoder: Elevating Large Language Models with Modular of Thought for Challenging Programming Tasks
von: Li, Jingyao, et al.
Veröffentlicht: (2023) -
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
von: Lai, Xin, et al.
Veröffentlicht: (2024)