On Giant's Shoulders: Effortless Weak to Strong by Dynamic Logits Fusion
Fuente:
arXiv
Salvato in:
| Autori principali: | Fan, Chenghao, Lu, Zhenyi, Wei, Wei, Tian, Jie, Qu, Xiaoye, Chen, Dangyang, Cheng, Yu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Twin-Merging: Dynamic Integration of Modular Expertise in Model Merging
di: Lu, Zhenyi, et al.
Pubblicazione: (2024)
di: Lu, Zhenyi, et al.
Pubblicazione: (2024)
Enhancing Low-Resource Relation Representations through Multi-View Decoupling
di: Fan, Chenghao, et al.
Pubblicazione: (2023)
di: Fan, Chenghao, et al.
Pubblicazione: (2023)
Mitigating Boundary Ambiguity and Inherent Bias for Text Classification in the Era of Large Language Models
di: Lu, Zhenyi, et al.
Pubblicazione: (2024)
di: Lu, Zhenyi, et al.
Pubblicazione: (2024)
Improving Pseudo Labels with Global-Local Denoising Framework for Cross-lingual Named Entity Recognition
di: Ding, Zhuojun, et al.
Pubblicazione: (2024)
di: Ding, Zhuojun, et al.
Pubblicazione: (2024)
Joint Multi-Facts Reasoning Network For Complex Temporal Question Answering Over Knowledge Graph
di: Huang, Rikui, et al.
Pubblicazione: (2024)
di: Huang, Rikui, et al.
Pubblicazione: (2024)
Intelligent Go-Explore: Standing on the Shoulders of Giant Foundation Models
di: Lu, Cong, et al.
Pubblicazione: (2024)
di: Lu, Cong, et al.
Pubblicazione: (2024)
Weak-to-Strong Elicitation via Mismatched Wrong Drafts
di: Deng, Wei
Pubblicazione: (2026)
di: Deng, Wei
Pubblicazione: (2026)
LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation
di: Zhang, Xuan, et al.
Pubblicazione: (2024)
di: Zhang, Xuan, et al.
Pubblicazione: (2024)
LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid
di: Sun, Weigao, et al.
Pubblicazione: (2025)
di: Sun, Weigao, et al.
Pubblicazione: (2025)
PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models
di: Song, Mingyang, et al.
Pubblicazione: (2025)
di: Song, Mingyang, et al.
Pubblicazione: (2025)
Reinforcement Learning with Token-level Feedback for Controllable Text Generation
di: Li, Wendi, et al.
Pubblicazione: (2024)
di: Li, Wendi, et al.
Pubblicazione: (2024)
Personalized Topic Selection Model for Topic-Grounded Dialogue
di: Fan, Shixuan, et al.
Pubblicazione: (2024)
di: Fan, Shixuan, et al.
Pubblicazione: (2024)
Spectral Logit Sculpting: Adaptive Low-Rank Logit Transformation for Controlled Text Generation
di: Li, Jin, et al.
Pubblicazione: (2025)
di: Li, Jin, et al.
Pubblicazione: (2025)
API Is Enough: Conformal Prediction for Large Language Models Without Logit-Access
di: Su, Jiayuan, et al.
Pubblicazione: (2024)
di: Su, Jiayuan, et al.
Pubblicazione: (2024)
Make LoRA Great Again: Boosting LoRA with Adaptive Singular Values and Mixture-of-Experts Optimization Alignment
di: Fan, Chenghao, et al.
Pubblicazione: (2025)
di: Fan, Chenghao, et al.
Pubblicazione: (2025)
Position Debiasing Fine-Tuning for Causal Perception in Long-Term Dialogue
di: Fan, Shixuan, et al.
Pubblicazione: (2024)
di: Fan, Shixuan, et al.
Pubblicazione: (2024)
Weak-to-Strong Search: Align Large Language Models via Searching over Small Language Models
di: Zhou, Zhanhui, et al.
Pubblicazione: (2024)
di: Zhou, Zhanhui, et al.
Pubblicazione: (2024)
Mitigating Multilingual Hallucination in Large Vision-Language Models
di: Qu, Xiaoye, et al.
Pubblicazione: (2024)
di: Qu, Xiaoye, et al.
Pubblicazione: (2024)
SLED: Self Logits Evolution Decoding for Improving Factuality in Large Language Models
di: Zhang, Jianyi, et al.
Pubblicazione: (2024)
di: Zhang, Jianyi, et al.
Pubblicazione: (2024)
Learning to Reason under Off-Policy Guidance
di: Yan, Jianhao, et al.
Pubblicazione: (2025)
di: Yan, Jianhao, et al.
Pubblicazione: (2025)
Inferring from Logits: Exploring Best Practices for Decoding-Free Generative Candidate Selection
di: Ma, Mingyu Derek, et al.
Pubblicazione: (2025)
di: Ma, Mingyu Derek, et al.
Pubblicazione: (2025)
Linear-MoE: Linear Sequence Modeling Meets Mixture-of-Experts
di: Sun, Weigao, et al.
Pubblicazione: (2025)
di: Sun, Weigao, et al.
Pubblicazione: (2025)
Pensieve Grader: An AI-Powered, Ready-to-Use Platform for Effortless Handwritten STEM Grading
di: Yang, Yoonseok, et al.
Pubblicazione: (2025)
di: Yang, Yoonseok, et al.
Pubblicazione: (2025)
Bayesian WeakS-to-Strong from Text Classification to Generation
di: Cui, Ziyun, et al.
Pubblicazione: (2024)
di: Cui, Ziyun, et al.
Pubblicazione: (2024)
ExGRPO: Learning to Reason from Experience
di: Zhan, Runzhe, et al.
Pubblicazione: (2025)
di: Zhan, Runzhe, et al.
Pubblicazione: (2025)
DC-W2S: Dual-Consensus Weak-to-Strong Training for Reliable Process Reward Modeling in Biological Reasoning
di: Chan, Chi-Min, et al.
Pubblicazione: (2026)
di: Chan, Chi-Min, et al.
Pubblicazione: (2026)
The Amazing Agent Race: Strong Tool Users, Weak Navigators
di: Kim, Zae Myung, et al.
Pubblicazione: (2026)
di: Kim, Zae Myung, et al.
Pubblicazione: (2026)
Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models
di: Chen, Zixiang, et al.
Pubblicazione: (2024)
di: Chen, Zixiang, et al.
Pubblicazione: (2024)
Min-$k$ Sampling: Decoupling Truncation from Temperature Scaling via Relative Logit Dynamics
di: Ding, Yuanhao, et al.
Pubblicazione: (2026)
di: Ding, Yuanhao, et al.
Pubblicazione: (2026)
Chinese-Vicuna: A Chinese Instruction-following Llama-based Model
di: Fan, Chenghao, et al.
Pubblicazione: (2025)
di: Fan, Chenghao, et al.
Pubblicazione: (2025)
Super(ficial)-alignment: Strong Models May Deceive Weak Models in Weak-to-Strong Generalization
di: Yang, Wenkai, et al.
Pubblicazione: (2024)
di: Yang, Wenkai, et al.
Pubblicazione: (2024)
Logits are All We Need to Adapt Closed Models
di: Hiranandani, Gaurush, et al.
Pubblicazione: (2025)
di: Hiranandani, Gaurush, et al.
Pubblicazione: (2025)
PCL-Reasoner-V1.5: Advancing Math Reasoning with Offline Reinforcement Learning
di: Lu, Yao, et al.
Pubblicazione: (2026)
di: Lu, Yao, et al.
Pubblicazione: (2026)
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models
di: Fu, Tingchen, et al.
Pubblicazione: (2025)
di: Fu, Tingchen, et al.
Pubblicazione: (2025)
How Do Latent Reasoning Methods Perform Under Weak and Strong Supervision?
di: Cui, Yingqian, et al.
Pubblicazione: (2026)
di: Cui, Yingqian, et al.
Pubblicazione: (2026)
Sequences of Logits Reveal the Low Rank Structure of Language Models
di: Golowich, Noah, et al.
Pubblicazione: (2025)
di: Golowich, Noah, et al.
Pubblicazione: (2025)
Earley-Driven Dynamic Pruning for Efficient Structured Decoding
di: Sun, Xintong, et al.
Pubblicazione: (2025)
di: Sun, Xintong, et al.
Pubblicazione: (2025)
Well Begun is Half Done: Low-resource Preference Alignment by Weak-to-Strong Decoding
di: Song, Feifan, et al.
Pubblicazione: (2025)
di: Song, Feifan, et al.
Pubblicazione: (2025)
AutoDetect: Towards a Unified Framework for Automated Weakness Detection in Large Language Models
di: Cheng, Jiale, et al.
Pubblicazione: (2024)
di: Cheng, Jiale, et al.
Pubblicazione: (2024)
DLO: Dynamic Layer Operation for Efficient Vertical Scaling of LLMs
di: Tan, Zhen, et al.
Pubblicazione: (2024)
di: Tan, Zhen, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Twin-Merging: Dynamic Integration of Modular Expertise in Model Merging
di: Lu, Zhenyi, et al.
Pubblicazione: (2024) -
Enhancing Low-Resource Relation Representations through Multi-View Decoupling
di: Fan, Chenghao, et al.
Pubblicazione: (2023) -
Mitigating Boundary Ambiguity and Inherent Bias for Text Classification in the Era of Large Language Models
di: Lu, Zhenyi, et al.
Pubblicazione: (2024) -
Improving Pseudo Labels with Global-Local Denoising Framework for Cross-lingual Named Entity Recognition
di: Ding, Zhuojun, et al.
Pubblicazione: (2024) -
Joint Multi-Facts Reasoning Network For Complex Temporal Question Answering Over Knowledge Graph
di: Huang, Rikui, et al.
Pubblicazione: (2024)