Cornerstones or Stumbling Blocks? Deciphering the Rock Tokens in On-Policy Distillation
Fuente:
arXiv
Salvato in:
| Autori principali: | Jiang, Yuxuan, Li, Runchao, Dipta, Shubhashis Roy, Li, Dawei, Yang, Zhao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Learning How to Use Tools, Not Just When: Pattern-Aware Tool-Integrated Reasoning
di: Xu, Ningning, et al.
Pubblicazione: (2025)
di: Xu, Ningning, et al.
Pubblicazione: (2025)
PromptGuard at BLP-2025 Task 1: A Few-Shot Classification Framework Using Majority Voting and Keyword Similarity for Bengali Hate Speech Detection
di: Hossan, Rakib, et al.
Pubblicazione: (2025)
di: Hossan, Rakib, et al.
Pubblicazione: (2025)
UMBCLU at SemEval-2024 Task 1A and 1C: Semantic Textual Relatedness with and without machine translation
di: Dipta, Shubhashis Roy, et al.
Pubblicazione: (2024)
di: Dipta, Shubhashis Roy, et al.
Pubblicazione: (2024)
HU at SemEval-2024 Task 8A: Can Contrastive Learning Learn Embeddings to Detect Machine-Generated Text?
di: Dipta, Shubhashis Roy, et al.
Pubblicazione: (2024)
di: Dipta, Shubhashis Roy, et al.
Pubblicazione: (2024)
BanglaTalk: Towards Real-Time Speech Assistance for Bengali Regional Dialects
di: Hasan, Jakir, et al.
Pubblicazione: (2025)
di: Hasan, Jakir, et al.
Pubblicazione: (2025)
FedMentor: Domain-Aware Differential Privacy for Heterogeneous Federated LLMs in Mental Health
di: Sarwar, Nobin, et al.
Pubblicazione: (2025)
di: Sarwar, Nobin, et al.
Pubblicazione: (2025)
DecomposeRL: Learning to Ask Useful, Informative, and Diverse Questions for Semi-Supervised, Traceable Claim Verification
di: Dipta, Shubhashis Roy, et al.
Pubblicazione: (2026)
di: Dipta, Shubhashis Roy, et al.
Pubblicazione: (2026)
GanitLLM: Difficulty-Aware Bengali Mathematical Reasoning through Curriculum-GRPO
di: Dipta, Shubhashis Roy, et al.
Pubblicazione: (2026)
di: Dipta, Shubhashis Roy, et al.
Pubblicazione: (2026)
PA3: Policy-Aware Agent Alignment through Chain-of-Thought
di: Dipta, Shubhashis Roy, et al.
Pubblicazione: (2026)
di: Dipta, Shubhashis Roy, et al.
Pubblicazione: (2026)
Breaking the Silence: A Dataset and Benchmark for Bangla Text-to-Gloss Translation
di: Abdullah, Sharif Mohammad, et al.
Pubblicazione: (2025)
di: Abdullah, Sharif Mohammad, et al.
Pubblicazione: (2025)
BanglaLlama: LLaMA for Bangla Language
di: Zehady, Abdullah Khan, et al.
Pubblicazione: (2024)
di: Zehady, Abdullah Khan, et al.
Pubblicazione: (2024)
If We May De-Presuppose: Robustly Verifying Claims through Presupposition-Free Question Decomposition
di: Dipta, Shubhashis Roy, et al.
Pubblicazione: (2025)
di: Dipta, Shubhashis Roy, et al.
Pubblicazione: (2025)
Bridging Reasoning Trajectories in On-Policy Distillation via Near-Future Guidance
di: Jiang, Yuxuan, et al.
Pubblicazione: (2026)
di: Jiang, Yuxuan, et al.
Pubblicazione: (2026)
Q2E: Query-to-Event Decomposition for Zero-Shot Multilingual Text-to-Video Retrieval
di: Dipta, Shubhashis Roy, et al.
Pubblicazione: (2025)
di: Dipta, Shubhashis Roy, et al.
Pubblicazione: (2025)
TRIAGE: Evaluating Prospective Metacognitive Control in LLMs under Resource Constraints
di: Nazi, Zabir Al, et al.
Pubblicazione: (2026)
di: Nazi, Zabir Al, et al.
Pubblicazione: (2026)
Cross-Lingual Sentiment Misalignment: Auditing Multilingual Language Models for Inversion Risk, Dialectal Representation, and Affective Stability
di: Lia, Nusrat Jahan, et al.
Pubblicazione: (2026)
di: Lia, Nusrat Jahan, et al.
Pubblicazione: (2026)
Towards Generalization of Block Attention via Automatic Segmentation and Block Distillation
di: Li, Shuaiyi, et al.
Pubblicazione: (2026)
di: Li, Shuaiyi, et al.
Pubblicazione: (2026)
Self-Distillation for Multi-Token Prediction
di: Zhao, Guoliang, et al.
Pubblicazione: (2026)
di: Zhao, Guoliang, et al.
Pubblicazione: (2026)
LLM-Based Section Identifiers Excel on Open Source but Stumble in Real World Applications
di: Krishnamoorthy, Saranya, et al.
Pubblicazione: (2024)
di: Krishnamoorthy, Saranya, et al.
Pubblicazione: (2024)
AlignDistil: Token-Level Language Model Alignment as Adaptive Policy Distillation
di: Zhang, Songming, et al.
Pubblicazione: (2025)
di: Zhang, Songming, et al.
Pubblicazione: (2025)
LLM-Oriented Token-Adaptive Knowledge Distillation
di: Xie, Xurong, et al.
Pubblicazione: (2025)
di: Xie, Xurong, et al.
Pubblicazione: (2025)
Contextualization Distillation from Large Language Model for Knowledge Graph Completion
di: Li, Dawei, et al.
Pubblicazione: (2024)
di: Li, Dawei, et al.
Pubblicazione: (2024)
Tokenization Matters! Degrading Large Language Models through Challenging Their Tokenization
di: Wang, Dixuan, et al.
Pubblicazione: (2024)
di: Wang, Dixuan, et al.
Pubblicazione: (2024)
MiniLLM: On-Policy Distillation of Large Language Models
di: Gu, Yuxian, et al.
Pubblicazione: (2023)
di: Gu, Yuxian, et al.
Pubblicazione: (2023)
LAWCAT: Efficient Distillation from Quadratic to Linear Attention with Convolution across Tokens for Long Context Modeling
di: Liu, Zeyu, et al.
Pubblicazione: (2025)
di: Liu, Zeyu, et al.
Pubblicazione: (2025)
VC-Inspector: Advancing Reference-free Evaluation of Video Captions with Factual Analysis
di: Dipta, Shubhashis Roy, et al.
Pubblicazione: (2025)
di: Dipta, Shubhashis Roy, et al.
Pubblicazione: (2025)
IGOT: Information Gain Optimized Tokenizer on Domain Adaptive Pretraining
di: Feng, Dawei, et al.
Pubblicazione: (2024)
di: Feng, Dawei, et al.
Pubblicazione: (2024)
SafeSteer: Localized On-Policy Distillation for Efficient Safety Alignment
di: Li, Hao, et al.
Pubblicazione: (2026)
di: Li, Hao, et al.
Pubblicazione: (2026)
†DAGGER: Distractor-Aware Graph Generation for Executable Reasoning in Math Problems
di: Nazi, Zabir Al, et al.
Pubblicazione: (2026)
di: Nazi, Zabir Al, et al.
Pubblicazione: (2026)
AdaSwitch: Balancing Exploration and Guidance in Knowledge Distillation via Adaptive Switching
di: Peng, Jingyu, et al.
Pubblicazione: (2025)
di: Peng, Jingyu, et al.
Pubblicazione: (2025)
Black-Box On-Policy Distillation of Large Language Models
di: Ye, Tianzhu, et al.
Pubblicazione: (2025)
di: Ye, Tianzhu, et al.
Pubblicazione: (2025)
Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes
di: Fu, Yuqian, et al.
Pubblicazione: (2026)
di: Fu, Yuqian, et al.
Pubblicazione: (2026)
Deciphering the Impact of Pretraining Data on Large Language Models through Machine Unlearning
di: Zhao, Yang, et al.
Pubblicazione: (2024)
di: Zhao, Yang, et al.
Pubblicazione: (2024)
Hybrid Policy Distillation for LLMs
di: Zhu, Wenhong, et al.
Pubblicazione: (2026)
di: Zhu, Wenhong, et al.
Pubblicazione: (2026)
From Next-Token to Next-Block: A Principled Adaptation Path for Diffusion LLMs
di: Tian, Yuchuan, et al.
Pubblicazione: (2025)
di: Tian, Yuchuan, et al.
Pubblicazione: (2025)
SelecTKD: Selective Token-Weighted Knowledge Distillation for LLMs
di: Huang, Haiduo, et al.
Pubblicazione: (2025)
di: Huang, Haiduo, et al.
Pubblicazione: (2025)
OPSDL: On-Policy Self-Distillation for Long-Context Language Models
di: Zhang, Xinsen, et al.
Pubblicazione: (2026)
di: Zhang, Xinsen, et al.
Pubblicazione: (2026)
StyleDecipher: Robust and Explainable Detection of LLM-Generated Texts with Stylistic Analysis
di: Li, Siyuan, et al.
Pubblicazione: (2025)
di: Li, Siyuan, et al.
Pubblicazione: (2025)
A Token is Worth over 1,000 Tokens: Efficient Knowledge Distillation through Low-Rank Clone
di: Hao, Jitai, et al.
Pubblicazione: (2025)
di: Hao, Jitai, et al.
Pubblicazione: (2025)
Transformer-based Joint Modelling for Automatic Essay Scoring and Off-Topic Detection
di: Das, Sourya Dipta, et al.
Pubblicazione: (2024)
di: Das, Sourya Dipta, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Learning How to Use Tools, Not Just When: Pattern-Aware Tool-Integrated Reasoning
di: Xu, Ningning, et al.
Pubblicazione: (2025) -
PromptGuard at BLP-2025 Task 1: A Few-Shot Classification Framework Using Majority Voting and Keyword Similarity for Bengali Hate Speech Detection
di: Hossan, Rakib, et al.
Pubblicazione: (2025) -
UMBCLU at SemEval-2024 Task 1A and 1C: Semantic Textual Relatedness with and without machine translation
di: Dipta, Shubhashis Roy, et al.
Pubblicazione: (2024) -
HU at SemEval-2024 Task 8A: Can Contrastive Learning Learn Embeddings to Detect Machine-Generated Text?
di: Dipta, Shubhashis Roy, et al.
Pubblicazione: (2024) -
BanglaTalk: Towards Real-Time Speech Assistance for Bengali Regional Dialects
di: Hasan, Jakir, et al.
Pubblicazione: (2025)