TAD: Temporal-Aware Trajectory Self-Distillation for Fast and Accurate Diffusion LLM
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhou, Haoyang, Kong, Li, Ren, Shijie, Wang, Xiting, Liang, Shuang, Wang, Guowei, Pan, Zhenxuan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents
von: Wang, Hao, et al.
Veröffentlicht: (2026)
von: Wang, Hao, et al.
Veröffentlicht: (2026)
$A^3$: Attention-Aware Accurate KV Cache Fusion for Fast Large Language Model Serving
von: Zhou, Yuechi, et al.
Veröffentlicht: (2025)
von: Zhou, Yuechi, et al.
Veröffentlicht: (2025)
Enhancing Safety of Large Language Models via Embedding Space Separation
von: Zhao, Xu, et al.
Veröffentlicht: (2026)
von: Zhao, Xu, et al.
Veröffentlicht: (2026)
Less Is More: Fast and Accurate Reasoning with Cross-Head Unified Sparse Attention
von: Yang, Lijie, et al.
Veröffentlicht: (2025)
von: Yang, Lijie, et al.
Veröffentlicht: (2025)
Temporal-Aware Heterogeneous Graph Reasoning with Multi-View Fusion for Temporal Question Answering
von: Wen, Wuzhenghong, et al.
Veröffentlicht: (2026)
von: Wen, Wuzhenghong, et al.
Veröffentlicht: (2026)
EchoDistill:Alignment Noisy-to-Clean Self-Distillation for Robust Audio LLMs
von: Lin, Liang, et al.
Veröffentlicht: (2026)
von: Lin, Liang, et al.
Veröffentlicht: (2026)
Ultra-Fast Language Generation via Discrete Diffusion Divergence Instruct
von: Zheng, Haoyang, et al.
Veröffentlicht: (2025)
von: Zheng, Haoyang, et al.
Veröffentlicht: (2025)
Skill-Conditioned Gated Self-Distillation for LLM Reasoning
von: Huang, Jiazhen, et al.
Veröffentlicht: (2026)
von: Huang, Jiazhen, et al.
Veröffentlicht: (2026)
Distilling Multi-Scale Knowledge for Event Temporal Relation Extraction
von: Yao, Hao-Ren, et al.
Veröffentlicht: (2022)
von: Yao, Hao-Ren, et al.
Veröffentlicht: (2022)
TAD-Bench: A Comprehensive Benchmark for Embedding-Based Text Anomaly Detection
von: Cao, Yang, et al.
Veröffentlicht: (2025)
von: Cao, Yang, et al.
Veröffentlicht: (2025)
Self-signals Driven Multi-LLM Debate for Efficient and Accurate Reasoning
von: Chen, Xuhang, et al.
Veröffentlicht: (2025)
von: Chen, Xuhang, et al.
Veröffentlicht: (2025)
AgentHER: Hindsight Experience Replay for LLM Agent Trajectory Relabeling
von: Ding, Liang
Veröffentlicht: (2026)
von: Ding, Liang
Veröffentlicht: (2026)
TABES: Trajectory-Aware Backward-on-Entropy Steering for Masked Diffusion Models
von: Saini, Shreshth, et al.
Veröffentlicht: (2026)
von: Saini, Shreshth, et al.
Veröffentlicht: (2026)
SkillAdaptor: Self-Adapting Skills for LLM Agents from Trajectories
von: Yu, Zhuoyun, et al.
Veröffentlicht: (2026)
von: Yu, Zhuoyun, et al.
Veröffentlicht: (2026)
Pride and Prejudice: LLM Amplifies Self-Bias in Self-Refinement
von: Xu, Wenda, et al.
Veröffentlicht: (2024)
von: Xu, Wenda, et al.
Veröffentlicht: (2024)
GRASP: Gated Regression-Aware Skill Proposer for Self-Improving LLM Agents
von: Moll, Johannes, et al.
Veröffentlicht: (2026)
von: Moll, Johannes, et al.
Veröffentlicht: (2026)
Entropy-based Exploration Conduction for Multi-step Reasoning
von: Zhang, Jinghan, et al.
Veröffentlicht: (2025)
von: Zhang, Jinghan, et al.
Veröffentlicht: (2025)
STaRR: Spatial-Temporal Token-Dynamics-Aware Responsive Remasking for Diffusion Language Models
von: Sun, Xinhao, et al.
Veröffentlicht: (2025)
von: Sun, Xinhao, et al.
Veröffentlicht: (2025)
Harmony in Divergence: Towards Fast, Accurate, and Memory-efficient Zeroth-order LLM Fine-tuning
von: Tan, Qitao, et al.
Veröffentlicht: (2025)
von: Tan, Qitao, et al.
Veröffentlicht: (2025)
SkillFactory: Self-Distillation For Learning Cognitive Behaviors
von: Sprague, Zayne, et al.
Veröffentlicht: (2025)
von: Sprague, Zayne, et al.
Veröffentlicht: (2025)
d3LLM: Ultra-Fast Diffusion LLM using Pseudo-Trajectory Distillation
von: Qian, Yu-Yang, et al.
Veröffentlicht: (2026)
von: Qian, Yu-Yang, et al.
Veröffentlicht: (2026)
O1 Replication Journey -- Part 2: Surpassing O1-preview through Simple Distillation, Big Progress or Bitter Lesson?
von: Huang, Zhen, et al.
Veröffentlicht: (2024)
von: Huang, Zhen, et al.
Veröffentlicht: (2024)
FASTTRACK: Fast and Accurate Fact Tracing for LLMs
von: Chen, Si, et al.
Veröffentlicht: (2024)
von: Chen, Si, et al.
Veröffentlicht: (2024)
Building Accurate Translation-Tailored LLMs with Language Aware Instruction Tuning
von: Zan, Changtong, et al.
Veröffentlicht: (2024)
von: Zan, Changtong, et al.
Veröffentlicht: (2024)
TidalDecode: Fast and Accurate LLM Decoding with Position Persistent Sparse Attention
von: Yang, Lijie, et al.
Veröffentlicht: (2024)
von: Yang, Lijie, et al.
Veröffentlicht: (2024)
Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL
von: Liu, Che, et al.
Veröffentlicht: (2025)
von: Liu, Che, et al.
Veröffentlicht: (2025)
Thinking Slow, Fast: Scaling Inference Compute with Distilled Reasoners
von: Paliotta, Daniele, et al.
Veröffentlicht: (2025)
von: Paliotta, Daniele, et al.
Veröffentlicht: (2025)
LLM-Oriented Token-Adaptive Knowledge Distillation
von: Xie, Xurong, et al.
Veröffentlicht: (2025)
von: Xie, Xurong, et al.
Veröffentlicht: (2025)
Extreme Region Policy Distillation
von: Chen, Changyu, et al.
Veröffentlicht: (2026)
von: Chen, Changyu, et al.
Veröffentlicht: (2026)
OPSDL: On-Policy Self-Distillation for Long-Context Language Models
von: Zhang, Xinsen, et al.
Veröffentlicht: (2026)
von: Zhang, Xinsen, et al.
Veröffentlicht: (2026)
LLM-MRD: LLM-Guided Multi-View Reasoning Distillation for Fake News Detection
von: Zhou, Weilin, et al.
Veröffentlicht: (2026)
von: Zhou, Weilin, et al.
Veröffentlicht: (2026)
GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation
von: Li, Sijia, et al.
Veröffentlicht: (2026)
von: Li, Sijia, et al.
Veröffentlicht: (2026)
Distilling Closed-Source LLM's Knowledge for Locally Stable and Economic Biomedical Entity Linking
von: Ai, Yihao, et al.
Veröffentlicht: (2025)
von: Ai, Yihao, et al.
Veröffentlicht: (2025)
DB-LLM: Accurate Dual-Binarization for Efficient LLMs
von: Chen, Hong, et al.
Veröffentlicht: (2024)
von: Chen, Hong, et al.
Veröffentlicht: (2024)
StruEdit: Structured Outputs Enable the Fast and Accurate Knowledge Editing for Large Language Models
von: Bi, Baolong, et al.
Veröffentlicht: (2024)
von: Bi, Baolong, et al.
Veröffentlicht: (2024)
"The Whole Is Greater Than the Sum of Its Parts": A Compatibility-Aware Multi-Teacher CoT Distillation Framework
von: Cui, Jin, et al.
Veröffentlicht: (2026)
von: Cui, Jin, et al.
Veröffentlicht: (2026)
GUARDIAN: Safeguarding LLM Multi-Agent Collaborations with Temporal Graph Modeling
von: Zhou, Jialong, et al.
Veröffentlicht: (2025)
von: Zhou, Jialong, et al.
Veröffentlicht: (2025)
Fast and Accurate Factual Inconsistency Detection Over Long Documents
von: Lattimer, Barrett Martin, et al.
Veröffentlicht: (2023)
von: Lattimer, Barrett Martin, et al.
Veröffentlicht: (2023)
Self-Correction Distillation for Structured Data Question Answering
von: Zhu, Yushan, et al.
Veröffentlicht: (2025)
von: Zhu, Yushan, et al.
Veröffentlicht: (2025)
Confidence-aware Self-Semantic Distillation on Knowledge Graph Embedding
von: Liu, Yichen, et al.
Veröffentlicht: (2022)
von: Liu, Yichen, et al.
Veröffentlicht: (2022)
Ähnliche Einträge
-
Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents
von: Wang, Hao, et al.
Veröffentlicht: (2026) -
$A^3$: Attention-Aware Accurate KV Cache Fusion for Fast Large Language Model Serving
von: Zhou, Yuechi, et al.
Veröffentlicht: (2025) -
Enhancing Safety of Large Language Models via Embedding Space Separation
von: Zhao, Xu, et al.
Veröffentlicht: (2026) -
Less Is More: Fast and Accurate Reasoning with Cross-Head Unified Sparse Attention
von: Yang, Lijie, et al.
Veröffentlicht: (2025) -
Temporal-Aware Heterogeneous Graph Reasoning with Multi-View Fusion for Temporal Question Answering
von: Wen, Wuzhenghong, et al.
Veröffentlicht: (2026)