Switch Attention: Towards Dynamic and Fine-grained Hybrid Transformers
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Zhao, Yusheng, Li, Hourun, Wu, Bohan, Yin, Yichun, Shang, Lifeng, Yuan, Jingyang, Zhang, Meng, Zhang, Ming |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Discover and Prove: An Open-source Agentic Framework for Hard Mode Automated Theorem Proving in Lean 4
par: Liu, Chengwu, et autres
Publié: (2026)
par: Liu, Chengwu, et autres
Publié: (2026)
MMEvalPro: Calibrating Multimodal Benchmarks Towards Trustworthy and Efficient Evaluation
par: Huang, Jinsheng, et autres
Publié: (2024)
par: Huang, Jinsheng, et autres
Publié: (2024)
A Hybrid RAG System with Comprehensive Enhancement on Complex Reasoning
par: Yuan, Ye, et autres
Publié: (2024)
par: Yuan, Ye, et autres
Publié: (2024)
G-MAP: General Memory-Augmented Pre-trained Language Model for Domain Tasks
par: Wan, Zhongwei, et autres
Publié: (2022)
par: Wan, Zhongwei, et autres
Publié: (2022)
Improving Transformers with Dynamically Composable Multi-Head Attention
par: Xiao, Da, et autres
Publié: (2024)
par: Xiao, Da, et autres
Publié: (2024)
SDMPrune: Self-Distillation MLP Pruning for Efficient Large Language Models
par: Zhu, Hourun, et autres
Publié: (2025)
par: Zhu, Hourun, et autres
Publié: (2025)
Safe: Enhancing Mathematical Reasoning in Large Language Models via Retrospective Step-aware Formal Verification
par: Liu, Chengwu, et autres
Publié: (2025)
par: Liu, Chengwu, et autres
Publié: (2025)
FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models
par: Jiang, Yuxin, et autres
Publié: (2023)
par: Jiang, Yuxin, et autres
Publié: (2023)
The Harder The Better: Maintaining Supervised Fine-tuning Generalization with Less but Harder Data
par: Shang, Zhaoyang, et autres
Publié: (2025)
par: Shang, Zhaoyang, et autres
Publié: (2025)
RobustFT: Robust Supervised Fine-tuning for Large Language Models under Noisy Response
par: Luo, Junyu, et autres
Publié: (2024)
par: Luo, Junyu, et autres
Publié: (2024)
A Survey on Efficient Large Language Model Training: From Data-centric Perspectives
par: Luo, Junyu, et autres
Publié: (2025)
par: Luo, Junyu, et autres
Publié: (2025)
Entropy-Guided Reasoning Compression
par: Zhu, Hourun, et autres
Publié: (2025)
par: Zhu, Hourun, et autres
Publié: (2025)
Gaps or Hallucinations? Gazing into Machine-Generated Legal Analysis for Fine-grained Text Evaluations
par: Hou, Abe Bohan, et autres
Publié: (2024)
par: Hou, Abe Bohan, et autres
Publié: (2024)
FLAT-LLM: Fine-grained Low-rank Activation Space Transformation for Large Language Model Compression
par: Tian, Jiayi, et autres
Publié: (2025)
par: Tian, Jiayi, et autres
Publié: (2025)
Fovea Transformer: Efficient Long-Context Modeling with Structured Fine-to-Coarse Attention
par: He, Ziwei, et autres
Publié: (2023)
par: He, Ziwei, et autres
Publié: (2023)
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression
par: Zhang, Jiebin, et autres
Publié: (2024)
par: Zhang, Jiebin, et autres
Publié: (2024)
CAT: Causal Attention Tuning For Injecting Fine-grained Causal Knowledge into Large Language Models
par: Han, Kairong, et autres
Publié: (2025)
par: Han, Kairong, et autres
Publié: (2025)
Repurposing Synthetic Data for Fine-grained Search Agent Supervision
par: Zhao, Yida, et autres
Publié: (2025)
par: Zhao, Yida, et autres
Publié: (2025)
Large Language Model Agent: A Survey on Methodology, Applications and Challenges
par: Luo, Junyu, et autres
Publié: (2025)
par: Luo, Junyu, et autres
Publié: (2025)
ResoFilter: Fine-grained Synthetic Data Filtering for Large Language Models through Data-Parameter Resonance Analysis
par: Tu, Zeao, et autres
Publié: (2024)
par: Tu, Zeao, et autres
Publié: (2024)
MedPRMBench: A Fine-grained Benchmark for Process Reward Models in Medical Reasoning
par: Wu, Lingyan, et autres
Publié: (2026)
par: Wu, Lingyan, et autres
Publié: (2026)
UltraGen: Extremely Fine-grained Controllable Generation via Attribute Reconstruction and Global Preference Optimization
par: Yun, Longfei, et autres
Publié: (2025)
par: Yun, Longfei, et autres
Publié: (2025)
SwitchHead: Accelerating Transformers with Mixture-of-Experts Attention
par: Csordás, Róbert, et autres
Publié: (2023)
par: Csordás, Róbert, et autres
Publié: (2023)
KPEval: Towards Fine-Grained Semantic-Based Keyphrase Evaluation
par: Wu, Di, et autres
Publié: (2023)
par: Wu, Di, et autres
Publié: (2023)
EFUF: Efficient Fine-grained Unlearning Framework for Mitigating Hallucinations in Multimodal Large Language Models
par: Xing, Shangyu, et autres
Publié: (2024)
par: Xing, Shangyu, et autres
Publié: (2024)
Fine-grained and Explainable Factuality Evaluation for Multimodal Summarization
par: Zhang, Yue, et autres
Publié: (2024)
par: Zhang, Yue, et autres
Publié: (2024)
DynMoLE: Boosting Mixture of LoRA Experts Fine-Tuning with a Hybrid Routing Mechanism
par: Li, Dengchun, et autres
Publié: (2025)
par: Li, Dengchun, et autres
Publié: (2025)
CANTONMT: Investigating Back-Translation and Model-Switch Mechanisms for Cantonese-English Neural Machine Translation
par: Hong, Kung Yin, et autres
Publié: (2024)
par: Hong, Kung Yin, et autres
Publié: (2024)
MolReFlect: Towards In-Context Fine-grained Alignments between Molecules and Texts
par: Li, Jiatong, et autres
Publié: (2024)
par: Li, Jiatong, et autres
Publié: (2024)
Towards Fine-Grained Code-Switch Speech Translation with Semantic Space Alignment
par: Gao, Yan, et autres
Publié: (2025)
par: Gao, Yan, et autres
Publié: (2025)
Improving Large Language Models via Fine-grained Reinforcement Learning with Minimum Editing Constraint
par: Chen, Zhipeng, et autres
Publié: (2024)
par: Chen, Zhipeng, et autres
Publié: (2024)
TTPA: Token-level Tool-use Preference Alignment Training Framework with Fine-grained Evaluation
par: Huang, Chengrui, et autres
Publié: (2025)
par: Huang, Chengrui, et autres
Publié: (2025)
ME-Switch: A Memory-Efficient Expert Switching Framework for Large Language Models
par: Liu, Jing, et autres
Publié: (2024)
par: Liu, Jing, et autres
Publié: (2024)
MVAM: Multi-View Attention Method for Fine-grained Image-Text Matching
par: Cui, Wanqing, et autres
Publié: (2024)
par: Cui, Wanqing, et autres
Publié: (2024)
Lost in the Mix: Evaluating LLM Understanding of Code-Switched Text
par: Mohamed, Amr, et autres
Publié: (2025)
par: Mohamed, Amr, et autres
Publié: (2025)
From Chaos to Order: The Atomic Reasoner Framework for Fine-grained Reasoning in Large Language Models
par: Liu, Jinyi, et autres
Publié: (2025)
par: Liu, Jinyi, et autres
Publié: (2025)
Beyond Imitation: Leveraging Fine-grained Quality Signals for Alignment
par: Guo, Geyang, et autres
Publié: (2023)
par: Guo, Geyang, et autres
Publié: (2023)
RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-grained Correctional Human Feedback
par: Yu, Tianyu, et autres
Publié: (2023)
par: Yu, Tianyu, et autres
Publié: (2023)
ChiMDQA: Towards Comprehensive Chinese Document QA with Fine-grained Evaluation
par: Gao, Jing, et autres
Publié: (2025)
par: Gao, Jing, et autres
Publié: (2025)
Mask-DPO: Generalizable Fine-grained Factuality Alignment of LLMs
par: Gu, Yuzhe, et autres
Publié: (2025)
par: Gu, Yuzhe, et autres
Publié: (2025)
Documents similaires
-
Discover and Prove: An Open-source Agentic Framework for Hard Mode Automated Theorem Proving in Lean 4
par: Liu, Chengwu, et autres
Publié: (2026) -
MMEvalPro: Calibrating Multimodal Benchmarks Towards Trustworthy and Efficient Evaluation
par: Huang, Jinsheng, et autres
Publié: (2024) -
A Hybrid RAG System with Comprehensive Enhancement on Complex Reasoning
par: Yuan, Ye, et autres
Publié: (2024) -
G-MAP: General Memory-Augmented Pre-trained Language Model for Domain Tasks
par: Wan, Zhongwei, et autres
Publié: (2022) -
Improving Transformers with Dynamically Composable Multi-Head Attention
par: Xiao, Da, et autres
Publié: (2024)