ParaDySe: A Parallel-Strategy Switching Framework for Dynamic Sequence Lengths in Transformer
Fuente:
arXiv
Salvato in:
| Autori principali: | Ou, Zhixin, Liang, Peng, Han, Jianchen, Liu, Baihui, Qiao, Linbo |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A Survey on Memory-Efficient Transformer-Based Model Training in AI for Science
di: Tian, Kaiyuan, et al.
Pubblicazione: (2025)
di: Tian, Kaiyuan, et al.
Pubblicazione: (2025)
Alloc-MoE: Budget-Aware Expert Activation Allocation for Efficient Mixture-of-Experts Inference
di: Liu, Baihui, et al.
Pubblicazione: (2026)
di: Liu, Baihui, et al.
Pubblicazione: (2026)
Dy-mer: An Explainable DNA Sequence Representation Scheme using Dictionary Learning
di: Peng, Zhiyuan, et al.
Pubblicazione: (2024)
di: Peng, Zhiyuan, et al.
Pubblicazione: (2024)
Budget-aware Auto Optimizer Configurator
di: Liu, Kang, et al.
Pubblicazione: (2026)
di: Liu, Kang, et al.
Pubblicazione: (2026)
Exact Dual Geometry of SOC-ICNN Value Functions
di: Liu, Kang, et al.
Pubblicazione: (2026)
di: Liu, Kang, et al.
Pubblicazione: (2026)
DyMRL: Dynamic Multispace Representation Learning for Multimodal Event Forecasting in Knowledge Graph
di: Zhao, Feng, et al.
Pubblicazione: (2026)
di: Zhao, Feng, et al.
Pubblicazione: (2026)
DyTTP: Trajectory Prediction with Normalization-Free Transformers
di: Zhu, JianLin, et al.
Pubblicazione: (2025)
di: Zhu, JianLin, et al.
Pubblicazione: (2025)
EEG-DCNet: A Fast and Accurate MI-EEG Dilated CNN Classification Method
di: Peng, Wei, et al.
Pubblicazione: (2024)
di: Peng, Wei, et al.
Pubblicazione: (2024)
CoDy: Counterfactual Explainers for Dynamic Graphs
di: Qu, Zhan, et al.
Pubblicazione: (2024)
di: Qu, Zhan, et al.
Pubblicazione: (2024)
Approximated Likelihood Ratio: A Forward-Only and Parallel Framework for Boosting Neural Network Training
di: Zhang, Zeliang, et al.
Pubblicazione: (2024)
di: Zhang, Zeliang, et al.
Pubblicazione: (2024)
Clip Your Sequences Fairly: Enforcing Length Fairness for Sequence-Level RL
di: Mao, Hanyi, et al.
Pubblicazione: (2025)
di: Mao, Hanyi, et al.
Pubblicazione: (2025)
Memory Sequence Length of Data Sampling Impacts the Adaptation of Meta-Reinforcement Learning Agents
di: Zhang, Menglong, et al.
Pubblicazione: (2024)
di: Zhang, Menglong, et al.
Pubblicazione: (2024)
Arithmetic Transformers Can Length-Generalize in Both Operand Length and Count
di: Cho, Hanseul, et al.
Pubblicazione: (2024)
di: Cho, Hanseul, et al.
Pubblicazione: (2024)
Highly Parallelized Reinforcement Learning Training with Relaxed Assignment Dependencies
di: He, Zhouyu, et al.
Pubblicazione: (2025)
di: He, Zhouyu, et al.
Pubblicazione: (2025)
Rethinking the Comparison Unit in Sequence-Level Reinforcement Learning: An Equal-Length Paired Training Framework from Loss Correction to Sample Construction
di: Ding, Fei, et al.
Pubblicazione: (2026)
di: Ding, Fei, et al.
Pubblicazione: (2026)
On Vanishing Variance in Transformer Length Generalization
di: Li, Ruining, et al.
Pubblicazione: (2025)
di: Li, Ruining, et al.
Pubblicazione: (2025)
Adaptive Ensembles of Fine-Tuned Transformers for LLM-Generated Text Detection
di: Lai, Zhixin, et al.
Pubblicazione: (2024)
di: Lai, Zhixin, et al.
Pubblicazione: (2024)
Scheduling Parallel Optical Circuit Switches for AI Training
di: Liang, Kevin, et al.
Pubblicazione: (2026)
di: Liang, Kevin, et al.
Pubblicazione: (2026)
Evaluating the Sensitivity of BiLSTM Forecasting Models to Sequence Length and Input Noise
di: Albelali, Salma, et al.
Pubblicazione: (2025)
di: Albelali, Salma, et al.
Pubblicazione: (2025)
Action-Adaptive Continual Learning: Enabling Policy Generalization under Dynamic Action Spaces
di: Pan, Chaofan, et al.
Pubblicazione: (2025)
di: Pan, Chaofan, et al.
Pubblicazione: (2025)
GSPN-2: Efficient Parallel Sequence Modeling
di: Wang, Hongjun, et al.
Pubblicazione: (2025)
di: Wang, Hongjun, et al.
Pubblicazione: (2025)
DyCE: Dynamically Configurable Exiting for Deep Learning Compression and Real-time Scaling
di: Wang, Qingyuan, et al.
Pubblicazione: (2024)
di: Wang, Qingyuan, et al.
Pubblicazione: (2024)
DySCo: Dynamic Semantic Compression for Effective Long-term Time Series Forecasting
di: Ao, Xiang, et al.
Pubblicazione: (2026)
di: Ao, Xiang, et al.
Pubblicazione: (2026)
Adaptive Overclocking: Dynamic Control of Thinking Path Length via Real-Time Reasoning Signals
di: Jiang, Shuhao, et al.
Pubblicazione: (2025)
di: Jiang, Shuhao, et al.
Pubblicazione: (2025)
How Particle-System Random Batch Methods Enhance Graph Transformer: Memory Efficiency and Parallel Computing Strategy
di: Liu, Hanwen, et al.
Pubblicazione: (2025)
di: Liu, Hanwen, et al.
Pubblicazione: (2025)
In Search of Lost DNA Sequence Pretraining
di: Tang, Zhijiang, et al.
Pubblicazione: (2026)
di: Tang, Zhijiang, et al.
Pubblicazione: (2026)
Acceleration for Deep Reinforcement Learning using Parallel and Distributed Computing: A Survey
di: Liu, Zhihong, et al.
Pubblicazione: (2024)
di: Liu, Zhihong, et al.
Pubblicazione: (2024)
Transolver is a Linear Transformer: Revisiting Physics-Attention through the Lens of Linear Attention
di: Hu, Wenjie, et al.
Pubblicazione: (2025)
di: Hu, Wenjie, et al.
Pubblicazione: (2025)
Provable Length Generalization in Sequence Prediction via Spectral Filtering
di: Marsden, Annie, et al.
Pubblicazione: (2024)
di: Marsden, Annie, et al.
Pubblicazione: (2024)
USP: A Unified Sequence Parallelism Approach for Long Context Generative AI
di: Fang, Jiarui, et al.
Pubblicazione: (2024)
di: Fang, Jiarui, et al.
Pubblicazione: (2024)
AGDC: Autoregressive Generation of Variable-Length Sequences with Joint Discrete and Continuous Spaces
di: Shin, Yeonsang, et al.
Pubblicazione: (2026)
di: Shin, Yeonsang, et al.
Pubblicazione: (2026)
Rethinking Random Transformers as Adaptive Sequence Smoothers for Sleep Staging
di: Liu, Guisong, et al.
Pubblicazione: (2026)
di: Liu, Guisong, et al.
Pubblicazione: (2026)
The Role of Sparsity for Length Generalization in Transformers
di: Golowich, Noah, et al.
Pubblicazione: (2025)
di: Golowich, Noah, et al.
Pubblicazione: (2025)
On the Limitations and Capabilities of Position Embeddings for Length Generalization
di: Chen, Yang, et al.
Pubblicazione: (2025)
di: Chen, Yang, et al.
Pubblicazione: (2025)
EasyQuant: An Efficient Data-free Quantization Algorithm for LLMs
di: Tang, Hanlin, et al.
Pubblicazione: (2024)
di: Tang, Hanlin, et al.
Pubblicazione: (2024)
DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks
di: Zhu, Kaijie, et al.
Pubblicazione: (2023)
di: Zhu, Kaijie, et al.
Pubblicazione: (2023)
DyEdgeGAT: Dynamic Edge via Graph Attention for Early Fault Detection in IIoT Systems
di: Zhao, Mengjie, et al.
Pubblicazione: (2023)
di: Zhao, Mengjie, et al.
Pubblicazione: (2023)
DyJR: Preserving Diversity in Reinforcement Learning with Verifiable Rewards via Dynamic Jensen-Shannon Replay
di: Li, Long, et al.
Pubblicazione: (2026)
di: Li, Long, et al.
Pubblicazione: (2026)
Energy-Efficient Stochastic Computing (SC) Neural Networks for Internet of Things Devices With Layer-Wise Adjustable Sequence Length (ASL)
di: Wang, Ziheng, et al.
Pubblicazione: (2025)
di: Wang, Ziheng, et al.
Pubblicazione: (2025)
GFT: From Imitation to Reward Fine-Tuning with Unbiased Group Advantages and Dynamic Coefficient Rectification
di: Gan, Wangjie, et al.
Pubblicazione: (2026)
di: Gan, Wangjie, et al.
Pubblicazione: (2026)
Documenti analoghi
-
A Survey on Memory-Efficient Transformer-Based Model Training in AI for Science
di: Tian, Kaiyuan, et al.
Pubblicazione: (2025) -
Alloc-MoE: Budget-Aware Expert Activation Allocation for Efficient Mixture-of-Experts Inference
di: Liu, Baihui, et al.
Pubblicazione: (2026) -
Dy-mer: An Explainable DNA Sequence Representation Scheme using Dictionary Learning
di: Peng, Zhiyuan, et al.
Pubblicazione: (2024) -
Budget-aware Auto Optimizer Configurator
di: Liu, Kang, et al.
Pubblicazione: (2026) -
Exact Dual Geometry of SOC-ICNN Value Functions
di: Liu, Kang, et al.
Pubblicazione: (2026)