SeqPE: Transformer with Sequential Position Encoding
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Huayang, Liu, Yahui, Sun, Hongyu, Cai, Deng, Cui, Leyang, Bi, Wei, Zhao, Peilin, Watanabe, Taro |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mitigating Coordinate Prediction Bias from Positional Encoding Failures
by: Tao, Xingjian, et al.
Published: (2025)
by: Tao, Xingjian, et al.
Published: (2025)
Inferflow: an Efficient and Highly Configurable Inference Engine for Large Language Models
by: Shi, Shuming, et al.
Published: (2024)
by: Shi, Shuming, et al.
Published: (2024)
ID-Align: RoPE-Conscious Position Remapping for Dynamic High-Resolution Adaptation in Vision-Language Models
by: Li, Bozhou, et al.
Published: (2025)
by: Li, Bozhou, et al.
Published: (2025)
Set2Seq Transformer: Temporal and Position-Aware Set Representations for Sequential Multiple-Instance Learning
by: Efthymiou, Athanasios, et al.
Published: (2024)
by: Efthymiou, Athanasios, et al.
Published: (2024)
Diagnosing Vision Language Models' Perception by Leveraging Human Methods for Color Vision Deficiencies
by: Hayashi, Kazuki, et al.
Published: (2025)
by: Hayashi, Kazuki, et al.
Published: (2025)
Constructing Multilingual Visual-Text Datasets Revealing Visual Multilingual Ability of Vision Language Models
by: Atuhurra, Jesse, et al.
Published: (2024)
by: Atuhurra, Jesse, et al.
Published: (2024)
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding
by: Chen, Zhanpeng, et al.
Published: (2025)
by: Chen, Zhanpeng, et al.
Published: (2025)
Seq2Time: Sequential Knowledge Transfer for Video LLM Temporal Grounding
by: Deng, Andong, et al.
Published: (2024)
by: Deng, Andong, et al.
Published: (2024)
Forest2Seq: Revitalizing Order Prior for Sequential Indoor Scene Synthesis
by: Sun, Qi, et al.
Published: (2024)
by: Sun, Qi, et al.
Published: (2024)
LaneGraph2Seq: Lane Topology Extraction with Language Model via Vertex-Edge Encoding and Connectivity Enhancement
by: Peng, Renyuan, et al.
Published: (2024)
by: Peng, Renyuan, et al.
Published: (2024)
When Diffusion Breaks Constraints: Sequential Autoregressive Generation with RL and MCTS
by: Zhao, Zirui, et al.
Published: (2025)
by: Zhao, Zirui, et al.
Published: (2025)
Beyond Sequential Distance: Inter-Modal Distance Invariant Position Encoding
by: Chen, Lin, et al.
Published: (2026)
by: Chen, Lin, et al.
Published: (2026)
Cross-lingual Contextualized Phrase Retrieval
by: Li, Huayang, et al.
Published: (2024)
by: Li, Huayang, et al.
Published: (2024)
SCoPE VLM: Selective Context Processing for Efficient Document Navigation in Vision-Language Models
by: Lim, Gyubeum, et al.
Published: (2025)
by: Lim, Gyubeum, et al.
Published: (2025)
Skywork-R1V3 Technical Report
by: Shen, Wei, et al.
Published: (2025)
by: Shen, Wei, et al.
Published: (2025)
Advanced Image Segmentation Techniques for Neural Activity Detection via C-fos Immediate Early Gene Expression
by: Cai, Peilin
Published: (2023)
by: Cai, Peilin
Published: (2023)
The Hard Positive Truth about Vision-Language Compositionality
by: Kamath, Amita, et al.
Published: (2024)
by: Kamath, Amita, et al.
Published: (2024)
v-HUB: A Benchmark for Video Humor Understanding from Vision and Sound
by: Shi, Zhengpeng, et al.
Published: (2025)
by: Shi, Zhengpeng, et al.
Published: (2025)
RayRoPE: Projective Ray Positional Encoding for Multi-view Attention
by: Wu, Yu, et al.
Published: (2026)
by: Wu, Yu, et al.
Published: (2026)
TextTIGER: Text-based Intelligent Generation with Entity Prompt Refinement for Text-to-Image Generation
by: Ozaki, Shintaro, et al.
Published: (2025)
by: Ozaki, Shintaro, et al.
Published: (2025)
Octavius: Mitigating Task Interference in MLLMs via LoRA-MoE
by: Chen, Zeren, et al.
Published: (2023)
by: Chen, Zeren, et al.
Published: (2023)
TRACE: Textual Relevance Augmentation and Contextual Encoding for Multimodal Hate Detection
by: Koushik, Girish A., et al.
Published: (2025)
by: Koushik, Girish A., et al.
Published: (2025)
CLoVe: Encoding Compositional Language in Contrastive Vision-Language Models
by: Castro, Santiago, et al.
Published: (2024)
by: Castro, Santiago, et al.
Published: (2024)
Inference-Time Alignment of Diffusion Models via Trust-Region Iterative Twisted Sequential Monte Carlo
by: Wang, Weixin, et al.
Published: (2026)
by: Wang, Weixin, et al.
Published: (2026)
Toward Automatic Safe Driving Instruction: A Large-Scale Vision Language Model Approach
by: Sakajo, Haruki, et al.
Published: (2025)
by: Sakajo, Haruki, et al.
Published: (2025)
From Backward Spreading to Forward Replay: Revisiting Target Construction in LLM Parameter Editing
by: Liu, Wei, et al.
Published: (2026)
by: Liu, Wei, et al.
Published: (2026)
SeqCSIST: Sequential Closely-Spaced Infrared Small Target Unmixing
by: Zhai, Ximeng, et al.
Published: (2025)
by: Zhai, Ximeng, et al.
Published: (2025)
What Do Vision-Language Models Encode for Personalized Image Aesthetics Assessment?
by: Ryu, Koki, et al.
Published: (2026)
by: Ryu, Koki, et al.
Published: (2026)
LEDiT: Your Length-Extrapolatable Diffusion Transformer without Positional Encoding
by: Zhang, Shen, et al.
Published: (2025)
by: Zhang, Shen, et al.
Published: (2025)
Unified Camera Positional Encoding for Controlled Video Generation
by: Zhang, Cheng, et al.
Published: (2025)
by: Zhang, Cheng, et al.
Published: (2025)
On the Transformations across Reward Model, Parameter Update, and In-Context Prompt
by: Cai, Deng, et al.
Published: (2024)
by: Cai, Deng, et al.
Published: (2024)
Uni-SMART: Universal Science Multimodal Analysis and Research Transformer
by: Cai, Hengxing, et al.
Published: (2024)
by: Cai, Hengxing, et al.
Published: (2024)
Global-Local Tree Search in VLMs for 3D Indoor Scene Generation
by: Deng, Wei, et al.
Published: (2025)
by: Deng, Wei, et al.
Published: (2025)
SAISA: Towards Multimodal Large Language Models with Both Training and Inference Efficiency
by: Yuan, Qianhao, et al.
Published: (2025)
by: Yuan, Qianhao, et al.
Published: (2025)
Attention Grounded Enhancement for Visual Document Retrieval
by: Cui, Wanqing, et al.
Published: (2025)
by: Cui, Wanqing, et al.
Published: (2025)
Weierstrass Positional Encoding for Vision Transformers
by: Xin, Zhihang, et al.
Published: (2026)
by: Xin, Zhihang, et al.
Published: (2026)
History-Aware Reasoning for GUI Agents
by: Wang, Ziwei, et al.
Published: (2025)
by: Wang, Ziwei, et al.
Published: (2025)
MaVEn: An Effective Multi-granularity Hybrid Visual Encoding Framework for Multimodal Large Language Model
by: Jiang, Chaoya, et al.
Published: (2024)
by: Jiang, Chaoya, et al.
Published: (2024)
Exploiting GPT-4 Vision for Zero-shot Point Cloud Understanding
by: Sun, Qi, et al.
Published: (2024)
by: Sun, Qi, et al.
Published: (2024)
NeuralOS: Towards Simulating Operating Systems via Neural Generative Models
by: Rivard, Luke, et al.
Published: (2025)
by: Rivard, Luke, et al.
Published: (2025)
Similar Items
-
Mitigating Coordinate Prediction Bias from Positional Encoding Failures
by: Tao, Xingjian, et al.
Published: (2025) -
Inferflow: an Efficient and Highly Configurable Inference Engine for Large Language Models
by: Shi, Shuming, et al.
Published: (2024) -
ID-Align: RoPE-Conscious Position Remapping for Dynamic High-Resolution Adaptation in Vision-Language Models
by: Li, Bozhou, et al.
Published: (2025) -
Set2Seq Transformer: Temporal and Position-Aware Set Representations for Sequential Multiple-Instance Learning
by: Efthymiou, Athanasios, et al.
Published: (2024) -
Diagnosing Vision Language Models' Perception by Leveraging Human Methods for Color Vision Deficiencies
by: Hayashi, Kazuki, et al.
Published: (2025)