ID-Align: RoPE-Conscious Position Remapping for Dynamic High-Resolution Adaptation in Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Bozhou, Zhang, Wentao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ExtraVAR: Stage-Aware RoPE Remapping for Resolution Extrapolation in Visual Autoregressive Models
by: Yan, Feihong, et al.
Published: (2026)
by: Yan, Feihong, et al.
Published: (2026)
ReRoPE: Repurposing RoPE for Relative Camera Control
by: Li, Chunyang, et al.
Published: (2026)
by: Li, Chunyang, et al.
Published: (2026)
Are Bigger Encoders Always Better in Vision Large Models?
by: Li, Bozhou, et al.
Published: (2024)
by: Li, Bozhou, et al.
Published: (2024)
SynthVLM: Towards High-Quality and Efficient Synthesis of Image-Caption Datasets for Vision-Language Models
by: Liu, Zheng, et al.
Published: (2024)
by: Liu, Zheng, et al.
Published: (2024)
Spiral RoPE: Rotate Your Rotary Positional Embeddings in the 2D Plane
by: Liu, Haoyu, et al.
Published: (2026)
by: Liu, Haoyu, et al.
Published: (2026)
Diffusion MRI Transformer with a Diffusion Space Rotary Positional Embedding (D-RoPE)
by: Kung, Gustavo Chau Loo, et al.
Published: (2026)
by: Kung, Gustavo Chau Loo, et al.
Published: (2026)
A Circular Argument : Does RoPE need to be Equivariant for Vision?
by: van de Geijn, Chase, et al.
Published: (2025)
by: van de Geijn, Chase, et al.
Published: (2025)
Untwisting RoPE: Frequency Control for Shared Attention in DiTs
by: Mikaeili, Aryan, et al.
Published: (2026)
by: Mikaeili, Aryan, et al.
Published: (2026)
Infinity-RoPE: Action-Controllable Infinite Video Generation Emerges From Autoregressive Self-Rollout
by: Yesiltepe, Hidir, et al.
Published: (2025)
by: Yesiltepe, Hidir, et al.
Published: (2025)
RoPeSLR: 3D RoPE-driven Sparse-LowRank Attention for Efficient Diffusion Transformers
by: Liu, Yuxi, et al.
Published: (2026)
by: Liu, Yuxi, et al.
Published: (2026)
FreeFlux: Understanding and Exploiting Layer-Specific Roles in RoPE-Based MMDiT for Versatile Image Editing
by: Wei, Tianyi, et al.
Published: (2025)
by: Wei, Tianyi, et al.
Published: (2025)
Anchor Forcing: Anchor Memory and Tri-Region RoPE for Interactive Streaming Video Diffusion
by: Yang, Yang, et al.
Published: (2026)
by: Yang, Yang, et al.
Published: (2026)
Conflict Adaptation in Vision-Language Models
by: Hu, Xiaoyang
Published: (2025)
by: Hu, Xiaoyang
Published: (2025)
MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing
by: Niu, Junbo, et al.
Published: (2025)
by: Niu, Junbo, et al.
Published: (2025)
RoPECraft: Training-Free Motion Transfer with Trajectory-Guided RoPE Optimization on Diffusion Transformers
by: Gokmen, Ahmet Berke, et al.
Published: (2025)
by: Gokmen, Ahmet Berke, et al.
Published: (2025)
How Does Vision-Language Adaptation Impact the Safety of Vision Language Models?
by: Lee, Seongyun, et al.
Published: (2024)
by: Lee, Seongyun, et al.
Published: (2024)
Instruction-Aligned Visual Attention for Mitigating Hallucinations in Large Vision-Language Models
by: Li, Bin, et al.
Published: (2025)
by: Li, Bin, et al.
Published: (2025)
HSCR: Hierarchical Self-Contrastive Rewarding for Aligning Medical Vision Language Models
by: Jiang, Songtao, et al.
Published: (2025)
by: Jiang, Songtao, et al.
Published: (2025)
AlignMMBench: Evaluating Chinese Multimodal Alignment in Large Vision-Language Models
by: Wu, Yuhang, et al.
Published: (2024)
by: Wu, Yuhang, et al.
Published: (2024)
SCoPE VLM: Selective Context Processing for Efficient Document Navigation in Vision-Language Models
by: Lim, Gyubeum, et al.
Published: (2025)
by: Lim, Gyubeum, et al.
Published: (2025)
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
by: Fan, Zhiwen, et al.
Published: (2025)
by: Fan, Zhiwen, et al.
Published: (2025)
DyPE: Dynamic Position Extrapolation for Ultra High Resolution Diffusion
by: Issachar, Noam, et al.
Published: (2025)
by: Issachar, Noam, et al.
Published: (2025)
ETA: Evaluating Then Aligning Safety of Vision Language Models at Inference Time
by: Ding, Yi, et al.
Published: (2024)
by: Ding, Yi, et al.
Published: (2024)
TASO: Task-Aligned Sparse Optimization for Parameter-Efficient Model Adaptation
by: Miao, Daiye, et al.
Published: (2025)
by: Miao, Daiye, et al.
Published: (2025)
FGAIF: Aligning Large Vision-Language Models with Fine-grained AI Feedback
by: Jing, Liqiang, et al.
Published: (2024)
by: Jing, Liqiang, et al.
Published: (2024)
DaLPSR: Leverage Degradation-Aligned Language Prompt for Real-World Image Super-Resolution
by: Jiang, Aiwen, et al.
Published: (2024)
by: Jiang, Aiwen, et al.
Published: (2024)
Enhancing Vision-Language Few-Shot Adaptation with Negative Learning
by: Zhang, Ce, et al.
Published: (2024)
by: Zhang, Ce, et al.
Published: (2024)
Evaluating Vision Language Model Adaptations for Radiology Report Generation in Low-Resource Languages
by: Salmè, Marco, et al.
Published: (2025)
by: Salmè, Marco, et al.
Published: (2025)
NPHardEval4V: Dynamic Evaluation of Large Vision-Language Models with Effects of Vision
by: Li, Xiang, et al.
Published: (2024)
by: Li, Xiang, et al.
Published: (2024)
Vision-Language Modeling in PET/CT for Visual Grounding of Positive Findings
by: Huemann, Zachary, et al.
Published: (2025)
by: Huemann, Zachary, et al.
Published: (2025)
MSR-Align: Policy-Grounded Multimodal Alignment for Safety-Aware Reasoning in Vision-Language Models
by: Xia, Yinan, et al.
Published: (2025)
by: Xia, Yinan, et al.
Published: (2025)
CROPE: Evaluating In-Context Adaptation of Vision and Language Models to Culture-Specific Concepts
by: Nikandrou, Malvina, et al.
Published: (2024)
by: Nikandrou, Malvina, et al.
Published: (2024)
Rethinking Multilingual Vision-Language Translation: Dataset, Evaluation, and Adaptation
by: Wang, Xintong, et al.
Published: (2025)
by: Wang, Xintong, et al.
Published: (2025)
VideoRoPE: What Makes for Good Video Rotary Position Embedding?
by: Wei, Xilin, et al.
Published: (2025)
by: Wei, Xilin, et al.
Published: (2025)
Rethinking Misalignment in Vision-Language Model Adaptation from a Causal Perspective
by: Zhang, Yanan, et al.
Published: (2024)
by: Zhang, Yanan, et al.
Published: (2024)
Res-Bench: Benchmarking the Robustness of Multimodal Large Language Models to Dynamic Resolution Input
by: Li, Chenxu, et al.
Published: (2025)
by: Li, Chenxu, et al.
Published: (2025)
RoPE-LIME: RoPE-Space Locality + Sparse-K Sampling for Efficient LLM Attribution
by: Picov, Isaac, et al.
Published: (2026)
by: Picov, Isaac, et al.
Published: (2026)
Patch-Prompt Aligned Bayesian Prompt Tuning for Vision-Language Models
by: Liu, Xinyang, et al.
Published: (2023)
by: Liu, Xinyang, et al.
Published: (2023)
The Hard Positive Truth about Vision-Language Compositionality
by: Kamath, Amita, et al.
Published: (2024)
by: Kamath, Amita, et al.
Published: (2024)
Efficient Architectures for High Resolution Vision-Language Models
by: Carvalho, Miguel, et al.
Published: (2025)
by: Carvalho, Miguel, et al.
Published: (2025)
Similar Items
-
ExtraVAR: Stage-Aware RoPE Remapping for Resolution Extrapolation in Visual Autoregressive Models
by: Yan, Feihong, et al.
Published: (2026) -
ReRoPE: Repurposing RoPE for Relative Camera Control
by: Li, Chunyang, et al.
Published: (2026) -
Are Bigger Encoders Always Better in Vision Large Models?
by: Li, Bozhou, et al.
Published: (2024) -
SynthVLM: Towards High-Quality and Efficient Synthesis of Image-Caption Datasets for Vision-Language Models
by: Liu, Zheng, et al.
Published: (2024) -
Spiral RoPE: Rotate Your Rotary Positional Embeddings in the 2D Plane
by: Liu, Haoyu, et al.
Published: (2026)