HoPE: Hybrid of Position Embedding for Long Context Vision-Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Haoran, Qin, Yingjie, Ou, Baoyuan, Xu, Lai, Xu, Ruiwen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HoPE: A Novel Positional Encoding Without Long-Term Decay for Enhanced Context Awareness and Extrapolation
von: Chen, Yuhan, et al.
Veröffentlicht: (2024)
von: Chen, Yuhan, et al.
Veröffentlicht: (2024)
LongProLIP: A Probabilistic Vision-Language Model with Long Context Text
von: Chun, Sanghyuk, et al.
Veröffentlicht: (2025)
von: Chun, Sanghyuk, et al.
Veröffentlicht: (2025)
Rotary Position Embedding for Vision Transformer
von: Heo, Byeongho, et al.
Veröffentlicht: (2024)
von: Heo, Byeongho, et al.
Veröffentlicht: (2024)
BendVLM: Test-Time Debiasing of Vision-Language Embeddings
von: Gerych, Walter, et al.
Veröffentlicht: (2024)
von: Gerych, Walter, et al.
Veröffentlicht: (2024)
Efficient Matrix Implementation for Rotary Position Embedding
von: Minqi, Chen, et al.
Veröffentlicht: (2026)
von: Minqi, Chen, et al.
Veröffentlicht: (2026)
Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks
von: Ding, Yuhe, et al.
Veröffentlicht: (2024)
von: Ding, Yuhe, et al.
Veröffentlicht: (2024)
Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models
von: Li, Xu, et al.
Veröffentlicht: (2024)
von: Li, Xu, et al.
Veröffentlicht: (2024)
Maximizing the Position Embedding for Vision Transformers with Global Average Pooling
von: Lee, Wonjun, et al.
Veröffentlicht: (2025)
von: Lee, Wonjun, et al.
Veröffentlicht: (2025)
Conceptualizing Embeddings: Sparse Disentanglement for Vision-Language Models
von: Kubaty, Piotr, et al.
Veröffentlicht: (2026)
von: Kubaty, Piotr, et al.
Veröffentlicht: (2026)
Lever LM: Configuring In-Context Sequence to Lever Large Vision Language Models
von: Yang, Xu, et al.
Veröffentlicht: (2023)
von: Yang, Xu, et al.
Veröffentlicht: (2023)
Position: Do Not Explain Vision Models Without Context
von: Tomaszewska, Paulina, et al.
Veröffentlicht: (2024)
von: Tomaszewska, Paulina, et al.
Veröffentlicht: (2024)
Efficient and Long-Tailed Generalization for Pre-trained Vision-Language Model
von: Shi, Jiang-Xin, et al.
Veröffentlicht: (2024)
von: Shi, Jiang-Xin, et al.
Veröffentlicht: (2024)
LoGeR: Long-Context Geometric Reconstruction with Hybrid Memory
von: Zhang, Junyi, et al.
Veröffentlicht: (2026)
von: Zhang, Junyi, et al.
Veröffentlicht: (2026)
Spurious Feature Eraser: Stabilizing Test-Time Adaptation for Vision-Language Foundation Model
von: Ma, Huan, et al.
Veröffentlicht: (2024)
von: Ma, Huan, et al.
Veröffentlicht: (2024)
RayRoPE: Projective Ray Positional Encoding for Multi-view Attention
von: Wu, Yu, et al.
Veröffentlicht: (2026)
von: Wu, Yu, et al.
Veröffentlicht: (2026)
Cross-Modal Redundancy and the Geometry of Vision-Language Embeddings
von: Dhimoïla, Grégoire, et al.
Veröffentlicht: (2026)
von: Dhimoïla, Grégoire, et al.
Veröffentlicht: (2026)
Domain Adaptation with a Single Vision-Language Embedding
von: Fahes, Mohammad, et al.
Veröffentlicht: (2024)
von: Fahes, Mohammad, et al.
Veröffentlicht: (2024)
Transferable Adversarial Attacks on Black-Box Vision-Language Models
von: Hu, Kai, et al.
Veröffentlicht: (2025)
von: Hu, Kai, et al.
Veröffentlicht: (2025)
Continual Learning with Vision-Language Models via Semantic-Geometry Preservation
von: He, Chiyuan, et al.
Veröffentlicht: (2026)
von: He, Chiyuan, et al.
Veröffentlicht: (2026)
Probabilistic Embeddings for Frozen Vision-Language Models: Uncertainty Quantification with Gaussian Process Latent Variable Models
von: Venkataramanan, Aishwarya, et al.
Veröffentlicht: (2025)
von: Venkataramanan, Aishwarya, et al.
Veröffentlicht: (2025)
PracticalDG: Perturbation Distillation on Vision-Language Models for Hybrid Domain Generalization
von: Chen, Zining, et al.
Veröffentlicht: (2024)
von: Chen, Zining, et al.
Veröffentlicht: (2024)
Fair Context Learning for Evidence-Balanced Test-Time Adaptation in Vision-Language Models
von: Yun, Sanggeon, et al.
Veröffentlicht: (2026)
von: Yun, Sanggeon, et al.
Veröffentlicht: (2026)
V2PE: Improving Multimodal Long-Context Capability of Vision-Language Models with Variable Visual Position Encoding
von: Ge, Junqi, et al.
Veröffentlicht: (2024)
von: Ge, Junqi, et al.
Veröffentlicht: (2024)
PRISM: Reducing Spurious Implicit Biases in Vision-Language Models with LLM-Guided Embedding Projection
von: Molahasani, Mahdiyar, et al.
Veröffentlicht: (2025)
von: Molahasani, Mahdiyar, et al.
Veröffentlicht: (2025)
Towards Long-window Anchoring in Vision-Language Model Distillation
von: Zhou, Haoyi, et al.
Veröffentlicht: (2025)
von: Zhou, Haoyi, et al.
Veröffentlicht: (2025)
Scaling Vision Language Models for Pharmaceutical Long Form Video Reasoning on Industrial GenAI Platform
von: Mishra, Suyash, et al.
Veröffentlicht: (2026)
von: Mishra, Suyash, et al.
Veröffentlicht: (2026)
HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction
von: Bao, Chen, et al.
Veröffentlicht: (2024)
von: Bao, Chen, et al.
Veröffentlicht: (2024)
Model-agnostic Adversarial Attack and Defense for Vision-Language-Action Models
von: Xu, Haochuan, et al.
Veröffentlicht: (2025)
von: Xu, Haochuan, et al.
Veröffentlicht: (2025)
Benchmarking Convolutional, Transformer, Hybrid, and Vision Language Models for Multi Disease Retinal Screening
von: Dey, Durjoy, et al.
Veröffentlicht: (2026)
von: Dey, Durjoy, et al.
Veröffentlicht: (2026)
Structuring GUI Elements through Vision Language Models: Towards Action Space Generation
von: Xu, Yi, et al.
Veröffentlicht: (2025)
von: Xu, Yi, et al.
Veröffentlicht: (2025)
Unleashing the Power of Vision-Language Models for Long-Tailed Multi-Label Visual Recognition
von: Tang, Wei, et al.
Veröffentlicht: (2025)
von: Tang, Wei, et al.
Veröffentlicht: (2025)
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
von: Li, Xinhao, et al.
Veröffentlicht: (2024)
von: Li, Xinhao, et al.
Veröffentlicht: (2024)
Vision Foundation Model Embedding-Based Semantic Anomaly Detection
von: Ronecker, Max Peter, et al.
Veröffentlicht: (2025)
von: Ronecker, Max Peter, et al.
Veröffentlicht: (2025)
CAPT: Class-Aware Prompt Tuning for Federated Long-Tailed Learning with Vision-Language Model
von: Hou, Shihao, et al.
Veröffentlicht: (2025)
von: Hou, Shihao, et al.
Veröffentlicht: (2025)
Atlas: Multi-Scale Attention Improves Long Context Image Modeling
von: Agrawal, Kumar Krishna, et al.
Veröffentlicht: (2025)
von: Agrawal, Kumar Krishna, et al.
Veröffentlicht: (2025)
VisionPAD: A Vision-Centric Pre-training Paradigm for Autonomous Driving
von: Zhang, Haiming, et al.
Veröffentlicht: (2024)
von: Zhang, Haiming, et al.
Veröffentlicht: (2024)
Qwen Look Again: Guiding Vision-Language Reasoning Models to Re-attention Visual Information
von: Chu, Xu, et al.
Veröffentlicht: (2025)
von: Chu, Xu, et al.
Veröffentlicht: (2025)
SeqPE: Transformer with Sequential Position Encoding
von: Li, Huayang, et al.
Veröffentlicht: (2025)
von: Li, Huayang, et al.
Veröffentlicht: (2025)
Configuring Data Augmentations to Reduce Variance Shift in Positional Embedding of Vision Transformers
von: Kim, Bum Jun, et al.
Veröffentlicht: (2024)
von: Kim, Bum Jun, et al.
Veröffentlicht: (2024)
Light-weight Fine-tuning Method for Defending Adversarial Noise in Pre-trained Medical Vision-Language Models
von: Han, Xu, et al.
Veröffentlicht: (2024)
von: Han, Xu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
HoPE: A Novel Positional Encoding Without Long-Term Decay for Enhanced Context Awareness and Extrapolation
von: Chen, Yuhan, et al.
Veröffentlicht: (2024) -
LongProLIP: A Probabilistic Vision-Language Model with Long Context Text
von: Chun, Sanghyuk, et al.
Veröffentlicht: (2025) -
Rotary Position Embedding for Vision Transformer
von: Heo, Byeongho, et al.
Veröffentlicht: (2024) -
BendVLM: Test-Time Debiasing of Vision-Language Embeddings
von: Gerych, Walter, et al.
Veröffentlicht: (2024) -
Efficient Matrix Implementation for Rotary Position Embedding
von: Minqi, Chen, et al.
Veröffentlicht: (2026)