Saved in:
| Main Authors: | Wu, Kebin, Albreiki, Fatima |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2511.11216 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FineLIP: Extending CLIP's Reach via Fine-Grained Alignment with Longer Text Inputs
by: Asokan, Mothilal, et al.
Published: (2025)
by: Asokan, Mothilal, et al.
Published: (2025)
SynSeg: Feature Synergy for Multi-Category Contrastive Learning in End-to-End Open-Vocabulary Semantic Segmentation
by: Zhang, Weichen, et al.
Published: (2025)
by: Zhang, Weichen, et al.
Published: (2025)
VisCon-100K: Leveraging Contextual Web Data for Fine-tuning Vision Language Models
by: Kumar, Gokul Karthik, et al.
Published: (2025)
by: Kumar, Gokul Karthik, et al.
Published: (2025)
Bias at the End of the Score
by: Magid, Salma Abdel, et al.
Published: (2026)
by: Magid, Salma Abdel, et al.
Published: (2026)
PPE: Positional Preservation Embedding for Token Compression in Multimodal Large Language Models
by: Huang, Mouxiao, et al.
Published: (2025)
by: Huang, Mouxiao, et al.
Published: (2025)
PMPNet: Pixel Movement Prediction Network for Monocular Depth Estimation in Dynamic Scenes
by: Peng, Kebin, et al.
Published: (2024)
by: Peng, Kebin, et al.
Published: (2024)
Anatomical Positional Embeddings
by: Goncharov, Mikhail, et al.
Published: (2024)
by: Goncharov, Mikhail, et al.
Published: (2024)
Learning Robust Correlation with Foundation Model for Weakly-Supervised Few-Shot Segmentation
by: Huang, Xinyang, et al.
Published: (2024)
by: Huang, Xinyang, et al.
Published: (2024)
Active Learning from Scene Embeddings for End-to-End Autonomous Driving
by: Jiang, Wenhao, et al.
Published: (2025)
by: Jiang, Wenhao, et al.
Published: (2025)
An Effective End-to-End Solution for Multimodal Action Recognition
by: Wang, Songping, et al.
Published: (2025)
by: Wang, Songping, et al.
Published: (2025)
Adaptive Begin-of-Video Tokens for Autoregressive Video Diffusion Models
by: Cheng, Tianle, et al.
Published: (2025)
by: Cheng, Tianle, et al.
Published: (2025)
AccidentGPT: Large Multi-Modal Foundation Model for Traffic Accident Analysis
by: Wu, Kebin, et al.
Published: (2024)
by: Wu, Kebin, et al.
Published: (2024)
IPixMatch: Boost Semi-supervised Semantic Segmentation with Inter-Pixel Relation
by: Wu, Kebin, et al.
Published: (2024)
by: Wu, Kebin, et al.
Published: (2024)
Hallucination Begins Where Saliency Drops
by: Zhang, Xiaofeng, et al.
Published: (2026)
by: Zhang, Xiaofeng, et al.
Published: (2026)
Bias and Generalizability of Foundation Models across Datasets in Breast Mammography
by: Germani, Elodie, et al.
Published: (2025)
by: Germani, Elodie, et al.
Published: (2025)
VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks
by: Wu, Jiannan, et al.
Published: (2024)
by: Wu, Jiannan, et al.
Published: (2024)
Positional Embedding-Aware Activations
by: Shah, Kathan, et al.
Published: (2023)
by: Shah, Kathan, et al.
Published: (2023)
Identifying and Mitigating Position Bias of Multi-image Vision-Language Models
by: Tian, Xinyu, et al.
Published: (2025)
by: Tian, Xinyu, et al.
Published: (2025)
Multimodal Action Diffusion for Robust End-to-End Autonomous Driving
by: Rodríguez-Vidal, Jorge Daniel, et al.
Published: (2026)
by: Rodríguez-Vidal, Jorge Daniel, et al.
Published: (2026)
Uni-Mlip: Unified Self-supervision for Medical Vision Language Pre-training
by: Bawazir, Ameera, et al.
Published: (2024)
by: Bawazir, Ameera, et al.
Published: (2024)
Self-Supervised Anomaly Detection in the Wild: Favor Joint Embeddings Methods
by: Otero, Daniel, et al.
Published: (2024)
by: Otero, Daniel, et al.
Published: (2024)
Corner Cases: How Size and Position of Objects Challenge ImageNet-Trained Models
by: Fatima, Mishal, et al.
Published: (2025)
by: Fatima, Mishal, et al.
Published: (2025)
TimeSoccer: An End-to-End Multimodal Large Language Model for Soccer Commentary Generation
by: You, Ling, et al.
Published: (2025)
by: You, Ling, et al.
Published: (2025)
3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding
by: Xiong, Haomiao, et al.
Published: (2025)
by: Xiong, Haomiao, et al.
Published: (2025)
E2E-GMNER: End-to-End Generative Grounded Multimodal Named Entity Recognition
by: Zhang, Meng, et al.
Published: (2026)
by: Zhang, Meng, et al.
Published: (2026)
REMM:Rotation-Equivariant Framework for End-to-End Multimodal Image Matching
by: Nie, Han, et al.
Published: (2024)
by: Nie, Han, et al.
Published: (2024)
AdaptInfer: Adaptive Token Pruning for Vision-Language Model Inference with Dynamical Text Guidance
by: Zhang, Weichen, et al.
Published: (2025)
by: Zhang, Weichen, et al.
Published: (2025)
Revisiting Multimodal Positional Encoding in Vision-Language Models
by: Huang, Jie, et al.
Published: (2025)
by: Huang, Jie, et al.
Published: (2025)
End-to-End Probabilistic Geometry-Guided Regression for 6DoF Object Pose Estimation
by: Pöllabauer, Thomas, et al.
Published: (2024)
by: Pöllabauer, Thomas, et al.
Published: (2024)
Learning to Adapt to Position Bias in Vision Transformer Classifiers
by: Bruintjes, Robert-Jan, et al.
Published: (2025)
by: Bruintjes, Robert-Jan, et al.
Published: (2025)
BBQ-V: Benchmarking Visual Stereotype Bias in Large Multimodal Models
by: Narnaware, Vishal, et al.
Published: (2025)
by: Narnaware, Vishal, et al.
Published: (2025)
Geometry-Aware Rotary Position Embedding for Consistent Video World Model
by: Xiang, Chendong, et al.
Published: (2026)
by: Xiang, Chendong, et al.
Published: (2026)
ObjEmbed: Towards Universal Multimodal Object Embeddings
by: Fu, Shenghao, et al.
Published: (2026)
by: Fu, Shenghao, et al.
Published: (2026)
Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models
by: Xia, Hou, et al.
Published: (2025)
by: Xia, Hou, et al.
Published: (2025)
RT-DETRv3: Real-time End-to-End Object Detection with Hierarchical Dense Positive Supervision
by: Wang, Shuo, et al.
Published: (2024)
by: Wang, Shuo, et al.
Published: (2024)
PLUME: Latent Reasoning Based Universal Multimodal Embedding
by: He, Chenwei, et al.
Published: (2026)
by: He, Chenwei, et al.
Published: (2026)
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs
by: Cheng, Dabing, et al.
Published: (2025)
by: Cheng, Dabing, et al.
Published: (2025)
VisionSelector: End-to-End Learnable Visual Token Compression for Efficient Multimodal LLMs
by: Zhu, Jiaying, et al.
Published: (2025)
by: Zhu, Jiaying, et al.
Published: (2025)
E2E-MFD: Towards End-to-End Synchronous Multimodal Fusion Detection
by: Zhang, Jiaqing, et al.
Published: (2024)
by: Zhang, Jiaqing, et al.
Published: (2024)
Drive-JEPA: Video JEPA Meets Multimodal Trajectory Distillation for End-to-End Driving
by: Wang, Linhan, et al.
Published: (2026)
by: Wang, Linhan, et al.
Published: (2026)
Similar Items
-
FineLIP: Extending CLIP's Reach via Fine-Grained Alignment with Longer Text Inputs
by: Asokan, Mothilal, et al.
Published: (2025) -
SynSeg: Feature Synergy for Multi-Category Contrastive Learning in End-to-End Open-Vocabulary Semantic Segmentation
by: Zhang, Weichen, et al.
Published: (2025) -
VisCon-100K: Leveraging Contextual Web Data for Fine-tuning Vision Language Models
by: Kumar, Gokul Karthik, et al.
Published: (2025) -
Bias at the End of the Score
by: Magid, Salma Abdel, et al.
Published: (2026) -
PPE: Positional Preservation Embedding for Token Compression in Multimodal Large Language Models
by: Huang, Mouxiao, et al.
Published: (2025)