Unifying Specialized Visual Encoders for Video Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chung, Jihoon, Zhu, Tyler, Saez-Diez, Max Gonzalez, Niebles, Juan Carlos, Zhou, Honglu, Russakovsky, Olga |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Seeing Beyond the Scene: Analyzing and Mitigating Background Bias in Action Recognition
von: Zhou, Ellie, et al.
Veröffentlicht: (2025)
von: Zhou, Ellie, et al.
Veröffentlicht: (2025)
Video Models Reason Early: Exploiting Plan Commitment for Maze Solving
von: Newman, Kaleb, et al.
Veröffentlicht: (2026)
von: Newman, Kaleb, et al.
Veröffentlicht: (2026)
Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding
von: Wang, Ziyang, et al.
Veröffentlicht: (2025)
von: Wang, Ziyang, et al.
Veröffentlicht: (2025)
SMILE: A Composite Lexical-Semantic Metric for Question-Answering Evaluation
von: Kendre, Shrikant, et al.
Veröffentlicht: (2025)
von: Kendre, Shrikant, et al.
Veröffentlicht: (2025)
ViUniT: Visual Unit Tests for More Robust Visual Programming
von: Panagopoulou, Artemis, et al.
Veröffentlicht: (2024)
von: Panagopoulou, Artemis, et al.
Veröffentlicht: (2024)
xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs
von: Ryoo, Michael S., et al.
Veröffentlicht: (2024)
von: Ryoo, Michael S., et al.
Veröffentlicht: (2024)
Analyzing the Roles of Language and Vision in Learning from Limited Data
von: Chen, Allison, et al.
Veröffentlicht: (2024)
von: Chen, Allison, et al.
Veröffentlicht: (2024)
Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization
von: Jin, Yang, et al.
Veröffentlicht: (2024)
von: Jin, Yang, et al.
Veröffentlicht: (2024)
ICONS: Influence Consensus for Vision-Language Data Selection
von: Wu, Xindi, et al.
Veröffentlicht: (2024)
von: Wu, Xindi, et al.
Veröffentlicht: (2024)
D2D: Detector-to-Differentiable Critic for Improved Numeracy in Text-to-Image Generation
von: Yoo, Nobline, et al.
Veröffentlicht: (2025)
von: Yoo, Nobline, et al.
Veröffentlicht: (2025)
Attention IoU: Examining Biases in CelebA using Attention Maps
von: Serianni, Aaron, et al.
Veröffentlicht: (2025)
von: Serianni, Aaron, et al.
Veröffentlicht: (2025)
UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation
von: Lin, Bin, et al.
Veröffentlicht: (2025)
von: Lin, Bin, et al.
Veröffentlicht: (2025)
Linear Scaling Video VLMs for Long Video Understanding
von: Eyzaguirre, Cristobal, et al.
Veröffentlicht: (2026)
von: Eyzaguirre, Cristobal, et al.
Veröffentlicht: (2026)
AdaVid: Adaptive Video-Language Pretraining
von: Patel, Chaitanya, et al.
Veröffentlicht: (2025)
von: Patel, Chaitanya, et al.
Veröffentlicht: (2025)
LongVILA: Scaling Long-Context Visual Language Models for Long Videos
von: Chen, Yukang, et al.
Veröffentlicht: (2024)
von: Chen, Yukang, et al.
Veröffentlicht: (2024)
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data
von: Zhou, Honglu, et al.
Veröffentlicht: (2025)
von: Zhou, Honglu, et al.
Veröffentlicht: (2025)
Future Optical Flow Prediction Improves Robot Control & Video Generation
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2026)
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2026)
DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies
von: Song, Wei, et al.
Veröffentlicht: (2025)
von: Song, Wei, et al.
Veröffentlicht: (2025)
Vision-Language Dataset Distillation
von: Wu, Xindi, et al.
Veröffentlicht: (2023)
von: Wu, Xindi, et al.
Veröffentlicht: (2023)
Unifying Vision-Language Latents for Zero-label Image Caption Enhancement
von: Byun, Sanghyun, et al.
Veröffentlicht: (2025)
von: Byun, Sanghyun, et al.
Veröffentlicht: (2025)
Rethinking the Mixture of Vision Encoders Paradigm for Enhanced Visual Understanding in Multimodal LLMs
von: Azadani, Mozhgan Nasr, et al.
Veröffentlicht: (2025)
von: Azadani, Mozhgan Nasr, et al.
Veröffentlicht: (2025)
Beyond the Vision Encoder: Identifying and Mitigating Spatial Bias in Large Vision-Language Models
von: Zhu, Yingjie, et al.
Veröffentlicht: (2025)
von: Zhu, Yingjie, et al.
Veröffentlicht: (2025)
Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times
von: Loginova, Olga, et al.
Veröffentlicht: (2025)
von: Loginova, Olga, et al.
Veröffentlicht: (2025)
Frozen Transformers in Language Models Are Effective Visual Encoder Layers
von: Pang, Ziqi, et al.
Veröffentlicht: (2023)
von: Pang, Ziqi, et al.
Veröffentlicht: (2023)
Visual Compositional Tuning
von: Wu, Xindi, et al.
Veröffentlicht: (2025)
von: Wu, Xindi, et al.
Veröffentlicht: (2025)
A Simple Recipe for Contrastively Pre-training Video-First Encoders Beyond 16 Frames
von: Papalampidi, Pinelopi, et al.
Veröffentlicht: (2023)
von: Papalampidi, Pinelopi, et al.
Veröffentlicht: (2023)
ImageNet-OOD: Deciphering Modern Out-of-Distribution Detection Algorithms
von: Yang, William, et al.
Veröffentlicht: (2023)
von: Yang, William, et al.
Veröffentlicht: (2023)
Pose-Based Sign Language Spotting via an End-to-End Encoder Architecture
von: Johnny, Samuel Ebimobowei, et al.
Veröffentlicht: (2025)
von: Johnny, Samuel Ebimobowei, et al.
Veröffentlicht: (2025)
Vision-Language Modeling in PET/CT for Visual Grounding of Positive Findings
von: Huemann, Zachary, et al.
Veröffentlicht: (2025)
von: Huemann, Zachary, et al.
Veröffentlicht: (2025)
VideoXum: Cross-modal Visual and Textural Summarization of Videos
von: Lin, Jingyang, et al.
Veröffentlicht: (2023)
von: Lin, Jingyang, et al.
Veröffentlicht: (2023)
Modality-Specialized Synergizers for Interleaved Vision-Language Generalists
von: Xu, Zhiyang, et al.
Veröffentlicht: (2024)
von: Xu, Zhiyang, et al.
Veröffentlicht: (2024)
RoboUniView: Visual-Language Model with Unified View Representation for Robotic Manipulation
von: Liu, Fanfan, et al.
Veröffentlicht: (2024)
von: Liu, Fanfan, et al.
Veröffentlicht: (2024)
Exploring the Design Space of Visual Context Representation in Video MLLMs
von: Du, Yifan, et al.
Veröffentlicht: (2024)
von: Du, Yifan, et al.
Veröffentlicht: (2024)
Visual In-Context Learning for Large Vision-Language Models
von: Zhou, Yucheng, et al.
Veröffentlicht: (2024)
von: Zhou, Yucheng, et al.
Veröffentlicht: (2024)
X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning
von: Panagopoulou, Artemis, et al.
Veröffentlicht: (2023)
von: Panagopoulou, Artemis, et al.
Veröffentlicht: (2023)
UniVS: Unified and Universal Video Segmentation with Prompts as Queries
von: Li, Minghan, et al.
Veröffentlicht: (2024)
von: Li, Minghan, et al.
Veröffentlicht: (2024)
Bootstrapping Action-Grounded Visual Dynamics in Unified Vision-Language Models
von: Qiu, Yifu, et al.
Veröffentlicht: (2025)
von: Qiu, Yifu, et al.
Veröffentlicht: (2025)
UTPTrack: Towards Simple and Unified Token Pruning for Visual Tracking
von: Wu, Hao, et al.
Veröffentlicht: (2026)
von: Wu, Hao, et al.
Veröffentlicht: (2026)
MoTE: Reconciling Generalization with Specialization for Visual-Language to Video Knowledge Transfer
von: Zhu, Minghao, et al.
Veröffentlicht: (2024)
von: Zhu, Minghao, et al.
Veröffentlicht: (2024)
Layer-wise Alignment: Examining Safety Alignment Across Image Encoder Layers in Vision Language Models
von: Bachu, Saketh, et al.
Veröffentlicht: (2024)
von: Bachu, Saketh, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Seeing Beyond the Scene: Analyzing and Mitigating Background Bias in Action Recognition
von: Zhou, Ellie, et al.
Veröffentlicht: (2025) -
Video Models Reason Early: Exploiting Plan Commitment for Maze Solving
von: Newman, Kaleb, et al.
Veröffentlicht: (2026) -
Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding
von: Wang, Ziyang, et al.
Veröffentlicht: (2025) -
SMILE: A Composite Lexical-Semantic Metric for Question-Answering Evaluation
von: Kendre, Shrikant, et al.
Veröffentlicht: (2025) -
ViUniT: Visual Unit Tests for More Robust Visual Programming
von: Panagopoulou, Artemis, et al.
Veröffentlicht: (2024)