Video-Panda: Parameter-efficient Alignment for Encoder-free Video-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Yi, Jinhui, Wasim, Syed Talal, Luo, Yanan, Naseer, Muzammal, Gall, Juergen |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
StableMamba: Distillation-free Scaling of Large SSMs for Images and Videos
by: Suleman, Hamid, et al.
Published: (2024)
by: Suleman, Hamid, et al.
Published: (2024)
MixANT: Observation-dependent Memory Propagation for Stochastic Dense Action Anticipation
by: Wasim, Syed Talal, et al.
Published: (2025)
by: Wasim, Syed Talal, et al.
Published: (2025)
Video-GroundingDINO: Towards Open-Vocabulary Spatio-Temporal Video Grounding
by: Wasim, Syed Talal, et al.
Published: (2023)
by: Wasim, Syed Talal, et al.
Published: (2023)
GroupMamba: Efficient Group-Based Visual State Space Model
by: Shaker, Abdelrahman, et al.
Published: (2024)
by: Shaker, Abdelrahman, et al.
Published: (2024)
MV-Match: Multi-View Matching for Domain-Adaptive Identification of Plant Nutrient Deficiencies
by: Yi, Jinhui, et al.
Published: (2024)
by: Yi, Jinhui, et al.
Published: (2024)
STING-BEE: Towards Vision-Language Model for Real-World X-ray Baggage Security Inspection
by: Velayudhan, Divya, et al.
Published: (2025)
by: Velayudhan, Divya, et al.
Published: (2025)
Rethinking temporal self-similarity for repetitive action counting
by: Luo, Yanan, et al.
Published: (2024)
by: Luo, Yanan, et al.
Published: (2024)
Video Panels for Long Video Understanding
by: Doorenbos, Lars, et al.
Published: (2025)
by: Doorenbos, Lars, et al.
Published: (2025)
Robust-LLaVA: On the Effectiveness of Large-Scale Robust Image Encoders for Multi-modal Large Language Models
by: Malik, Hashmat Shadab, et al.
Published: (2025)
by: Malik, Hashmat Shadab, et al.
Published: (2025)
VANE-Bench: Video Anomaly Evaluation Benchmark for Conversational LMMs
by: Bharadwaj, Rohit, et al.
Published: (2024)
by: Bharadwaj, Rohit, et al.
Published: (2024)
Efficient Video Object Segmentation via Modulated Cross-Attention Memory
by: Shaker, Abdelrahman, et al.
Published: (2024)
by: Shaker, Abdelrahman, et al.
Published: (2024)
VURF: A General-purpose Reasoning and Self-refinement Framework for Video Understanding
by: Mahmood, Ahmad, et al.
Published: (2024)
by: Mahmood, Ahmad, et al.
Published: (2024)
Analysing the Robustness of Vision-Language-Models to Common Corruptions
by: Usama, Muhammad, et al.
Published: (2025)
by: Usama, Muhammad, et al.
Published: (2025)
CamC2V: Context-aware Controllable Video Generation
by: Denninger, Luis, et al.
Published: (2025)
by: Denninger, Luis, et al.
Published: (2025)
How Good is my Video LMM? Complex Video Reasoning and Robustness Evaluation Suite for Video-LMMs
by: Khattak, Muhammad Uzair, et al.
Published: (2024)
by: Khattak, Muhammad Uzair, et al.
Published: (2024)
FlowNar: Scalable Streaming Narration for Long-Form Videos
by: Zhong, Zeyun, et al.
Published: (2026)
by: Zhong, Zeyun, et al.
Published: (2026)
PromptSmooth: Certifying Robustness of Medical Vision-Language Models via Prompt Learning
by: Hussein, Noor, et al.
Published: (2024)
by: Hussein, Noor, et al.
Published: (2024)
Language Guided Domain Generalized Medical Image Segmentation
by: Kunhimon, Shahina, et al.
Published: (2024)
by: Kunhimon, Shahina, et al.
Published: (2024)
Composed Video Retrieval via Enriched Context and Discriminative Embeddings
by: Thawakar, Omkar, et al.
Published: (2024)
by: Thawakar, Omkar, et al.
Published: (2024)
SPARROW: Learning Spatial Precision and Temporal Referential Consistency in Pixel-Grounded Video MLLMs
by: Alansari, Mohamad, et al.
Published: (2026)
by: Alansari, Mohamad, et al.
Published: (2026)
Enhancing Video-Based Robot Failure Detection Using Task Knowledge
by: Thoduka, Santosh, et al.
Published: (2025)
by: Thoduka, Santosh, et al.
Published: (2025)
SyncVP: Joint Diffusion for Synchronous Multi-Modal Video Prediction
by: Pallotta, Enrico, et al.
Published: (2025)
by: Pallotta, Enrico, et al.
Published: (2025)
Parameter-free Video Segmentation for Vision and Language Understanding
by: Mahon, Louis, et al.
Published: (2025)
by: Mahon, Louis, et al.
Published: (2025)
Hierarchical Text-to-Vision Self Supervised Alignment for Improved Histopathology Representation Learning
by: Watawana, Hasindri, et al.
Published: (2024)
by: Watawana, Hasindri, et al.
Published: (2024)
Makeup-Guided Facial Privacy Protection via Untrained Neural Network Priors
by: Shamshad, Fahad, et al.
Published: (2024)
by: Shamshad, Fahad, et al.
Published: (2024)
AgriCLIP: Adapting CLIP for Agriculture and Livestock via Domain-Specialized Cross-Model Alignment
by: Nawaz, Umair, et al.
Published: (2024)
by: Nawaz, Umair, et al.
Published: (2024)
STRIVE: Structured Spatiotemporal Exploration for Reinforcement Learning in Video Question Answering
by: Bahrami, Emad, et al.
Published: (2026)
by: Bahrami, Emad, et al.
Published: (2026)
RedSage: A Cybersecurity Generalist LLM
by: Suryanto, Naufal, et al.
Published: (2026)
by: Suryanto, Naufal, et al.
Published: (2026)
Learning to Prompt with Text Only Supervision for Vision-Language Models
by: Khattak, Muhammad Uzair, et al.
Published: (2024)
by: Khattak, Muhammad Uzair, et al.
Published: (2024)
Alignment-free Raw Video Demoireing
by: Xu, Shuning, et al.
Published: (2024)
by: Xu, Shuning, et al.
Published: (2024)
VideoPanda: Video Panoramic Diffusion with Multi-view Attention
by: Xie, Kevin, et al.
Published: (2025)
by: Xie, Kevin, et al.
Published: (2025)
Sequence-Adaptive Video Prediction in Continuous Streams using Diffusion Noise Optimization
by: Azar, Sina Mokhtarzadeh, et al.
Published: (2025)
by: Azar, Sina Mokhtarzadeh, et al.
Published: (2025)
Probing the Efficacy of Federated Parameter-Efficient Fine-Tuning of Vision Transformers for Medical Image Classification
by: Alkhunaizi, Naif, et al.
Published: (2024)
by: Alkhunaizi, Naif, et al.
Published: (2024)
Cross-Modal Self-Training: Aligning Images and Pointclouds to Learn Classification without Labels
by: Dharmasiri, Amaya, et al.
Published: (2024)
by: Dharmasiri, Amaya, et al.
Published: (2024)
EgoControl: Controllable Egocentric Video Generation via 3D Full-Body Poses
by: Pallotta, Enrico, et al.
Published: (2025)
by: Pallotta, Enrico, et al.
Published: (2025)
READ: Recurrent Adapter with Partial Video-Language Alignment for Parameter-Efficient Transfer Learning in Low-Resource Video-Language Modeling
by: Nguyen, Thong, et al.
Published: (2023)
by: Nguyen, Thong, et al.
Published: (2023)
STEREO: A Two-Stage Framework for Adversarially Robust Concept Erasing from Text-to-Image Diffusion Models
by: Srivatsan, Koushik, et al.
Published: (2024)
by: Srivatsan, Koushik, et al.
Published: (2024)
Language-Image Alignment with Fixed Text Encoders
by: Yang, Jingfeng, et al.
Published: (2025)
by: Yang, Jingfeng, et al.
Published: (2025)
Align Your Prompts: Test-Time Prompting with Distribution Alignment for Zero-Shot Generalization
by: Hassan, Jameel, et al.
Published: (2023)
by: Hassan, Jameel, et al.
Published: (2023)
DyCON: Dynamic Uncertainty-aware Consistency and Contrastive Learning for Semi-supervised Medical Image Segmentation
by: Assefa, Maregu, et al.
Published: (2025)
by: Assefa, Maregu, et al.
Published: (2025)
Similar Items
-
StableMamba: Distillation-free Scaling of Large SSMs for Images and Videos
by: Suleman, Hamid, et al.
Published: (2024) -
MixANT: Observation-dependent Memory Propagation for Stochastic Dense Action Anticipation
by: Wasim, Syed Talal, et al.
Published: (2025) -
Video-GroundingDINO: Towards Open-Vocabulary Spatio-Temporal Video Grounding
by: Wasim, Syed Talal, et al.
Published: (2023) -
GroupMamba: Efficient Group-Based Visual State Space Model
by: Shaker, Abdelrahman, et al.
Published: (2024) -
MV-Match: Multi-View Matching for Domain-Adaptive Identification of Plant Nutrient Deficiencies
by: Yi, Jinhui, et al.
Published: (2024)