REEF: Relevance-Aware and Efficient LLM Adapter for Video Understanding
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Reza, Sakib, Song, Xiyun, Yu, Heather, Lin, Zongfang, Moghaddam, Mohsen, Camps, Octavia |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
HAT: History-Augmented Anchor Transformer for Online Temporal Action Localization
par: Reza, Sakib, et autres
Publié: (2024)
par: Reza, Sakib, et autres
Publié: (2024)
ePBR: Extended PBR Materials in Image Synthesis
par: Guo, Yu, et autres
Publié: (2025)
par: Guo, Yu, et autres
Publié: (2025)
OT-Talk: Animating 3D Talking Head with Optimal Transportation
par: Wang, Xinmu, et autres
Publié: (2025)
par: Wang, Xinmu, et autres
Publié: (2025)
Inverse Rendering for High-Genus Surface Meshes from Multi-View Images
par: Gao, Xiang, et autres
Publié: (2025)
par: Gao, Xiang, et autres
Publié: (2025)
3D-HGS: 3D Half-Gaussian Splatting
par: Li, Haolin, et autres
Publié: (2024)
par: Li, Haolin, et autres
Publié: (2024)
Solving Masked Jigsaw Puzzles with Diffusion Vision Transformers
par: Liu, Jinyang, et autres
Publié: (2024)
par: Liu, Jinyang, et autres
Publié: (2024)
SIRR-LMM: Single-image Reflection Removal via Large Multimodal Model
par: Guo, Yu, et autres
Publié: (2026)
par: Guo, Yu, et autres
Publié: (2026)
MotionAdapter: Video Motion Transfer via Content-Aware Attention Customization
par: Zhang, Zhexin, et autres
Publié: (2026)
par: Zhang, Zhexin, et autres
Publié: (2026)
Frame2Freq: Spectral Adapters for Fine-Grained Video Understanding
par: Ponbagavathi, Thinesh Thiyakesan, et autres
Publié: (2026)
par: Ponbagavathi, Thinesh Thiyakesan, et autres
Publié: (2026)
Detection of Intoxicated Individuals from Facial Video Sequences via a Recurrent Fusion Model
par: Baroutian, Bita, et autres
Publié: (2025)
par: Baroutian, Bita, et autres
Publié: (2025)
RAP: Efficient Text-Video Retrieval with Sparse-and-Correlated Adapter
par: Cao, Meng, et autres
Publié: (2024)
par: Cao, Meng, et autres
Publié: (2024)
Towards Efficient Partially Relevant Video Retrieval with Active Moment Discovering
par: Song, Peipei, et autres
Publié: (2025)
par: Song, Peipei, et autres
Publié: (2025)
RAG-Adapter: A Plug-and-Play RAG-enhanced Framework for Long Video Understanding
par: Tan, Xichen, et autres
Publié: (2025)
par: Tan, Xichen, et autres
Publié: (2025)
When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding
par: Fang, Pengcheng, et autres
Publié: (2025)
par: Fang, Pengcheng, et autres
Publié: (2025)
A Race Bias Free Face Aging Model for Reliable Kinship Verification
par: Nazari, Ali, et autres
Publié: (2025)
par: Nazari, Ali, et autres
Publié: (2025)
Generalized Disguise Makeup Presentation Attack Detection Using an Attention-Guided Patch-Based Framework
par: Taraghi, Fateme, et autres
Publié: (2026)
par: Taraghi, Fateme, et autres
Publié: (2026)
CAD: Memory Efficient Convolutional Adapter for Segment Anything
par: Kim, Joohyeok, et autres
Publié: (2024)
par: Kim, Joohyeok, et autres
Publié: (2024)
PMA: Towards Parameter-Efficient Point Cloud Understanding via Point Mamba Adapter
par: Zha, Yaohua, et autres
Publié: (2025)
par: Zha, Yaohua, et autres
Publié: (2025)
AdaFocus: Adaptive Relevance-Diversity Sampling with Zero-Cache Look-back for Efficient Long Video Understanding
par: Yang, Xiao, et autres
Publié: (2026)
par: Yang, Xiao, et autres
Publié: (2026)
Hierarchical Awareness Adapters with Hybrid Pyramid Feature Fusion for Dense Depth Prediction
par: Su, Wuqi, et autres
Publié: (2026)
par: Su, Wuqi, et autres
Publié: (2026)
Test-Time Temporal Sampling for Efficient MLLM Video Understanding
par: Wang, Kaibin, et autres
Publié: (2025)
par: Wang, Kaibin, et autres
Publié: (2025)
Dyn-Adapter: Towards Disentangled Representation for Efficient Visual Recognition
par: Zhang, Yurong, et autres
Publié: (2024)
par: Zhang, Yurong, et autres
Publié: (2024)
VideoDetective: Clue Hunting via both Extrinsic Query and Intrinsic Relevance for Long Video Understanding
par: Yang, Ruoliu, et autres
Publié: (2026)
par: Yang, Ruoliu, et autres
Publié: (2026)
Q-Adapter: Visual Query Adapter for Extracting Textually-related Features in Video Captioning
par: Chen, Junan, et autres
Publié: (2025)
par: Chen, Junan, et autres
Publié: (2025)
I2V-Adapter: A General Image-to-Video Adapter for Diffusion Models
par: Guo, Xun, et autres
Publié: (2023)
par: Guo, Xun, et autres
Publié: (2023)
MoMA: Multimodal LLM Adapter for Fast Personalized Image Generation
par: Song, Kunpeng, et autres
Publié: (2024)
par: Song, Kunpeng, et autres
Publié: (2024)
VideoMamba: State Space Model for Efficient Video Understanding
par: Li, Kunchang, et autres
Publié: (2024)
par: Li, Kunchang, et autres
Publié: (2024)
VADER: Towards Causal Video Anomaly Understanding with Relation-Aware Large Language Models
par: Cheng, Ying, et autres
Publié: (2025)
par: Cheng, Ying, et autres
Publié: (2025)
TAMM: TriAdapter Multi-Modal Learning for 3D Shape Understanding
par: Zhang, Zhihao, et autres
Publié: (2024)
par: Zhang, Zhihao, et autres
Publié: (2024)
Thinking with Drafts: Speculative Temporal Reasoning for Efficient Long Video Understanding
par: Hu, Pengfei, et autres
Publié: (2025)
par: Hu, Pengfei, et autres
Publié: (2025)
UrbanSAM: Learning Invariance-Inspired Adapters for Segment Anything Models in Urban Construction
par: Li, Chenyu, et autres
Publié: (2025)
par: Li, Chenyu, et autres
Publié: (2025)
VideoAgent2: Enhancing the LLM-Based Agent System for Long-Form Video Understanding by Uncertainty-Aware CoT
par: Zhi, Zhuo, et autres
Publié: (2025)
par: Zhi, Zhuo, et autres
Publié: (2025)
HIERAMP: Coarse-to-Fine Autoregressive Amplification for Generative Dataset Distillation
par: Zhao, Lin, et autres
Publié: (2026)
par: Zhao, Lin, et autres
Publié: (2026)
AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding
par: Li, Handong, et autres
Publié: (2026)
par: Li, Handong, et autres
Publié: (2026)
M-LLM Based Video Frame Selection for Efficient Video Understanding
par: Hu, Kai, et autres
Publié: (2025)
par: Hu, Kai, et autres
Publié: (2025)
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding
par: Zheng, Duo, et autres
Publié: (2024)
par: Zheng, Duo, et autres
Publié: (2024)
CROME: Cross-Modal Adapters for Efficient Multimodal LLM
par: Ebrahimi, Sayna, et autres
Publié: (2024)
par: Ebrahimi, Sayna, et autres
Publié: (2024)
PickStyle: Video-to-Video Style Transfer with Context-Style Adapters
par: Mehraban, Soroush, et autres
Publié: (2025)
par: Mehraban, Soroush, et autres
Publié: (2025)
MV-Adapter: Multimodal Video Transfer Learning for Video Text Retrieval
par: Jin, Xiaojie, et autres
Publié: (2023)
par: Jin, Xiaojie, et autres
Publié: (2023)
BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning
par: Liu, Ruyang, et autres
Publié: (2023)
par: Liu, Ruyang, et autres
Publié: (2023)
Documents similaires
-
HAT: History-Augmented Anchor Transformer for Online Temporal Action Localization
par: Reza, Sakib, et autres
Publié: (2024) -
ePBR: Extended PBR Materials in Image Synthesis
par: Guo, Yu, et autres
Publié: (2025) -
OT-Talk: Animating 3D Talking Head with Optimal Transportation
par: Wang, Xinmu, et autres
Publié: (2025) -
Inverse Rendering for High-Genus Surface Meshes from Multi-View Images
par: Gao, Xiang, et autres
Publié: (2025) -
3D-HGS: 3D Half-Gaussian Splatting
par: Li, Haolin, et autres
Publié: (2024)