ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task
Fuente:
arXiv
Saved in:
| Main Authors: | Khalil, Ahmad, Khalil, Mahmoud, Ngom, Alioune |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ResNetVLLM-2: Addressing ResNetVLLM's Multi-Modal Hallucinations
by: Khalil, Ahmad, et al.
Published: (2025)
by: Khalil, Ahmad, et al.
Published: (2025)
Representation Learning with Adaptive Superpixel Coding
by: Khalil, Mahmoud, et al.
Published: (2025)
by: Khalil, Mahmoud, et al.
Published: (2025)
IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes
by: Liang, Yujia, et al.
Published: (2025)
by: Liang, Yujia, et al.
Published: (2025)
ViGText: Deepfake Image Detection with Vision-Language Model Explanations and Graph Neural Networks
by: ALBarqawi, Ahmad, et al.
Published: (2025)
by: ALBarqawi, Ahmad, et al.
Published: (2025)
PhyVLLM: Physics-Guided Video Language Model with Motion-Appearance Disentanglement
by: Zhan, Yu-Wei, et al.
Published: (2025)
by: Zhan, Yu-Wei, et al.
Published: (2025)
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens
by: Lu, Zhuqiang, et al.
Published: (2024)
by: Lu, Zhuqiang, et al.
Published: (2024)
Expand VSR Benchmark for VLLM to Expertize in Spatial Rules
by: Xie, Peijin, et al.
Published: (2024)
by: Xie, Peijin, et al.
Published: (2024)
Multimodal Chaptering for Long-Form TV Newscast Video
by: Guetari, Khalil, et al.
Published: (2024)
by: Guetari, Khalil, et al.
Published: (2024)
TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types
by: Chen, Jiankang, et al.
Published: (2025)
by: Chen, Jiankang, et al.
Published: (2025)
MFI-ResNet: Efficient ResNet Architecture Optimization via MeanFlow Compression and Selective Incubation
by: Sun, Nuolin, et al.
Published: (2025)
by: Sun, Nuolin, et al.
Published: (2025)
Tackling the Abstraction and Reasoning Corpus with Vision Transformers: the Importance of 2D Representation, Positions, and Objects
by: Li, Wenhao, et al.
Published: (2024)
by: Li, Wenhao, et al.
Published: (2024)
M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding
by: Liu, Shenxi, et al.
Published: (2025)
by: Liu, Shenxi, et al.
Published: (2025)
CapsFake: A Multimodal Capsule Network for Detecting Instruction-Guided Deepfakes
by: Nguyen, Tuan, et al.
Published: (2025)
by: Nguyen, Tuan, et al.
Published: (2025)
Knowledge-based learning in Text-RAG and Image-RAG
by: Shim, Alexander, et al.
Published: (2026)
by: Shim, Alexander, et al.
Published: (2026)
GazeVLM: A Vision-Language Model for Multi-Task Gaze Understanding
by: Mathew, Athul M., et al.
Published: (2025)
by: Mathew, Athul M., et al.
Published: (2025)
Res2NetFuse: A Novel Res2Net-based Fusion Method for Infrared and Visible Images
by: Song, Xu, et al.
Published: (2021)
by: Song, Xu, et al.
Published: (2021)
Beyond Static Vision: Scene Dynamic Field Unlocks Intuitive Physics Understanding in Multi-modal Large Language Models
by: Li, Nanxi, et al.
Published: (2026)
by: Li, Nanxi, et al.
Published: (2026)
Quotient Network -- A Network Similar to ResNet but Learning Quotients
by: Hui, Peng, et al.
Published: (2025)
by: Hui, Peng, et al.
Published: (2025)
Interpreting ResNet-based CLIP via Neuron-Attention Decomposition
by: Bu, Edmund, et al.
Published: (2025)
by: Bu, Edmund, et al.
Published: (2025)
Wavelet-based GAN Fingerprint Detection using ResNet50
by: Erukude, Sai Teja, et al.
Published: (2025)
by: Erukude, Sai Teja, et al.
Published: (2025)
Concept Visualization: Explaining the CLIP Multi-modal Embedding Using WordNet
by: Giulivi, Loris, et al.
Published: (2024)
by: Giulivi, Loris, et al.
Published: (2024)
Adaptation of Multi-modal Representation Models for Multi-task Surgical Computer Vision
by: Walimbe, Soham, et al.
Published: (2025)
by: Walimbe, Soham, et al.
Published: (2025)
Exploring Efficient Foundational Multi-modal Models for Video Summarization
by: Samel, Karan, et al.
Published: (2024)
by: Samel, Karan, et al.
Published: (2024)
ResNet: Enabling Deep Convolutional Neural Networks through Residual Learning
by: Liu, Xingyu, et al.
Published: (2025)
by: Liu, Xingyu, et al.
Published: (2025)
ZERO: Industry-ready Vision Foundation Model with Multi-modal Prompts
by: Choi, Sangbum, et al.
Published: (2025)
by: Choi, Sangbum, et al.
Published: (2025)
Explaining Multi-modal Large Language Models by Analyzing their Vision Perception
by: Giulivi, Loris, et al.
Published: (2024)
by: Giulivi, Loris, et al.
Published: (2024)
M-LLM Based Video Frame Selection for Efficient Video Understanding
by: Hu, Kai, et al.
Published: (2025)
by: Hu, Kai, et al.
Published: (2025)
Learning Multi-modal Representations by Watching Hundreds of Surgical Video Lectures
by: Yuan, Kun, et al.
Published: (2023)
by: Yuan, Kun, et al.
Published: (2023)
VISion On Request: Enhanced VLLM efficiency with sparse, dynamically selected, vision-language interactions
by: Bulat, Adrian, et al.
Published: (2026)
by: Bulat, Adrian, et al.
Published: (2026)
Research on Brain Tumor Classification Method Based on Improved ResNet34 Network
by: Li, Yufeng, et al.
Published: (2025)
by: Li, Yufeng, et al.
Published: (2025)
SpatialLLM: From Multi-modality Data to Urban Spatial Intelligence
by: Chen, Jiabin, et al.
Published: (2025)
by: Chen, Jiabin, et al.
Published: (2025)
Transfer Learning for Wildlife Classification: Evaluating YOLOv8 against DenseNet, ResNet, and VGGNet on a Custom Dataset
by: Sharma, Subek, et al.
Published: (2024)
by: Sharma, Subek, et al.
Published: (2024)
Multi-Phase Automated Segmentation of Dental Structures in CBCT Using a Lightweight Auto3DSeg and SegResNet Implementation
by: LaBella, Dominic, et al.
Published: (2025)
by: LaBella, Dominic, et al.
Published: (2025)
PerspectiveNet: Multi-View Perception for Dynamic Scene Understanding
by: Nguyen, Vinh
Published: (2024)
by: Nguyen, Vinh
Published: (2024)
Learning Emergent Modular Representations in Multi-modality Medical Vision Foundation Models
by: He, Yuting, et al.
Published: (2026)
by: He, Yuting, et al.
Published: (2026)
Overview of the NLPCC 2025 Shared Task 4: Multi-modal, Multilingual, and Multi-hop Medical Instructional Video Question Answering Challenge
by: Li, Bin, et al.
Published: (2025)
by: Li, Bin, et al.
Published: (2025)
VideoWebArena: Evaluating Long Context Multimodal Agents with Video Understanding Web Tasks
by: Jang, Lawrence, et al.
Published: (2024)
by: Jang, Lawrence, et al.
Published: (2024)
Task-Aware KV Compression For Cost-Effective Long Video Understanding
by: Qin, Minghao, et al.
Published: (2025)
by: Qin, Minghao, et al.
Published: (2025)
Solution for SMART-101 Challenge of CVPR Multi-modal Algorithmic Reasoning Task 2024
by: Ahn, Jinwoo, et al.
Published: (2024)
by: Ahn, Jinwoo, et al.
Published: (2024)
XAI-Driven Skin Disease Classification: Leveraging GANs to Augment ResNet-50 Performance
by: Villanueva, Kim Gerard A., et al.
Published: (2025)
by: Villanueva, Kim Gerard A., et al.
Published: (2025)
Similar Items
-
ResNetVLLM-2: Addressing ResNetVLLM's Multi-Modal Hallucinations
by: Khalil, Ahmad, et al.
Published: (2025) -
Representation Learning with Adaptive Superpixel Coding
by: Khalil, Mahmoud, et al.
Published: (2025) -
IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes
by: Liang, Yujia, et al.
Published: (2025) -
ViGText: Deepfake Image Detection with Vision-Language Model Explanations and Graph Neural Networks
by: ALBarqawi, Ahmad, et al.
Published: (2025) -
PhyVLLM: Physics-Guided Video Language Model with Motion-Appearance Disentanglement
by: Zhan, Yu-Wei, et al.
Published: (2025)