Saved in:
| Main Authors: | Baharlouei, Elaheh, Shafaei, Mahsa, Zhang, Yigeng, Escalante, Hugo Jair, Solorio, Thamar |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2406.07841 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Positive and Risky Message Assessment for Music Products
by: Zhang, Yigeng, et al.
Published: (2023)
by: Zhang, Yigeng, et al.
Published: (2023)
Interpreting Themes from Educational Stories
by: Zhang, Yigeng, et al.
Published: (2024)
by: Zhang, Yigeng, et al.
Published: (2024)
Question-Instructed Visual Descriptions for Zero-Shot Video Question Answering
by: Romero, David, et al.
Published: (2024)
by: Romero, David, et al.
Published: (2024)
Humor in Pixels: Benchmarking Large Multimodal Models Understanding of Online Comics
by: Ryan, Yuriel, et al.
Published: (2025)
by: Ryan, Yuriel, et al.
Published: (2025)
Enhancing NER Performance in Low-Resource Pakistani Languages using Cross-Lingual Data Augmentation
by: Ehsan, Toqeer, et al.
Published: (2025)
by: Ehsan, Toqeer, et al.
Published: (2025)
ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention
by: Liu, Wenjie, et al.
Published: (2026)
by: Liu, Wenjie, et al.
Published: (2026)
CHiP: Cross-modal Hierarchical Direct Preference Optimization for Multimodal LLMs
by: Fu, Jinlan, et al.
Published: (2025)
by: Fu, Jinlan, et al.
Published: (2025)
HyperLoader: Integrating Hypernetwork-Based LoRA and Adapter Layers into Multi-Task Transformers for Sequence Labelling
by: Ortiz-Barajas, Jesus-German, et al.
Published: (2024)
by: Ortiz-Barajas, Jesus-German, et al.
Published: (2024)
Context-aware Adversarial Attack on Named Entity Recognition
by: Chen, Shuguang, et al.
Published: (2023)
by: Chen, Shuguang, et al.
Published: (2023)
Optimizing Multimodal Language Models through Attention-based Interpretability
by: Sergeev, Alexander, et al.
Published: (2025)
by: Sergeev, Alexander, et al.
Published: (2025)
VideoStudio: Generating Consistent-Content and Multi-Scene Videos
by: Long, Fuchen, et al.
Published: (2024)
by: Long, Fuchen, et al.
Published: (2024)
HieraVid: Hierarchical Token Pruning for Fast Video Large Language Models
by: Guo, Yansong, et al.
Published: (2026)
by: Guo, Yansong, et al.
Published: (2026)
Multimodal Abstractive Summarization of Instructional Videos with Vision-Language Models
by: Nazir, Maham, et al.
Published: (2026)
by: Nazir, Maham, et al.
Published: (2026)
A Culturally-diverse Multilingual Multimodal Video Benchmark & Model
by: Shafique, Bhuiyan Sanjid, et al.
Published: (2025)
by: Shafique, Bhuiyan Sanjid, et al.
Published: (2025)
Multimodal LLM Enhanced Cross-lingual Cross-modal Retrieval
by: Wang, Yabing, et al.
Published: (2024)
by: Wang, Yabing, et al.
Published: (2024)
Hierarchical Multimodal Pre-training for Visually Rich Webpage Understanding
by: Xu, Hongshen, et al.
Published: (2024)
by: Xu, Hongshen, et al.
Published: (2024)
Online Misinformation Detection in Live Streaming Videos
by: Cao, Rui
Published: (2025)
by: Cao, Rui
Published: (2025)
Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding
by: Li, Yun, et al.
Published: (2025)
by: Li, Yun, et al.
Published: (2025)
VideoARM: Agentic Reasoning over Hierarchical Memory for Long-Form Video Understanding
by: Yin, Yufei, et al.
Published: (2025)
by: Yin, Yufei, et al.
Published: (2025)
Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models
by: Van, Minh-Hao, et al.
Published: (2025)
by: Van, Minh-Hao, et al.
Published: (2025)
Think While Watching: Online Streaming Segment-Level Memory for Multi-Turn Video Reasoning in Multimodal Large Language Models
by: Wang, Lu, et al.
Published: (2026)
by: Wang, Lu, et al.
Published: (2026)
VideoXum: Cross-modal Visual and Textural Summarization of Videos
by: Lin, Jingyang, et al.
Published: (2023)
by: Lin, Jingyang, et al.
Published: (2023)
ICM-Assistant: Instruction-tuning Multimodal Large Language Models for Rule-based Explainable Image Content Moderation
by: Wu, Mengyang, et al.
Published: (2024)
by: Wu, Mengyang, et al.
Published: (2024)
Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm
by: Tong, Jingqi, et al.
Published: (2025)
by: Tong, Jingqi, et al.
Published: (2025)
Multimodal Integration of Human-Like Attention in Visual Question Answering
by: Sood, Ekta, et al.
Published: (2021)
by: Sood, Ekta, et al.
Published: (2021)
Tell me Habibi, is it Real or Fake?
by: Kuckreja, Kartik, et al.
Published: (2025)
by: Kuckreja, Kartik, et al.
Published: (2025)
Multimodal Transformer for Comics Text-Cloze
by: Vivoli, Emanuele, et al.
Published: (2024)
by: Vivoli, Emanuele, et al.
Published: (2024)
Xuanwu: Evolving General Multimodal Models into an Industrial-Grade Foundation for Content Ecosystems
by: Zhang, Zhiqian, et al.
Published: (2026)
by: Zhang, Zhiqian, et al.
Published: (2026)
Multimodal LLM With Hierarchical Mixture-of-Experts for VQA on 3D Brain MRI
by: Vepa, Arvind Murari, et al.
Published: (2025)
by: Vepa, Arvind Murari, et al.
Published: (2025)
Toward Robust Incomplete Multimodal Sentiment Analysis via Hierarchical Representation Learning
by: Li, Mingcheng, et al.
Published: (2024)
by: Li, Mingcheng, et al.
Published: (2024)
Attend to Evidence: Evidence-Anchored Spatial Attention Supervision for Multimodal RLVR
by: Hu, Ruina, et al.
Published: (2026)
by: Hu, Ruina, et al.
Published: (2026)
Large Content And Behavior Models To Understand, Simulate, And Optimize Content And Behavior
by: Khandelwal, Ashmit, et al.
Published: (2023)
by: Khandelwal, Ashmit, et al.
Published: (2023)
VidCoM: Fast Video Comprehension through Large Language Models with Multimodal Tools
by: Qi, Ji, et al.
Published: (2023)
by: Qi, Ji, et al.
Published: (2023)
Adaptive Cross-lingual Text Classification through In-Context One-Shot Demonstrations
by: Villa-Cueva, Emilio, et al.
Published: (2024)
by: Villa-Cueva, Emilio, et al.
Published: (2024)
Cross-modal Information Flow in Multimodal Large Language Models
by: Zhang, Zhi, et al.
Published: (2024)
by: Zhang, Zhi, et al.
Published: (2024)
From Videos to Indexed Knowledge Graphs -- Framework to Marry Methods for Multimodal Content Analysis and Understanding
by: Rizk, Basem, et al.
Published: (2025)
by: Rizk, Basem, et al.
Published: (2025)
OmChat: A Recipe to Train Multimodal Language Models with Strong Long Context and Video Understanding
by: Zhao, Tiancheng, et al.
Published: (2024)
by: Zhao, Tiancheng, et al.
Published: (2024)
MultiClimate: Multimodal Stance Detection on Climate Change Videos
by: Wang, Jiawen, et al.
Published: (2024)
by: Wang, Jiawen, et al.
Published: (2024)
Less Is More? Selective Visual Attention to High-Importance Regions for Multimodal Radiology Summarization
by: Naznin, Mst. Fahmida Sultana, et al.
Published: (2026)
by: Naznin, Mst. Fahmida Sultana, et al.
Published: (2026)
Make LVLMs Focus: Context-Aware Attention Modulation for Better Multimodal In-Context Learning
by: Li, Yanshu, et al.
Published: (2025)
by: Li, Yanshu, et al.
Published: (2025)
Similar Items
-
Positive and Risky Message Assessment for Music Products
by: Zhang, Yigeng, et al.
Published: (2023) -
Interpreting Themes from Educational Stories
by: Zhang, Yigeng, et al.
Published: (2024) -
Question-Instructed Visual Descriptions for Zero-Shot Video Question Answering
by: Romero, David, et al.
Published: (2024) -
Humor in Pixels: Benchmarking Large Multimodal Models Understanding of Online Comics
by: Ryan, Yuriel, et al.
Published: (2025) -
Enhancing NER Performance in Low-Resource Pakistani Languages using Cross-Lingual Data Augmentation
by: Ehsan, Toqeer, et al.
Published: (2025)