ViSIL: Unified Evaluation of Information Loss in Multimodal Video Captioning
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Po-han, Chen, Shenghui, Topcu, Ufuk, Chinchali, Sandeep |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VIBE: Annotation-Free Video-to-Text Information Bottleneck Evaluation for TL;DR
by: Chen, Shenghui, et al.
Published: (2025)
by: Chen, Shenghui, et al.
Published: (2025)
Human-Agent Coordination in Games under Incomplete Information via Multi-Step Intent
by: Chen, Shenghui, et al.
Published: (2024)
by: Chen, Shenghui, et al.
Published: (2024)
CSA: Data-efficient Mapping of Unimodal Features to Multimodal Features
by: Li, Po-han, et al.
Published: (2024)
by: Li, Po-han, et al.
Published: (2024)
Human-Agent Cooperation in Games under Incomplete Information through Natural Language Communication
by: Chen, Shenghui, et al.
Published: (2024)
by: Chen, Shenghui, et al.
Published: (2024)
Any2Any: Incomplete Multimodal Retrieval with Conformal Prediction
by: Li, Po-han, et al.
Published: (2024)
by: Li, Po-han, et al.
Published: (2024)
NarrativeBridge: Enhancing Video Captioning with Causal-Temporal Narrative
by: Nadeem, Asmar, et al.
Published: (2024)
by: Nadeem, Asmar, et al.
Published: (2024)
ViEEG: Hierarchical Visual Neural Representation for EEG Brain Decoding
by: Liu, Minxu, et al.
Published: (2025)
by: Liu, Minxu, et al.
Published: (2025)
Siamese-Driven Optimization for Low-Resolution Image Latent Embedding in Image Captioning
by: Tan, Jing Jie, et al.
Published: (2025)
by: Tan, Jing Jie, et al.
Published: (2025)
Evaluating Human Trust in LLM-Based Planners: A Preliminary Study
by: Chen, Shenghui, et al.
Published: (2025)
by: Chen, Shenghui, et al.
Published: (2025)
GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents
by: Ouyang, Mingyu, et al.
Published: (2026)
by: Ouyang, Mingyu, et al.
Published: (2026)
"It's trained by non-disabled people": Evaluating How Image Quality Affects Product Captioning with Vision-Language Models
by: Garg, Kapil, et al.
Published: (2025)
by: Garg, Kapil, et al.
Published: (2025)
VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation
by: Zhao, Yiming, et al.
Published: (2026)
by: Zhao, Yiming, et al.
Published: (2026)
ViT-Explainer: An Interactive Walkthrough of the Vision Transformer Pipeline
by: Hernandez, Juan Manuel, et al.
Published: (2026)
by: Hernandez, Juan Manuel, et al.
Published: (2026)
Seeing, Hearing, and Knowing Together: Multimodal Strategies in Deepfake Videos Detection
by: Chen, Chen, et al.
Published: (2026)
by: Chen, Chen, et al.
Published: (2026)
UI-UG: A Unified MLLM for UI Understanding and Generation
by: Yang, Hao, et al.
Published: (2025)
by: Yang, Hao, et al.
Published: (2025)
Five Years of SciCap: What We Learned and Future Directions for Scientific Figure Captioning
by: Huang, Ting-Hao 'Kenneth', et al.
Published: (2025)
by: Huang, Ting-Hao 'Kenneth', et al.
Published: (2025)
ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Scientific Workflows
by: Sun, Qiushi, et al.
Published: (2025)
by: Sun, Qiushi, et al.
Published: (2025)
EEG-based Multimodal Representation Learning for Emotion Recognition
by: Yin, Kang, et al.
Published: (2024)
by: Yin, Kang, et al.
Published: (2024)
DAVE: Distribution-aware Attribution via ViT Gradient Decomposition
by: Wróbel, Adam, et al.
Published: (2026)
by: Wróbel, Adam, et al.
Published: (2026)
GazeLLM: Multimodal LLMs incorporating Human Visual Attention
by: Rekimoto, Jun
Published: (2025)
by: Rekimoto, Jun
Published: (2025)
Automated ARAT Scoring Using Multimodal Video Analysis, Multi-View Fusion, and Hierarchical Bayesian Models: A Clinician Study
by: Ahmed, Tamim, et al.
Published: (2025)
by: Ahmed, Tamim, et al.
Published: (2025)
Towards a Multimodal Document-grounded Conversational AI System for Education
by: Taneja, Karan, et al.
Published: (2025)
by: Taneja, Karan, et al.
Published: (2025)
OmniResponse: Online Multimodal Conversational Response Generation in Dyadic Interactions
by: Luo, Cheng, et al.
Published: (2025)
by: Luo, Cheng, et al.
Published: (2025)
A Multimodal Dataset of Student Oral Presentations with Sensors and Evaluation Data
by: Becerra, Alvaro, et al.
Published: (2026)
by: Becerra, Alvaro, et al.
Published: (2026)
Acoustic Field Video for Multimodal Scene Understanding
by: Kim, Daehwa, et al.
Published: (2026)
by: Kim, Daehwa, et al.
Published: (2026)
Milmer: a Framework for Multiple Instance Learning based Multimodal Emotion Recognition
by: Wang, Zaitian, et al.
Published: (2025)
by: Wang, Zaitian, et al.
Published: (2025)
CG-MER: A Card Game-based Multimodal dataset for Emotion Recognition
by: Farhat, Nessrine, et al.
Published: (2025)
by: Farhat, Nessrine, et al.
Published: (2025)
"I Can See Forever!": Evaluating Real-time VideoLLMs for Assisting Individuals with Visual Impairments
by: Zhang, Ziyi, et al.
Published: (2025)
by: Zhang, Ziyi, et al.
Published: (2025)
A Picture is Worth a Thousand (Correct) Captions: A Vision-Guided Judge-Corrector System for Multimodal Machine Translation
by: Betala, Siddharth, et al.
Published: (2025)
by: Betala, Siddharth, et al.
Published: (2025)
Learning Spatio-Temporal Feature Representations for Video-Based Gaze Estimation
by: Personnic, Alexandre, et al.
Published: (2025)
by: Personnic, Alexandre, et al.
Published: (2025)
See-Control: A Multimodal Agent Framework for Smartphone Interaction with a Robotic Arm
by: Zhao, Haoyu, et al.
Published: (2025)
by: Zhao, Haoyu, et al.
Published: (2025)
GarmentLab: A Unified Simulation and Benchmark for Garment Manipulation
by: Lu, Haoran, et al.
Published: (2024)
by: Lu, Haoran, et al.
Published: (2024)
Simulating Clinical AI Assistance using Multimodal LLMs: A Case Study in Diabetic Retinopathy
by: Barakat, Nadim, et al.
Published: (2025)
by: Barakat, Nadim, et al.
Published: (2025)
Quantitative Movement Testing: Measuring Patient Movements from a Single Smartphone Video
by: Mahajan, Pranav, et al.
Published: (2026)
by: Mahajan, Pranav, et al.
Published: (2026)
PedaCo-Gen: Scaffolding Pedagogical Agency in Human-AI Collaborative Video Authoring
by: Baek, Injun, et al.
Published: (2026)
by: Baek, Injun, et al.
Published: (2026)
Enhancing Online Learning by Integrating Biosensors and Multimodal Learning Analytics for Detecting and Predicting Student Behavior: A Review
by: Becerra, Alvaro, et al.
Published: (2025)
by: Becerra, Alvaro, et al.
Published: (2025)
Predicting 3D Motion from 2D Video for Behavior-Based VR Biometrics
by: Li, Mingjun, et al.
Published: (2025)
by: Li, Mingjun, et al.
Published: (2025)
Benchmarking XAI Explanations with Human-Aligned Evaluations
by: Kazmierczak, Rémi, et al.
Published: (2024)
by: Kazmierczak, Rémi, et al.
Published: (2024)
ImageTalk: Designing a Multimodal AAC Text Generation System Driven by Image Recognition and Natural Language Generation
by: Yang, Boyin, et al.
Published: (2025)
by: Yang, Boyin, et al.
Published: (2025)
See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding
by: Sun, Boyuan, et al.
Published: (2026)
by: Sun, Boyuan, et al.
Published: (2026)
Similar Items
-
VIBE: Annotation-Free Video-to-Text Information Bottleneck Evaluation for TL;DR
by: Chen, Shenghui, et al.
Published: (2025) -
Human-Agent Coordination in Games under Incomplete Information via Multi-Step Intent
by: Chen, Shenghui, et al.
Published: (2024) -
CSA: Data-efficient Mapping of Unimodal Features to Multimodal Features
by: Li, Po-han, et al.
Published: (2024) -
Human-Agent Cooperation in Games under Incomplete Information through Natural Language Communication
by: Chen, Shenghui, et al.
Published: (2024) -
Any2Any: Incomplete Multimodal Retrieval with Conformal Prediction
by: Li, Po-han, et al.
Published: (2024)