Optimizing Multimodal LLMs for Egocentric Video Understanding: A Solution for the HD-EPIC VQA Challenge
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Sicheng, Huang, Yukai, Sun, Shitong, Cai, Weitong, Deng, Jiankang, Song, Jifei, Zhang, Zhensong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Plug-and-Play Clarifier: A Zero-Shot Multimodal Framework for Egocentric Intent Disambiguation
by: Yang, Sicheng, et al.
Published: (2025)
by: Yang, Sicheng, et al.
Published: (2025)
Color When It Counts: Grayscale-Guided Online Triggering for Always-On Streaming Video Sensing
by: Cai, Weitong, et al.
Published: (2026)
by: Cai, Weitong, et al.
Published: (2026)
Memory-Anchored Multimodal Reasoning for Explainable Video Forensics
by: Chen, Chen, et al.
Published: (2025)
by: Chen, Chen, et al.
Published: (2025)
CAMP-VQA: Caption-Embedded Multimodal Perception for No-Reference Quality Assessment of Compressed Video
by: Wang, Xinyi, et al.
Published: (2025)
by: Wang, Xinyi, et al.
Published: (2025)
QoE Optimization for Semantic Self-Correcting Video Transmission in Multi-UAV Networks
by: Chen, Xuyang, et al.
Published: (2025)
by: Chen, Xuyang, et al.
Published: (2025)
Towards User-level QoE: Large-scale Practice in Personalized Optimization of Adaptive Video Streaming
by: Jia, Lianchen, et al.
Published: (2025)
by: Jia, Lianchen, et al.
Published: (2025)
HippoMM: Hippocampal-inspired Multimodal Memory for Long Audiovisual Event Understanding
by: Lin, Yueqian, et al.
Published: (2025)
by: Lin, Yueqian, et al.
Published: (2025)
Optimizing Mobile-Friendly Viewport Prediction for Live 360-Degree Video Streaming
by: Zhang, Lei, et al.
Published: (2024)
by: Zhang, Lei, et al.
Published: (2024)
Symmetric Entropy-Constrained Video Coding for Machines
by: Sun, Yuxiao, et al.
Published: (2025)
by: Sun, Yuxiao, et al.
Published: (2025)
DIVA-VQA: Detecting Inter-frame Variations in UGC Video Quality
by: Wang, Xinyi, et al.
Published: (2025)
by: Wang, Xinyi, et al.
Published: (2025)
Interactive $360^{\circ}$ Video Streaming Using FoV-Adaptive Coding with Temporal Prediction
by: Mao, Yixiang, et al.
Published: (2024)
by: Mao, Yixiang, et al.
Published: (2024)
RoWSFormer: A Robust Watermarking Framework with Swin Transformer for Enhanced Geometric Attack Resilience
by: Chen, Weitong, et al.
Published: (2024)
by: Chen, Weitong, et al.
Published: (2024)
Perception-Aware Video Semantic Communication
by: Huang, Yinhuan, et al.
Published: (2026)
by: Huang, Yinhuan, et al.
Published: (2026)
A Versatile Depth Video Encoding Scheme Based on Low-rank Tensor Modeling for Free Viewpoint Video
by: Sharma, Mansi, et al.
Published: (2021)
by: Sharma, Mansi, et al.
Published: (2021)
Prompt-based Multimodal Semantic Communication for Multi-spectral Image Segmentation
by: Zhang, Haoshuo, et al.
Published: (2025)
by: Zhang, Haoshuo, et al.
Published: (2025)
Tube-Structured Incremental Semantic HARQ for Generative Video Receivers
by: Wang, Xuesong, et al.
Published: (2026)
by: Wang, Xuesong, et al.
Published: (2026)
Fast Multirate Encoding for 360° Video in OMAF Streaming Workflows
by: Premkumar, Amritha, et al.
Published: (2026)
by: Premkumar, Amritha, et al.
Published: (2026)
Adaptive Resolution and Chroma Subsampling for Energy-Efficient Video Coding
by: Premkumar, Amritha, et al.
Published: (2026)
by: Premkumar, Amritha, et al.
Published: (2026)
Convex-hull Estimation using XPSNR for Versatile Video Coding
by: Menon, Vignesh V, et al.
Published: (2024)
by: Menon, Vignesh V, et al.
Published: (2024)
Efficient Sub-pixel Motion Compensation in Learned Video Codecs
by: Ladune, Théo, et al.
Published: (2025)
by: Ladune, Théo, et al.
Published: (2025)
ReLaX-VQA: Residual Fragment and Layer Stack Extraction for Enhancing Video Quality Assessment
by: Wang, Xinyi, et al.
Published: (2024)
by: Wang, Xinyi, et al.
Published: (2024)
Rate-Quality or Energy-Quality Pareto Fronts for Adaptive Video Streaming?
by: Katsenou, Angeliki, et al.
Published: (2024)
by: Katsenou, Angeliki, et al.
Published: (2024)
Smaller is Better: Generative Models Can Power Short Video Preloading
by: Liu, Liming, et al.
Published: (2026)
by: Liu, Liming, et al.
Published: (2026)
Audio-Visual Cross-Modal Compression for Generative Face Video Coding
by: Xu, Youmin, et al.
Published: (2025)
by: Xu, Youmin, et al.
Published: (2025)
Rip Current Detection in Nearshore Areas through UAV Video Analysis with Almost Local-Isometric Embedding Techniques on Sphere
by: Sun, Anchen, et al.
Published: (2023)
by: Sun, Anchen, et al.
Published: (2023)
Encoding Time and Energy Model for SVT-AV1 based on Video Complexity
by: Eichermüller, Lena, et al.
Published: (2024)
by: Eichermüller, Lena, et al.
Published: (2024)
Neural Compression of 360-Degree Equirectangular Videos using Quality Parameter Adaptation
by: Arai, Daichi, et al.
Published: (2025)
by: Arai, Daichi, et al.
Published: (2025)
Progressive Frame Patching for FoV-based Point Cloud Video Streaming
by: Zong, Tongyu, et al.
Published: (2023)
by: Zong, Tongyu, et al.
Published: (2023)
Content-Driven Frame-Level Bit Prediction for Rate Control in Versatile Video Coding
by: Premkumar, Amritha, et al.
Published: (2026)
by: Premkumar, Amritha, et al.
Published: (2026)
DiV-INR: Extreme Low-Bitrate Diffusion Video Compression with INR Conditioning
by: Çetin, Eren, et al.
Published: (2026)
by: Çetin, Eren, et al.
Published: (2026)
Generative Flow Networks for Personalized Multimedia Systems: A Case Study on Short Video Feeds
by: Jin, Yili, et al.
Published: (2025)
by: Jin, Yili, et al.
Published: (2025)
H.265/HEVC Video Steganalysis Based on CU Block Structure Gradients and IPM Mapping
by: Zhang, Xiang, et al.
Published: (2026)
by: Zhang, Xiang, et al.
Published: (2026)
DQ-Ladder: A Deep Reinforcement Learning-based Bitrate Ladder for Adaptive Video Streaming
by: Farahani, Reza, et al.
Published: (2026)
by: Farahani, Reza, et al.
Published: (2026)
Camel: Frame-Level Bandwidth Estimation for Low-Latency Live Streaming under Video Bitrate Undershooting
by: Liu, Liming, et al.
Published: (2026)
by: Liu, Liming, et al.
Published: (2026)
Video Compression Beyond VVC: Quantitative Analysis of Intra Coding Tools in Enhanced Compression Model (ECM)
by: Abdoli, Mohsen, et al.
Published: (2024)
by: Abdoli, Mohsen, et al.
Published: (2024)
EvEnhancer: Empowering Effectiveness, Efficiency and Generalizability for Continuous Space-Time Video Super-Resolution with Events
by: Wei, Shuoyan, et al.
Published: (2025)
by: Wei, Shuoyan, et al.
Published: (2025)
A H.265/HEVC Fine-Grained ROI Video Encryption Algorithm Based on Coding Unit and Prompt Segmentation
by: Zhang, Xiang, et al.
Published: (2026)
by: Zhang, Xiang, et al.
Published: (2026)
Subjective and Objective Quality-of-Experience Evaluation Study for Live Video Streaming
by: Zhu, Zehao, et al.
Published: (2024)
by: Zhu, Zehao, et al.
Published: (2024)
Unravelling the Power of Single-Pass Look-Ahead in Modern Codecs for Optimized Transcoding Deployment
by: Vibhoothi, Vibhoothi, et al.
Published: (2024)
by: Vibhoothi, Vibhoothi, et al.
Published: (2024)
Learning Perceptual Representations for Gaming NR-VQA with Multi-Task FR Signals
by: Chen, Yu-Chih, et al.
Published: (2026)
by: Chen, Yu-Chih, et al.
Published: (2026)
Similar Items
-
Plug-and-Play Clarifier: A Zero-Shot Multimodal Framework for Egocentric Intent Disambiguation
by: Yang, Sicheng, et al.
Published: (2025) -
Color When It Counts: Grayscale-Guided Online Triggering for Always-On Streaming Video Sensing
by: Cai, Weitong, et al.
Published: (2026) -
Memory-Anchored Multimodal Reasoning for Explainable Video Forensics
by: Chen, Chen, et al.
Published: (2025) -
CAMP-VQA: Caption-Embedded Multimodal Perception for No-Reference Quality Assessment of Compressed Video
by: Wang, Xinyi, et al.
Published: (2025) -
QoE Optimization for Semantic Self-Correcting Video Transmission in Multi-UAV Networks
by: Chen, Xuyang, et al.
Published: (2025)