Retrieval-Augmented Egocentric Video Captioning
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Jilan, Huang, Yifei, Hou, Junlin, Chen, Guo, Zhang, Yuejie, Feng, Rui, Xie, Weidi |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EgoExo-Gen: Ego-centric Video Prediction by Watching Exo-centric Videos
by: Xu, Jilan, et al.
Published: (2025)
by: Xu, Jilan, et al.
Published: (2025)
Modeling Fine-Grained Hand-Object Dynamics for Egocentric Video Representation Learning
by: Pei, Baoqi, et al.
Published: (2025)
by: Pei, Baoqi, et al.
Published: (2025)
QMix: Quality-aware Learning with Mixed Noise for Robust Retinal Disease Diagnosis
by: Hou, Junlin, et al.
Published: (2024)
by: Hou, Junlin, et al.
Published: (2024)
Grounded Question-Answering in Long Egocentric Videos
by: Di, Shangzhe, et al.
Published: (2023)
by: Di, Shangzhe, et al.
Published: (2023)
Domain Adaptation Using Pseudo Labels for COVID-19 Detection
by: Yuan, Runtian, et al.
Published: (2024)
by: Yuan, Runtian, et al.
Published: (2024)
Advancing COVID-19 Detection in 3D CT Scans
by: Li, Qingqiu, et al.
Published: (2024)
by: Li, Qingqiu, et al.
Published: (2024)
Advancing Lung Disease Diagnosis in 3D CT Scans
by: Li, Qingqiu, et al.
Published: (2025)
by: Li, Qingqiu, et al.
Published: (2025)
Multi-Source COVID-19 Detection via Variance Risk Extrapolation
by: Yuan, Runtian, et al.
Published: (2025)
by: Yuan, Runtian, et al.
Published: (2025)
Concept-Attention Whitening for Interpretable Skin Lesion Diagnosis
by: Hou, Junlin, et al.
Published: (2024)
by: Hou, Junlin, et al.
Published: (2024)
Grounded Multi-Hop VideoQA in Long-Form Egocentric Videos
by: Chen, Qirui, et al.
Published: (2024)
by: Chen, Qirui, et al.
Published: (2024)
EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation
by: Pei, Baoqi, et al.
Published: (2024)
by: Pei, Baoqi, et al.
Published: (2024)
Learning Streaming Video Representation via Multitask Training
by: Yan, Yibin, et al.
Published: (2025)
by: Yan, Yibin, et al.
Published: (2025)
Clinical Priors Guided Lung Disease Detection in 3D CT Scans
by: Lu, Kejin, et al.
Published: (2026)
by: Lu, Kejin, et al.
Published: (2026)
Vision-Language Model Based Multi-Expert Fusion for CT Image Classification
by: Bai, Jianfa, et al.
Published: (2026)
by: Bai, Jianfa, et al.
Published: (2026)
EgoThinker: Unveiling Egocentric Reasoning with Spatio-Temporal CoT
by: Pei, Baoqi, et al.
Published: (2025)
by: Pei, Baoqi, et al.
Published: (2025)
Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric Vision
by: He, Yuping, et al.
Published: (2025)
by: He, Yuping, et al.
Published: (2025)
OmniStream: Mastering Perception, Reconstruction and Action in Continuous Streams
by: Yan, Yibin, et al.
Published: (2026)
by: Yan, Yibin, et al.
Published: (2026)
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs
by: He, Yuping, et al.
Published: (2025)
by: He, Yuping, et al.
Published: (2025)
Text-Promptable Propagation for Referring Medical Image Sequence Segmentation
by: Yuan, Runtian, et al.
Published: (2025)
by: Yuan, Runtian, et al.
Published: (2025)
FACE-net: Factual Calibration and Emotion Augmentation for Retrieval-enhanced Emotional Video Captioning
by: Chen, Weidong, et al.
Published: (2026)
by: Chen, Weidong, et al.
Published: (2026)
VQ-Jarvis: Retrieval-Augmented Video Restoration Agent with Sharp Vision and Fast Thought
by: Zhang, Xuanyu, et al.
Published: (2026)
by: Zhang, Xuanyu, et al.
Published: (2026)
Masked Video and Body-worn IMU Autoencoder for Egocentric Action Recognition
by: Zhang, Mingfang, et al.
Published: (2024)
by: Zhang, Mingfang, et al.
Published: (2024)
Video Mamba Suite: State Space Model as a Versatile Alternative for Video Understanding
by: Chen, Guo, et al.
Published: (2024)
by: Chen, Guo, et al.
Published: (2024)
CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
by: Chen, Guo, et al.
Published: (2024)
by: Chen, Guo, et al.
Published: (2024)
Anatomical Structure-Guided Medical Vision-Language Pre-training
by: Li, Qingqiu, et al.
Published: (2024)
by: Li, Qingqiu, et al.
Published: (2024)
The N-Body Problem: Parallel Execution from Single-Person Egocentric Video
by: Zhu, Zhifan, et al.
Published: (2025)
by: Zhu, Zhifan, et al.
Published: (2025)
Understanding Retrieval Robustness for Retrieval-Augmented Image Captioning
by: Li, Wenyan, et al.
Published: (2024)
by: Li, Wenyan, et al.
Published: (2024)
Text-Only Training for Image Captioning with Retrieval Augmentation and Modality Gap Correction
by: Fonseca, Rui, et al.
Published: (2025)
by: Fonseca, Rui, et al.
Published: (2025)
DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding
by: Wu, Hao, et al.
Published: (2024)
by: Wu, Hao, et al.
Published: (2024)
Diffusion Based Augmentation for Captioning and Retrieval in Cultural Heritage
by: Cioni, Dario, et al.
Published: (2023)
by: Cioni, Dario, et al.
Published: (2023)
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model
by: Huang, Yifei, et al.
Published: (2024)
by: Huang, Yifei, et al.
Published: (2024)
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking
by: Meng, Desen, et al.
Published: (2025)
by: Meng, Desen, et al.
Published: (2025)
Enhancing Video-LLM Reasoning via Agent-of-Thoughts Distillation
by: Shi, Yudi, et al.
Published: (2024)
by: Shi, Yudi, et al.
Published: (2024)
Sali4Vid: Saliency-Aware Video Reweighting and Adaptive Caption Retrieval for Dense Video Captioning
by: Jeon, MinJu, et al.
Published: (2025)
by: Jeon, MinJu, et al.
Published: (2025)
Dual Semantic-Aware Network for Noise Suppressed Ultrasound Video Segmentation
by: Zhou, Ling, et al.
Published: (2025)
by: Zhou, Ling, et al.
Published: (2025)
MeaCap: Memory-Augmented Zero-shot Image Captioning
by: Zeng, Zequn, et al.
Published: (2024)
by: Zeng, Zequn, et al.
Published: (2024)
A Sanity Check on Composed Image Retrieval
by: Liu, Yikun, et al.
Published: (2026)
by: Liu, Yikun, et al.
Published: (2026)
Video Enriched Retrieval Augmented Generation Using Aligned Video Captions
by: Rosa, Kevin Dela
Published: (2024)
by: Rosa, Kevin Dela
Published: (2024)
MOSMOS: Multi-organ segmentation facilitated by medical report supervision
by: Tian, Weiwei, et al.
Published: (2024)
by: Tian, Weiwei, et al.
Published: (2024)
Zero-shot Composed Text-Image Retrieval
by: Liu, Yikun, et al.
Published: (2023)
by: Liu, Yikun, et al.
Published: (2023)
Similar Items
-
EgoExo-Gen: Ego-centric Video Prediction by Watching Exo-centric Videos
by: Xu, Jilan, et al.
Published: (2025) -
Modeling Fine-Grained Hand-Object Dynamics for Egocentric Video Representation Learning
by: Pei, Baoqi, et al.
Published: (2025) -
QMix: Quality-aware Learning with Mixed Noise for Robust Retinal Disease Diagnosis
by: Hou, Junlin, et al.
Published: (2024) -
Grounded Question-Answering in Long Egocentric Videos
by: Di, Shangzhe, et al.
Published: (2023) -
Domain Adaptation Using Pseudo Labels for COVID-19 Detection
by: Yuan, Runtian, et al.
Published: (2024)