Unsupervised Ego- and Exo-centric Dense Procedural Activity Captioning via Gaze Consensus Adaptation
Fuente:
arXiv
Saved in:
| Main Authors: | Shi, Zhaofeng, Qiu, Heqian, Wang, Lanxiao, Wu, Qingbo, Meng, Fanman, Li, Hongliang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Test-time Ego-Exo-centric Adaptation for Action Anticipation via Multi-Label Prototype Growing and Dual-Clue Consistency
by: Shi, Zhaofeng, et al.
Published: (2026)
by: Shi, Zhaofeng, et al.
Published: (2026)
Cognition Transferring and Decoupling for Text-supervised Egocentric Semantic Segmentation
by: Shi, Zhaofeng, et al.
Published: (2024)
by: Shi, Zhaofeng, et al.
Published: (2024)
SAVA-X: Ego-to-Exo Imitation Error Detection via Scene-Adaptive View Alignment and Bidirectional Cross View Fusion
by: Li, Xiang, et al.
Published: (2026)
by: Li, Xiang, et al.
Published: (2026)
Leveraging Pre-Trained Models for Multimodal Class-Incremental Learning under Adaptive Fusion
by: Chen, Yukun, et al.
Published: (2025)
by: Chen, Yukun, et al.
Published: (2025)
EgoMe: A New Dataset and Challenge for Following Me via Egocentric View in Real World
by: Qiu, Heqian, et al.
Published: (2025)
by: Qiu, Heqian, et al.
Published: (2025)
Attention-disentangled Uniform Orthogonal Feature Space Optimization for Few-shot Object Detection
by: Zhao, Taijin, et al.
Published: (2025)
by: Zhao, Taijin, et al.
Published: (2025)
Cross-modal Cognitive Consensus guided Audio-Visual Segmentation
by: Shi, Zhaofeng, et al.
Published: (2023)
by: Shi, Zhaofeng, et al.
Published: (2023)
Ego3DT: Tracking Every 3D Object in Ego-centric Videos
by: Hao, Shengyu, et al.
Published: (2024)
by: Hao, Shengyu, et al.
Published: (2024)
Distribution-Level Memory Recall for Continual Learning: Preserving Knowledge and Avoiding Confusion
by: Cheng, Shaoxu, et al.
Published: (2024)
by: Cheng, Shaoxu, et al.
Published: (2024)
MCF-VC: Mitigate Catastrophic Forgetting in Class-Incremental Learning for Multimodal Video Captioning
by: Xiong, Huiyu, et al.
Published: (2024)
by: Xiong, Huiyu, et al.
Published: (2024)
Multi Agents Semantic Emotion Aligned Music to Image Generation with Music Derived Captions
by: Shi, Junchang, et al.
Published: (2025)
by: Shi, Junchang, et al.
Published: (2025)
EgoExoLearn: A Dataset for Bridging Asynchronous Ego- and Exo-centric View of Procedural Activities in Real World
by: Huang, Yifei, et al.
Published: (2024)
by: Huang, Yifei, et al.
Published: (2024)
DIBS: Enhancing Dense Video Captioning with Unlabeled Videos via Pseudo Boundary Enrichment and Online Refinement
by: Wu, Hao, et al.
Published: (2024)
by: Wu, Hao, et al.
Published: (2024)
GRSDet: Learning to Generate Local Reverse Samples for Few-shot Object Detection
by: Mei, Hefei, et al.
Published: (2023)
by: Mei, Hefei, et al.
Published: (2023)
Hallucination Localization in Video Captioning
by: Nakada, Shota, et al.
Published: (2025)
by: Nakada, Shota, et al.
Published: (2025)
EgoBlind: Towards Egocentric Visual Assistance for the Blind
by: Xiao, Junbin, et al.
Published: (2025)
by: Xiao, Junbin, et al.
Published: (2025)
Implicit Location-Caption Alignment via Complementary Masking for Weakly-Supervised Dense Video Captioning
by: Ge, Shiping, et al.
Published: (2024)
by: Ge, Shiping, et al.
Published: (2024)
Cap2Sum: Learning to Summarize Videos by Generating Captions
by: Zhao, Cairong, et al.
Published: (2024)
by: Zhao, Cairong, et al.
Published: (2024)
EyEar: Learning Audio Synchronized Human Gaze Trajectory Based on Physics-Informed Dynamics
by: Liu, Xiaochuan, et al.
Published: (2025)
by: Liu, Xiaochuan, et al.
Published: (2025)
DetailVerifyBench: A Benchmark for Dense Hallucination Localization in Long Image Captions
by: Wang, Xinran, et al.
Published: (2026)
by: Wang, Xinran, et al.
Published: (2026)
Universal Organizer of SAM for Unsupervised Semantic Segmentation
by: Li, Tingting, et al.
Published: (2024)
by: Li, Tingting, et al.
Published: (2024)
Exo2EgoDVC: Dense Video Captioning of Egocentric Procedural Activities Using Web Instructional Videos
by: Ohkawa, Takehiko, et al.
Published: (2023)
by: Ohkawa, Takehiko, et al.
Published: (2023)
EgoExo-Gen: Ego-centric Video Prediction by Watching Exo-centric Videos
by: Xu, Jilan, et al.
Published: (2025)
by: Xu, Jilan, et al.
Published: (2025)
EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering
by: Zhou, Sheng, et al.
Published: (2025)
by: Zhou, Sheng, et al.
Published: (2025)
Embedded Heterogeneous Attention Transformer for Cross-lingual Image Captioning
by: Song, Zijie, et al.
Published: (2023)
by: Song, Zijie, et al.
Published: (2023)
PolySmart @ TRECVid 2024 Video Captioning (VTT)
by: Wu, Jiaxin, et al.
Published: (2024)
by: Wu, Jiaxin, et al.
Published: (2024)
Voices, Faces, and Feelings: Multi-modal Emotion-Cognition Captioning for Mental Health Understanding
by: Zhou, Zhiyuan, et al.
Published: (2026)
by: Zhou, Zhiyuan, et al.
Published: (2026)
NewsCaption: Named-Entity aware Captioning for Out-of-Context Media
by: Singh, Anurag, et al.
Published: (2024)
by: Singh, Anurag, et al.
Published: (2024)
EgoForge: Goal-Directed Egocentric World Simulator
by: Shen, Yifan, et al.
Published: (2026)
by: Shen, Yifan, et al.
Published: (2026)
Mitigating Image Captioning Hallucinations in Vision-Language Models
by: Zhao, Fei, et al.
Published: (2025)
by: Zhao, Fei, et al.
Published: (2025)
UPDA: Unsupervised Progressive Domain Adaptation for No-Reference Point Cloud Quality Assessment
by: Xie, Bingxu, et al.
Published: (2026)
by: Xie, Bingxu, et al.
Published: (2026)
Semantic-Guided Unsupervised Video Summarization
by: Liu, Haizhou, et al.
Published: (2026)
by: Liu, Haizhou, et al.
Published: (2026)
EyeNexus: Adaptive Gaze-Driven Quality and Bitrate Streaming for Seamless VR Cloud Gaming Experiences
by: Wu, Ze, et al.
Published: (2025)
by: Wu, Ze, et al.
Published: (2025)
Towards Generating Diverse Audio Captions via Adversarial Training
by: Mei, Xinhao, et al.
Published: (2022)
by: Mei, Xinhao, et al.
Published: (2022)
Sound-VECaps: Improving Audio Generation with Visual Enhanced Captions
by: Yuan, Yi, et al.
Published: (2024)
by: Yuan, Yi, et al.
Published: (2024)
Group-based Distinctive Image Captioning with Memory Difference Encoding and Attention
by: Wang, Jiuniu, et al.
Published: (2025)
by: Wang, Jiuniu, et al.
Published: (2025)
Layer-wise Model Merging for Unsupervised Domain Adaptation in Segmentation Tasks
by: Alcover-Couso, Roberto, et al.
Published: (2024)
by: Alcover-Couso, Roberto, et al.
Published: (2024)
DesCLIP: Robust Continual Learning via General Attribute Descriptions for VLM-Based Visual Recognition
by: He, Chiyuan, et al.
Published: (2025)
by: He, Chiyuan, et al.
Published: (2025)
Robust Live Streaming over LEO Satellite Constellations: Measurement, Analysis, and Handover-Aware Adaptation
by: Fang, Hao, et al.
Published: (2025)
by: Fang, Hao, et al.
Published: (2025)
Modulating CNN Features with Pre-Trained ViT Representations for Open-Vocabulary Object Detection
by: Gao, Xiangyu, et al.
Published: (2025)
by: Gao, Xiangyu, et al.
Published: (2025)
Similar Items
-
Test-time Ego-Exo-centric Adaptation for Action Anticipation via Multi-Label Prototype Growing and Dual-Clue Consistency
by: Shi, Zhaofeng, et al.
Published: (2026) -
Cognition Transferring and Decoupling for Text-supervised Egocentric Semantic Segmentation
by: Shi, Zhaofeng, et al.
Published: (2024) -
SAVA-X: Ego-to-Exo Imitation Error Detection via Scene-Adaptive View Alignment and Bidirectional Cross View Fusion
by: Li, Xiang, et al.
Published: (2026) -
Leveraging Pre-Trained Models for Multimodal Class-Incremental Learning under Adaptive Fusion
by: Chen, Yukun, et al.
Published: (2025) -
EgoMe: A New Dataset and Challenge for Following Me via Egocentric View in Real World
by: Qiu, Heqian, et al.
Published: (2025)