EgoCVR: An Egocentric Benchmark for Fine-Grained Composed Video Retrieval
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hummel, Thomas, Karthik, Shyamgopal, Georgescu, Mariana-Iuliana, Akata, Zeynep |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Road Obstacle Video Segmentation
von: Rai, Shyam Nandan, et al.
Veröffentlicht: (2025)
von: Rai, Shyam Nandan, et al.
Veröffentlicht: (2025)
X-Aligner: Composed Visual Retrieval without the Bells and Whistles
von: Zheng, Yuqian, et al.
Veröffentlicht: (2026)
von: Zheng, Yuqian, et al.
Veröffentlicht: (2026)
Vision-by-Language for Training-Free Compositional Image Retrieval
von: Karthik, Shyamgopal, et al.
Veröffentlicht: (2023)
von: Karthik, Shyamgopal, et al.
Veröffentlicht: (2023)
FLAIR: VLM with Fine-grained Language-informed Image Representations
von: Xiao, Rui, et al.
Veröffentlicht: (2024)
von: Xiao, Rui, et al.
Veröffentlicht: (2024)
ReNO: Enhancing One-step Text-to-Image Models through Reward-based Noise Optimization
von: Eyring, Luca, et al.
Veröffentlicht: (2024)
von: Eyring, Luca, et al.
Veröffentlicht: (2024)
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training
von: Kim, Sanghwan, et al.
Veröffentlicht: (2024)
von: Kim, Sanghwan, et al.
Veröffentlicht: (2024)
Noise Hypernetworks: Amortizing Test-Time Compute in Diffusion Models
von: Eyring, Luca, et al.
Veröffentlicht: (2025)
von: Eyring, Luca, et al.
Veröffentlicht: (2025)
Scalable Ranked Preference Optimization for Text-to-Image Generation
von: Karthik, Shyamgopal, et al.
Veröffentlicht: (2024)
von: Karthik, Shyamgopal, et al.
Veröffentlicht: (2024)
Sparse Autoencoders Learn Monosemantic Features in Vision-Language Models
von: Pach, Mateusz, et al.
Veröffentlicht: (2025)
von: Pach, Mateusz, et al.
Veröffentlicht: (2025)
Concept-Guided Interpretability via Neural Chunking
von: Wu, Shuchen, et al.
Veröffentlicht: (2025)
von: Wu, Shuchen, et al.
Veröffentlicht: (2025)
EgoSound: Benchmarking Sound Understanding in Egocentric Videos
von: Zhu, Bingwen, et al.
Veröffentlicht: (2026)
von: Zhu, Bingwen, et al.
Veröffentlicht: (2026)
UniCVR: From Alignment to Reranking for Unified Zero-Shot Composed Visual Retrieval
von: Wen, Haokun, et al.
Veröffentlicht: (2026)
von: Wen, Haokun, et al.
Veröffentlicht: (2026)
From Play to Replay: Composed Video Retrieval for Temporally Fine-Grained Videos
von: Gupta, Animesh, et al.
Veröffentlicht: (2025)
von: Gupta, Animesh, et al.
Veröffentlicht: (2025)
Estimating Ego-Body Pose from Doubly Sparse Egocentric Video Data
von: Chi, Seunggeun, et al.
Veröffentlicht: (2024)
von: Chi, Seunggeun, et al.
Veröffentlicht: (2024)
Post-hoc Probabilistic Vision-Language Models
von: Baumann, Anton, et al.
Veröffentlicht: (2024)
von: Baumann, Anton, et al.
Veröffentlicht: (2024)
Q-Former Autoencoder: A Modern Framework for Medical Anomaly Detection
von: Dalmonte, Francesco, et al.
Veröffentlicht: (2025)
von: Dalmonte, Francesco, et al.
Veröffentlicht: (2025)
LaDe: Unified Multi-Layered Graphic Media Generation and Decomposition
von: Lungu-Stan, Vlad-Constantin, et al.
Veröffentlicht: (2026)
von: Lungu-Stan, Vlad-Constantin, et al.
Veröffentlicht: (2026)
TeleEgo: Benchmarking Egocentric AI Assistants in the Wild
von: Yan, Jiaqi, et al.
Veröffentlicht: (2025)
von: Yan, Jiaqi, et al.
Veröffentlicht: (2025)
EgoPoints: Advancing Point Tracking for Egocentric Videos
von: Darkhalil, Ahmad, et al.
Veröffentlicht: (2024)
von: Darkhalil, Ahmad, et al.
Veröffentlicht: (2024)
EgoToM: Benchmarking Theory of Mind Reasoning from Egocentric Videos
von: Li, Yuxuan, et al.
Veröffentlicht: (2025)
von: Li, Yuxuan, et al.
Veröffentlicht: (2025)
Subspace-Boosted Model Merging
von: Skorobogat, Ronald, et al.
Veröffentlicht: (2025)
von: Skorobogat, Ronald, et al.
Veröffentlicht: (2025)
Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos
von: Plizzari, Chiara, et al.
Veröffentlicht: (2025)
von: Plizzari, Chiara, et al.
Veröffentlicht: (2025)
EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation
von: Pei, Baoqi, et al.
Veröffentlicht: (2024)
von: Pei, Baoqi, et al.
Veröffentlicht: (2024)
EgoVITA: Learning to Plan and Verify for Egocentric Video Reasoning
von: Kulkarni, Yogesh, et al.
Veröffentlicht: (2025)
von: Kulkarni, Yogesh, et al.
Veröffentlicht: (2025)
EgoLCD: Egocentric Video Generation with Long Context Diffusion
von: Zhang, Liuzhou, et al.
Veröffentlicht: (2025)
von: Zhang, Liuzhou, et al.
Veröffentlicht: (2025)
EgoGraph: Temporal Knowledge Graph for Egocentric Video Understanding
von: Sun, Shitong, et al.
Veröffentlicht: (2026)
von: Sun, Shitong, et al.
Veröffentlicht: (2026)
EgoPro-Bench: Benchmarking Personalized Proactive Interaction in Egocentric Video Streams
von: Ran, Dongchuan, et al.
Veröffentlicht: (2026)
von: Ran, Dongchuan, et al.
Veröffentlicht: (2026)
EgoEsportsQA: An Egocentric Video Benchmark for Perception and Reasoning in Esports
von: Ma, Jianzhe, et al.
Veröffentlicht: (2026)
von: Ma, Jianzhe, et al.
Veröffentlicht: (2026)
EgoInstruct: An Egocentric Video Dataset of Face-to-face Instructional Interactions with Multi-modal LLM Benchmarking
von: Sakai, Yuki, et al.
Veröffentlicht: (2025)
von: Sakai, Yuki, et al.
Veröffentlicht: (2025)
EgoX: Egocentric Video Generation from a Single Exocentric Video
von: Kang, Taewoong, et al.
Veröffentlicht: (2025)
von: Kang, Taewoong, et al.
Veröffentlicht: (2025)
Ego2Web: A Web Agent Benchmark Grounded in Egocentric Videos
von: Yu, Shoubin, et al.
Veröffentlicht: (2026)
von: Yu, Shoubin, et al.
Veröffentlicht: (2026)
EgoInteract: Synthetic Egocentric Videos Generation for Interaction Understanding and Anticipation
von: Leonardi, Rosario, et al.
Veröffentlicht: (2026)
von: Leonardi, Rosario, et al.
Veröffentlicht: (2026)
EgoEdit: Dataset, Real-Time Streaming Model, and Benchmark for Egocentric Video Editing
von: Li, Runjia, et al.
Veröffentlicht: (2025)
von: Li, Runjia, et al.
Veröffentlicht: (2025)
Heterogeneous Uncertainty-Guided Composed Image Retrieval with Fine-Grained Probabilistic Learning
von: Tang, Haomiao, et al.
Veröffentlicht: (2026)
von: Tang, Haomiao, et al.
Veröffentlicht: (2026)
EgoMimic: Scaling Imitation Learning via Egocentric Video
von: Kareer, Simar, et al.
Veröffentlicht: (2024)
von: Kareer, Simar, et al.
Veröffentlicht: (2024)
Minerva-Ego: Spatiotemporal Hints for Egocentric Video Understanding
von: Nagrani, Arsha, et al.
Veröffentlicht: (2026)
von: Nagrani, Arsha, et al.
Veröffentlicht: (2026)
EgoVLM: Policy Optimization for Egocentric Video Understanding
von: Vinod, Ashwin, et al.
Veröffentlicht: (2025)
von: Vinod, Ashwin, et al.
Veröffentlicht: (2025)
Modeling Fine-Grained Hand-Object Dynamics for Egocentric Video Representation Learning
von: Pei, Baoqi, et al.
Veröffentlicht: (2025)
von: Pei, Baoqi, et al.
Veröffentlicht: (2025)
EgoIntrospect: An Egocentric Dataset and Benchmark for User-Centric Internal State Reasoning
von: Wang, Zeyu, et al.
Veröffentlicht: (2026)
von: Wang, Zeyu, et al.
Veröffentlicht: (2026)
EgoIntent: An Egocentric Step-level Benchmark for Understanding What, Why, and Next
von: Pan, Ye, et al.
Veröffentlicht: (2026)
von: Pan, Ye, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Road Obstacle Video Segmentation
von: Rai, Shyam Nandan, et al.
Veröffentlicht: (2025) -
X-Aligner: Composed Visual Retrieval without the Bells and Whistles
von: Zheng, Yuqian, et al.
Veröffentlicht: (2026) -
Vision-by-Language for Training-Free Compositional Image Retrieval
von: Karthik, Shyamgopal, et al.
Veröffentlicht: (2023) -
FLAIR: VLM with Fine-grained Language-informed Image Representations
von: Xiao, Rui, et al.
Veröffentlicht: (2024) -
ReNO: Enhancing One-step Text-to-Image Models through Reward-based Noise Optimization
von: Eyring, Luca, et al.
Veröffentlicht: (2024)