Leveraging Modality Tags for Enhanced Cross-Modal Video Retrieval
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fragomeni, Adriano, Damen, Dima, Wray, Michael |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Leveraging Auxiliary Information in Text-to-Video Retrieval: A Review
von: Fragomeni, Adriano, et al.
Veröffentlicht: (2025)
von: Fragomeni, Adriano, et al.
Veröffentlicht: (2025)
Video Editing for Video Retrieval
von: Zhu, Bin, et al.
Veröffentlicht: (2024)
von: Zhu, Bin, et al.
Veröffentlicht: (2024)
Moment of Untruth: Dealing with Negative Queries in Video Moment Retrieval
von: Flanagan, Kevin, et al.
Veröffentlicht: (2025)
von: Flanagan, Kevin, et al.
Veröffentlicht: (2025)
Beyond Caption-Based Queries for Video Moment Retrieval
von: Pujol-Perich, David, et al.
Veröffentlicht: (2026)
von: Pujol-Perich, David, et al.
Veröffentlicht: (2026)
HOI-Ref: Hand-Object Interaction Referral in Egocentric Vision
von: Bansal, Siddhant, et al.
Veröffentlicht: (2024)
von: Bansal, Siddhant, et al.
Veröffentlicht: (2024)
GenHowTo: Learning to Generate Actions and State Transformations from Instructional Videos
von: Souček, Tomáš, et al.
Veröffentlicht: (2023)
von: Souček, Tomáš, et al.
Veröffentlicht: (2023)
Get a Grip: Reconstructing Hand-Object Stable Grasps in Egocentric Videos
von: Zhu, Zhifan, et al.
Veröffentlicht: (2023)
von: Zhu, Zhifan, et al.
Veröffentlicht: (2023)
ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions
von: Souček, Tomáš, et al.
Veröffentlicht: (2024)
von: Souček, Tomáš, et al.
Veröffentlicht: (2024)
Reconstructing Objects along Hand Interaction Timelines in Egocentric Video
von: Zhu, Zhifan, et al.
Veröffentlicht: (2025)
von: Zhu, Zhifan, et al.
Veröffentlicht: (2025)
It's Just Another Day: Unique Video Captioning by Discriminative Prompting
von: Perrett, Toby, et al.
Veröffentlicht: (2024)
von: Perrett, Toby, et al.
Veröffentlicht: (2024)
The N-Body Problem: Parallel Execution from Single-Person Egocentric Video
von: Zhu, Zhifan, et al.
Veröffentlicht: (2025)
von: Zhu, Zhifan, et al.
Veröffentlicht: (2025)
Every Shot Counts: Using Exemplars for Repetition Counting in Videos
von: Sinha, Saptarshi, et al.
Veröffentlicht: (2024)
von: Sinha, Saptarshi, et al.
Veröffentlicht: (2024)
EgoPoints: Advancing Point Tracking for Egocentric Videos
von: Darkhalil, Ahmad, et al.
Veröffentlicht: (2024)
von: Darkhalil, Ahmad, et al.
Veröffentlicht: (2024)
COM3D: Leveraging Cross-View Correspondence and Cross-Modal Mining for 3D Retrieval
von: Wu, Hao, et al.
Veröffentlicht: (2024)
von: Wu, Hao, et al.
Veröffentlicht: (2024)
StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval
von: Wang, Shaokun, et al.
Veröffentlicht: (2026)
von: Wang, Shaokun, et al.
Veröffentlicht: (2026)
Do You Remember? Dense Video Captioning with Cross-Modal Memory Retrieval
von: Kim, Minkuk, et al.
Veröffentlicht: (2024)
von: Kim, Minkuk, et al.
Veröffentlicht: (2024)
PointSt3R: Point Tracking through 3D Grounded Correspondence
von: Guerrier, Rhodri, et al.
Veröffentlicht: (2025)
von: Guerrier, Rhodri, et al.
Veröffentlicht: (2025)
CAR-MFL: Cross-Modal Augmentation by Retrieval for Multimodal Federated Learning with Missing Modalities
von: Poudel, Pranav, et al.
Veröffentlicht: (2024)
von: Poudel, Pranav, et al.
Veröffentlicht: (2024)
Prototype-Enhanced Confidence Modeling for Cross-Modal Medical Image-Report Retrieval
von: Gowda, Shreyank N, et al.
Veröffentlicht: (2025)
von: Gowda, Shreyank N, et al.
Veröffentlicht: (2025)
SpatiaLoc: Leveraging Multi-Level Spatial Enhanced Descriptors for Cross-Modal Localization
von: Shang, Tianyi, et al.
Veröffentlicht: (2026)
von: Shang, Tianyi, et al.
Veröffentlicht: (2026)
Spatial Cognition from Egocentric Video: Out of Sight, Not Out of Mind
von: Plizzari, Chiara, et al.
Veröffentlicht: (2024)
von: Plizzari, Chiara, et al.
Veröffentlicht: (2024)
Cross-Modal and Uni-Modal Soft-Label Alignment for Image-Text Retrieval
von: Huang, Hailang, et al.
Veröffentlicht: (2024)
von: Huang, Hailang, et al.
Veröffentlicht: (2024)
The Invisible EgoHand: 3D Hand Forecasting through EgoBody Pose Estimation
von: Hatano, Masashi, et al.
Veröffentlicht: (2025)
von: Hatano, Masashi, et al.
Veröffentlicht: (2025)
AMEGO: Active Memory from long EGOcentric videos
von: Goletto, Gabriele, et al.
Veröffentlicht: (2024)
von: Goletto, Gabriele, et al.
Veröffentlicht: (2024)
Leveraging Multi-Modal Information to Enhance Dataset Distillation
von: Li, Zhe, et al.
Veröffentlicht: (2025)
von: Li, Zhe, et al.
Veröffentlicht: (2025)
TCMA: Text-Conditioned Multi-granularity Alignment for Drone Cross-Modal Text-Video Retrieval
von: Zhao, Zixu, et al.
Veröffentlicht: (2025)
von: Zhao, Zixu, et al.
Veröffentlicht: (2025)
MLLM as Video Narrator: Mitigating Modality Imbalance in Video Moment Retrieval
von: Cai, Weitong, et al.
Veröffentlicht: (2024)
von: Cai, Weitong, et al.
Veröffentlicht: (2024)
Distill CLIP (DCLIP): Enhancing Image-Text Retrieval via Cross-Modal Transformer Distillation
von: Csizmadia, Daniel, et al.
Veröffentlicht: (2025)
von: Csizmadia, Daniel, et al.
Veröffentlicht: (2025)
Segmenting Collision Sound Sources in Egocentric Videos
von: Parida, Kranti Kumar, et al.
Veröffentlicht: (2025)
von: Parida, Kranti Kumar, et al.
Veröffentlicht: (2025)
NeighborRetr: Balancing Hub Centrality in Cross-Modal Retrieval
von: Lin, Zengrong, et al.
Veröffentlicht: (2025)
von: Lin, Zengrong, et al.
Veröffentlicht: (2025)
Distilling Vision-Language Pretraining for Efficient Cross-Modal Retrieval
von: Jang, Young Kyun, et al.
Veröffentlicht: (2024)
von: Jang, Young Kyun, et al.
Veröffentlicht: (2024)
Beyond Cross-Modal Alignment: Measuring and Leveraging Modality Gap in Vision-Language Models
von: Yan, Hanqi, et al.
Veröffentlicht: (2025)
von: Yan, Hanqi, et al.
Veröffentlicht: (2025)
Combating Visual Neglect and Semantic Drift in Large Multimodal Models for Enhanced Cross-Modal Retrieval
von: Zhang, Guosheng, et al.
Veröffentlicht: (2026)
von: Zhang, Guosheng, et al.
Veröffentlicht: (2026)
Video, How Do Your Tokens Merge?
von: Pollard, Sam, et al.
Veröffentlicht: (2025)
von: Pollard, Sam, et al.
Veröffentlicht: (2025)
A Video Is Not Worth a Thousand Words
von: Pollard, Sam, et al.
Veröffentlicht: (2025)
von: Pollard, Sam, et al.
Veröffentlicht: (2025)
Perception Test 2024: Challenge Summary and a Novel Hour-Long VideoQA Benchmark
von: Heyward, Joseph, et al.
Veröffentlicht: (2024)
von: Heyward, Joseph, et al.
Veröffentlicht: (2024)
Learning Modality Knowledge Alignment for Cross-Modality Transfer
von: Ma, Wenxuan, et al.
Veröffentlicht: (2024)
von: Ma, Wenxuan, et al.
Veröffentlicht: (2024)
Referring Video Object Segmentation with Cross-Modality Proxy Queries
von: Sun, Baoli, et al.
Veröffentlicht: (2025)
von: Sun, Baoli, et al.
Veröffentlicht: (2025)
Federated Cross-Modal Retrieval with Missing Modalities via Semantic Routing and Adapter Personalization
von: Zhou, Hefeng, et al.
Veröffentlicht: (2026)
von: Zhou, Hefeng, et al.
Veröffentlicht: (2026)
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval
von: Huang, Jinghao, et al.
Veröffentlicht: (2025)
von: Huang, Jinghao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Leveraging Auxiliary Information in Text-to-Video Retrieval: A Review
von: Fragomeni, Adriano, et al.
Veröffentlicht: (2025) -
Video Editing for Video Retrieval
von: Zhu, Bin, et al.
Veröffentlicht: (2024) -
Moment of Untruth: Dealing with Negative Queries in Video Moment Retrieval
von: Flanagan, Kevin, et al.
Veröffentlicht: (2025) -
Beyond Caption-Based Queries for Video Moment Retrieval
von: Pujol-Perich, David, et al.
Veröffentlicht: (2026) -
HOI-Ref: Hand-Object Interaction Referral in Egocentric Vision
von: Bansal, Siddhant, et al.
Veröffentlicht: (2024)