R^3: Composed Video Retrieval via Reasoning-Guided Recalling and Re-ranking
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Li, Zixu, Hu, Yupeng, Fu, Zhiheng, Chen, Zhiwei, Guan, Weili, Nie, Liqiang |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
HUD: Hierarchical Uncertainty-Aware Disambiguation Network for Composed Video Retrieval
par: Chen, Zhiwei, et autres
Publié: (2025)
par: Chen, Zhiwei, et autres
Publié: (2025)
OFFSET: Segmentation-based Focus Shift Revision for Composed Image Retrieval
par: Chen, Zhiwei, et autres
Publié: (2025)
par: Chen, Zhiwei, et autres
Publié: (2025)
TEMA: Anchor the Image, Follow the Text for Multi-Modification Composed Image Retrieval
par: Li, Zixu, et autres
Publié: (2026)
par: Li, Zixu, et autres
Publié: (2026)
OmniEgo-R$^2$: A Routed Reasoning Framework for the 1st Cross-Domain EgoCross Challenge at CVPR 2026
par: Li, Zixu, et autres
Publié: (2026)
par: Li, Zixu, et autres
Publié: (2026)
ConeSep: Cone-based Robust Noise-Unlearning Compositional Network for Composed Image Retrieval
par: Li, Zixu, et autres
Publié: (2026)
par: Li, Zixu, et autres
Publié: (2026)
FineCIR: Explicit Parsing of Fine-Grained Modification Semantics for Composed Image Retrieval
par: Li, Zixu, et autres
Publié: (2025)
par: Li, Zixu, et autres
Publié: (2025)
TempRet: Temporal Enhancement and Two-Stage Reranking for CVPR 2026 EPIC-KITCHENS-100 Multi-Instance Retrieval Challenge
par: Li, Zixu, et autres
Publié: (2026)
par: Li, Zixu, et autres
Publié: (2026)
ReTrack: Evidence-Driven Dual-Stream Directional Anchor Calibration Network for Composed Video Retrieval
par: Li, Zixu, et autres
Publié: (2026)
par: Li, Zixu, et autres
Publié: (2026)
EgoAdapt: A Multi-Scene Egocentric Adaptation Method for CVPR 2026 HD-EPIC VQA Challenge
par: Chen, Zhiwei, et autres
Publié: (2026)
par: Chen, Zhiwei, et autres
Publié: (2026)
EgoAction: Egocentric Action Composition with Reliability-Aware Temporal Fusion for the EPIC-KITCHENS Action Detection Challenge at CVPR 2026
par: Fu, Zhiheng, et autres
Publié: (2026)
par: Fu, Zhiheng, et autres
Publié: (2026)
Air-Know: Arbiter-Calibrated Knowledge-Internalizing Robust Network for Composed Image Retrieval
par: Fu, Zhiheng, et autres
Publié: (2026)
par: Fu, Zhiheng, et autres
Publié: (2026)
HINT: Composed Image Retrieval with Dual-path Compositional Contextualized Network
par: Zhang, Mingyu, et autres
Publié: (2026)
par: Zhang, Mingyu, et autres
Publié: (2026)
INTENT: Invariance and Discrimination-aware Noise Mitigation for Robust Composed Image Retrieval
par: Chen, Zhiwei, et autres
Publié: (2026)
par: Chen, Zhiwei, et autres
Publié: (2026)
HABIT: Chrono-Synergia Robust Progressive Learning Framework for Composed Image Retrieval
par: Li, Zixu, et autres
Publié: (2026)
par: Li, Zixu, et autres
Publié: (2026)
MELT: Improve Composed Image Retrieval via the Modification Frequentation-Rarity Balance Network
par: Qiu, Guozhi, et autres
Publié: (2026)
par: Qiu, Guozhi, et autres
Publié: (2026)
StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval
par: Wang, Shaokun, et autres
Publié: (2026)
par: Wang, Shaokun, et autres
Publié: (2026)
UniCVR: From Alignment to Reranking for Unified Zero-Shot Composed Visual Retrieval
par: Wen, Haokun, et autres
Publié: (2026)
par: Wen, Haokun, et autres
Publié: (2026)
Fine-grained Textual Inversion Network for Zero-Shot Composed Image Retrieval
par: Lin, Haoqiang, et autres
Publié: (2025)
par: Lin, Haoqiang, et autres
Publié: (2025)
Reason, Retrieve, Re-rank: A Zero-Shot Reasoning-Aware Framework for Composed Video Retrieval
par: Alavi, Ali
Publié: (2026)
par: Alavi, Ali
Publié: (2026)
Uncovering Hidden Connections: Iterative Search and Reasoning for Video-grounded Dialog
par: Zhang, Haoyu, et autres
Publié: (2023)
par: Zhang, Haoyu, et autres
Publié: (2023)
CoVR-R:Reason-Aware Composed Video Retrieval
par: Thawakar, Omkar, et autres
Publié: (2026)
par: Thawakar, Omkar, et autres
Publié: (2026)
Advancing Complex Video Object Segmentation via Tracking-Enhanced Prompt: The 1st Winner for 5th PVUW MOSE Challenge
par: Zhang, Jinrong, et autres
Publié: (2026)
par: Zhang, Jinrong, et autres
Publié: (2026)
PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning
par: Lyu, Yibo, et autres
Publié: (2025)
par: Lyu, Yibo, et autres
Publié: (2025)
Object-Shot Enhanced Grounding Network for Egocentric Video
par: Feng, Yisen, et autres
Publié: (2025)
par: Feng, Yisen, et autres
Publié: (2025)
The 1st Winner for 5th PVUW MeViS-Text Challenge: Strong MLLMs Meet SAM3 for Referring Video Object Segmentation
par: He, Xusheng, et autres
Publié: (2026)
par: He, Xusheng, et autres
Publié: (2026)
Curriculum Coarse-to-Fine Selection for High-IPC Dataset Distillation
par: Chen, Yanda, et autres
Publié: (2025)
par: Chen, Yanda, et autres
Publié: (2025)
MoME: Mixture of Multimodal Experts for Generalist Multimodal Large Language Models
par: Shen, Leyang, et autres
Publié: (2024)
par: Shen, Leyang, et autres
Publié: (2024)
ViSAGE @ NTIRE 2026 Challenge on Video Saliency Prediction
par: Wang, Kun, et autres
Publié: (2026)
par: Wang, Kun, et autres
Publié: (2026)
V-STaR: Benchmarking Video-LLMs on Video Spatio-Temporal Reasoning
par: Cheng, Zixu, et autres
Publié: (2025)
par: Cheng, Zixu, et autres
Publié: (2025)
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers
par: Zhang, Renshan, et autres
Publié: (2025)
par: Zhang, Renshan, et autres
Publié: (2025)
Token-level Correlation-guided Compression for Efficient Multimodal Document Understanding
par: Zhang, Renshan, et autres
Publié: (2024)
par: Zhang, Renshan, et autres
Publié: (2024)
A Comprehensive Survey on Composed Image Retrieval
par: Song, Xuemeng, et autres
Publié: (2025)
par: Song, Xuemeng, et autres
Publié: (2025)
Spatial Understanding from Videos: Structured Prompts Meet Simulation Data
par: Zhang, Haoyu, et autres
Publié: (2025)
par: Zhang, Haoyu, et autres
Publié: (2025)
Detecting Deepfakes via Hamiltonian Dynamics
par: Cheng, Harry, et autres
Publié: (2026)
par: Cheng, Harry, et autres
Publié: (2026)
CoVA: Text-Guided Composed Video Retrieval for Audio-Visual Content
par: Han, Gyuwon, et autres
Publié: (2026)
par: Han, Gyuwon, et autres
Publié: (2026)
Composed Video Retrieval via Enriched Context and Discriminative Embeddings
par: Thawakar, Omkar, et autres
Publié: (2024)
par: Thawakar, Omkar, et autres
Publié: (2024)
Handling Imbalanced Pseudolabels for Vision-Language Models with Concept Alignment and Confusion-Aware Calibrated Margin
par: Wang, Yuchen, et autres
Publié: (2025)
par: Wang, Yuchen, et autres
Publié: (2025)
Reasoning in the Dark: Interleaved Vision-Text Reasoning in Latent Space
par: Chen, Chao, et autres
Publié: (2025)
par: Chen, Chao, et autres
Publié: (2025)
Composed Object Retrieval: Object-level Retrieval via Composed Expressions
par: Wang, Tong, et autres
Publié: (2025)
par: Wang, Tong, et autres
Publié: (2025)
Grounding Video Reasoning in Physical Signals
par: Osmanli, Alibay, et autres
Publié: (2026)
par: Osmanli, Alibay, et autres
Publié: (2026)
Documents similaires
-
HUD: Hierarchical Uncertainty-Aware Disambiguation Network for Composed Video Retrieval
par: Chen, Zhiwei, et autres
Publié: (2025) -
OFFSET: Segmentation-based Focus Shift Revision for Composed Image Retrieval
par: Chen, Zhiwei, et autres
Publié: (2025) -
TEMA: Anchor the Image, Follow the Text for Multi-Modification Composed Image Retrieval
par: Li, Zixu, et autres
Publié: (2026) -
OmniEgo-R$^2$: A Routed Reasoning Framework for the 1st Cross-Domain EgoCross Challenge at CVPR 2026
par: Li, Zixu, et autres
Publié: (2026) -
ConeSep: Cone-based Robust Noise-Unlearning Compositional Network for Composed Image Retrieval
par: Li, Zixu, et autres
Publié: (2026)