StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Wang, Shaokun, Guan, Weili, Han, Jizhou, Wu, Jianlong, Hu, Yupeng, Nie, Liqiang |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
R^3: Composed Video Retrieval via Reasoning-Guided Recalling and Re-ranking
par: Li, Zixu, et autres
Publié: (2026)
par: Li, Zixu, et autres
Publié: (2026)
The 1st Winner for 5th PVUW MeViS-Text Challenge: Strong MLLMs Meet SAM3 for Referring Video Object Segmentation
par: He, Xusheng, et autres
Publié: (2026)
par: He, Xusheng, et autres
Publié: (2026)
Advancing Complex Video Object Segmentation via Tracking-Enhanced Prompt: The 1st Winner for 5th PVUW MOSE Challenge
par: Zhang, Jinrong, et autres
Publié: (2026)
par: Zhang, Jinrong, et autres
Publié: (2026)
Self-Enhanced Image Clustering with Cross-Modal Semantic Consistency
par: Li, Zihan, et autres
Publié: (2025)
par: Li, Zihan, et autres
Publié: (2025)
OmniEgo-R$^2$: A Routed Reasoning Framework for the 1st Cross-Domain EgoCross Challenge at CVPR 2026
par: Li, Zixu, et autres
Publié: (2026)
par: Li, Zixu, et autres
Publié: (2026)
TempRet: Temporal Enhancement and Two-Stage Reranking for CVPR 2026 EPIC-KITCHENS-100 Multi-Instance Retrieval Challenge
par: Li, Zixu, et autres
Publié: (2026)
par: Li, Zixu, et autres
Publié: (2026)
UniCVR: From Alignment to Reranking for Unified Zero-Shot Composed Visual Retrieval
par: Wen, Haokun, et autres
Publié: (2026)
par: Wen, Haokun, et autres
Publié: (2026)
PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning
par: Lyu, Yibo, et autres
Publié: (2025)
par: Lyu, Yibo, et autres
Publié: (2025)
Cross-Modal and Uni-Modal Soft-Label Alignment for Image-Text Retrieval
par: Huang, Hailang, et autres
Publié: (2024)
par: Huang, Hailang, et autres
Publié: (2024)
TEMA: Anchor the Image, Follow the Text for Multi-Modification Composed Image Retrieval
par: Li, Zixu, et autres
Publié: (2026)
par: Li, Zixu, et autres
Publié: (2026)
EgoAdapt: A Multi-Scene Egocentric Adaptation Method for CVPR 2026 HD-EPIC VQA Challenge
par: Chen, Zhiwei, et autres
Publié: (2026)
par: Chen, Zhiwei, et autres
Publié: (2026)
EgoAction: Egocentric Action Composition with Reliability-Aware Temporal Fusion for the EPIC-KITCHENS Action Detection Challenge at CVPR 2026
par: Fu, Zhiheng, et autres
Publié: (2026)
par: Fu, Zhiheng, et autres
Publié: (2026)
PersonalAlign: Hierarchical Implicit Intent Alignment for Personalized GUI Agent with Long-Term User-Centric Records
par: Lyu, Yibo, et autres
Publié: (2026)
par: Lyu, Yibo, et autres
Publié: (2026)
HUD: Hierarchical Uncertainty-Aware Disambiguation Network for Composed Video Retrieval
par: Chen, Zhiwei, et autres
Publié: (2025)
par: Chen, Zhiwei, et autres
Publié: (2025)
TCMA: Text-Conditioned Multi-granularity Alignment for Drone Cross-Modal Text-Video Retrieval
par: Zhao, Zixu, et autres
Publié: (2025)
par: Zhao, Zixu, et autres
Publié: (2025)
Object-Shot Enhanced Grounding Network for Egocentric Video
par: Feng, Yisen, et autres
Publié: (2025)
par: Feng, Yisen, et autres
Publié: (2025)
GOAL: Geometrically Optimal Alignment for Continual Generalized Category Discovery
par: Han, Jizhou, et autres
Publié: (2026)
par: Han, Jizhou, et autres
Publié: (2026)
Continuous Knowledge-Preserving Decomposition with Adaptive Layer Selection for Few-Shot Class-Incremental Learning
par: Li, Xiaojie, et autres
Publié: (2025)
par: Li, Xiaojie, et autres
Publié: (2025)
Spatial Understanding from Videos: Structured Prompts Meet Simulation Data
par: Zhang, Haoyu, et autres
Publié: (2025)
par: Zhang, Haoyu, et autres
Publié: (2025)
Handling Imbalanced Pseudolabels for Vision-Language Models with Concept Alignment and Confusion-Aware Calibrated Margin
par: Wang, Yuchen, et autres
Publié: (2025)
par: Wang, Yuchen, et autres
Publié: (2025)
OFFSET: Segmentation-based Focus Shift Revision for Composed Image Retrieval
par: Chen, Zhiwei, et autres
Publié: (2025)
par: Chen, Zhiwei, et autres
Publié: (2025)
MoME: Mixture of Multimodal Experts for Generalist Multimodal Large Language Models
par: Shen, Leyang, et autres
Publié: (2024)
par: Shen, Leyang, et autres
Publié: (2024)
Curriculum Coarse-to-Fine Selection for High-IPC Dataset Distillation
par: Chen, Yanda, et autres
Publié: (2025)
par: Chen, Yanda, et autres
Publié: (2025)
Uncovering Hidden Connections: Iterative Search and Reasoning for Video-grounded Dialog
par: Zhang, Haoyu, et autres
Publié: (2023)
par: Zhang, Haoyu, et autres
Publié: (2023)
ConeSep: Cone-based Robust Noise-Unlearning Compositional Network for Composed Image Retrieval
par: Li, Zixu, et autres
Publié: (2026)
par: Li, Zixu, et autres
Publié: (2026)
ViSAGE @ NTIRE 2026 Challenge on Video Saliency Prediction
par: Wang, Kun, et autres
Publié: (2026)
par: Wang, Kun, et autres
Publié: (2026)
Fine-grained Textual Inversion Network for Zero-Shot Composed Image Retrieval
par: Lin, Haoqiang, et autres
Publié: (2025)
par: Lin, Haoqiang, et autres
Publié: (2025)
Consistent Supervised-Unsupervised Alignment for Generalized Category Discovery
par: Han, Jizhou, et autres
Publié: (2025)
par: Han, Jizhou, et autres
Publié: (2025)
IOTA: Corrective Knowledge-Guided Prompt Learning via Black-White Box Framework
par: Wang, Shaokun, et autres
Publié: (2026)
par: Wang, Shaokun, et autres
Publié: (2026)
Multi-Modal Cross-Domain Alignment Network for Video Moment Retrieval
par: Fang, Xiang, et autres
Publié: (2022)
par: Fang, Xiang, et autres
Publié: (2022)
FineCIR: Explicit Parsing of Fine-Grained Modification Semantics for Composed Image Retrieval
par: Li, Zixu, et autres
Publié: (2025)
par: Li, Zixu, et autres
Publié: (2025)
A Survey on Video Temporal Grounding with Multimodal Large Language Model
par: Wu, Jianlong, et autres
Publié: (2025)
par: Wu, Jianlong, et autres
Publié: (2025)
Token-level Correlation-guided Compression for Efficient Multimodal Document Understanding
par: Zhang, Renshan, et autres
Publié: (2024)
par: Zhang, Renshan, et autres
Publié: (2024)
ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understanding
par: Wang, Xiao, et autres
Publié: (2024)
par: Wang, Xiao, et autres
Publié: (2024)
Leveraging Modality Tags for Enhanced Cross-Modal Video Retrieval
par: Fragomeni, Adriano, et autres
Publié: (2025)
par: Fragomeni, Adriano, et autres
Publié: (2025)
Video DataFlywheel: Resolving the Impossible Data Trinity in Video-Language Understanding
par: Wang, Xiao, et autres
Publié: (2024)
par: Wang, Xiao, et autres
Publié: (2024)
Decoupled Cross-Modal Alignment Network for Text-RGBT Person Retrieval and A High-Quality Benchmark
par: Deng, Yifei, et autres
Publié: (2025)
par: Deng, Yifei, et autres
Publié: (2025)
AdaReTaKe: Adaptive Redundancy Reduction to Perceive Longer for Video-language Understanding
par: Wang, Xiao, et autres
Publié: (2025)
par: Wang, Xiao, et autres
Publié: (2025)
DecAlign: Hierarchical Cross-Modal Alignment for Decoupled Multimodal Representation Learning
par: Qian, Chengxuan, et autres
Publié: (2025)
par: Qian, Chengxuan, et autres
Publié: (2025)
AlignGen: Boosting Personalized Image Generation with Cross-Modality Prior Alignment
par: Lin, Yiheng, et autres
Publié: (2025)
par: Lin, Yiheng, et autres
Publié: (2025)
Documents similaires
-
R^3: Composed Video Retrieval via Reasoning-Guided Recalling and Re-ranking
par: Li, Zixu, et autres
Publié: (2026) -
The 1st Winner for 5th PVUW MeViS-Text Challenge: Strong MLLMs Meet SAM3 for Referring Video Object Segmentation
par: He, Xusheng, et autres
Publié: (2026) -
Advancing Complex Video Object Segmentation via Tracking-Enhanced Prompt: The 1st Winner for 5th PVUW MOSE Challenge
par: Zhang, Jinrong, et autres
Publié: (2026) -
Self-Enhanced Image Clustering with Cross-Modal Semantic Consistency
par: Li, Zihan, et autres
Publié: (2025) -
OmniEgo-R$^2$: A Routed Reasoning Framework for the 1st Cross-Domain EgoCross Challenge at CVPR 2026
par: Li, Zixu, et autres
Publié: (2026)