REMAP: Regularized Matching and Partial Alignment of Video Embeddings
Fuente:
arXiv
Saved in:
| Main Authors: | Chandra, Soumyadeep, Roy, Kaushik |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Visual Syntactical Understanding
by: Chowdhury, Sayeed Shafayet, et al.
Published: (2024)
by: Chowdhury, Sayeed Shafayet, et al.
Published: (2024)
Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
by: Alwis, Praditha, et al.
Published: (2026)
by: Alwis, Praditha, et al.
Published: (2026)
Temporal Alignment-Free Video Matching for Few-shot Action Recognition
by: Lee, SuBeen, et al.
Published: (2025)
by: Lee, SuBeen, et al.
Published: (2025)
Panoptic Diffusion Models: co-generation of images and segmentation maps
by: Long, Yinghan, et al.
Published: (2024)
by: Long, Yinghan, et al.
Published: (2024)
Unsupervised Domain Adaptation for Action Recognition via Self-Ensembling and Conditional Embedding Alignment
by: Ghosh, Indrajeet, et al.
Published: (2024)
by: Ghosh, Indrajeet, et al.
Published: (2024)
On Inherent Adversarial Robustness of Active Vision Systems
by: Mukherjee, Amitangshu, et al.
Published: (2024)
by: Mukherjee, Amitangshu, et al.
Published: (2024)
Progressive Prompt Detailing for Improved Alignment in Text-to-Image Generative Models
by: Saichandran, Ketan Suhaas, et al.
Published: (2025)
by: Saichandran, Ketan Suhaas, et al.
Published: (2025)
Mitigating Hallucination in Vision-Language Models through Barrier-Regulated Adaptive Closed-form Steering
by: Jana, Soumyadeep, et al.
Published: (2026)
by: Jana, Soumyadeep, et al.
Published: (2026)
Ambiguity-Restrained Text-Video Representation Learning for Partially Relevant Video Retrieval
by: Cho, CH, et al.
Published: (2025)
by: Cho, CH, et al.
Published: (2025)
Mitigating Semantic Collapse in Partially Relevant Video Retrieval
by: Moon, WonJun, et al.
Published: (2025)
by: Moon, WonJun, et al.
Published: (2025)
Uneven Event Modeling for Partially Relevant Video Retrieval
by: Zhu, Sa, et al.
Published: (2025)
by: Zhu, Sa, et al.
Published: (2025)
Non-Contrastive Vision-Language Learning with Predictive Embedding Alignment
by: Kuhn, Lukas, et al.
Published: (2026)
by: Kuhn, Lukas, et al.
Published: (2026)
STR-Match: Matching SpatioTemporal Relevance Score for Training-Free Video Editing
by: Lee, Junsung, et al.
Published: (2025)
by: Lee, Junsung, et al.
Published: (2025)
Temporal Regularization Makes Your Video Generator Stronger
by: Chen, Harold Haodong, et al.
Published: (2025)
by: Chen, Harold Haodong, et al.
Published: (2025)
Towards Two-Stream Foveation-based Active Vision Learning
by: Ibrayev, Timur, et al.
Published: (2024)
by: Ibrayev, Timur, et al.
Published: (2024)
VidVec: Unlocking Video MLLM Embeddings for Video-Text Retrieval
by: Tzachor, Issar, et al.
Published: (2026)
by: Tzachor, Issar, et al.
Published: (2026)
Flowception: Temporally Expansive Flow Matching for Video Generation
by: Ifriqi, Tariq Berrada, et al.
Published: (2025)
by: Ifriqi, Tariq Berrada, et al.
Published: (2025)
MoCA-Video: Motion-Aware Concept Alignment for Consistent Video Editing
by: Zhang, Tong, et al.
Published: (2025)
by: Zhang, Tong, et al.
Published: (2025)
BehAVE: Behaviour Alignment of Video Game Encodings
by: Rašajski, Nemanja, et al.
Published: (2024)
by: Rašajski, Nemanja, et al.
Published: (2024)
Prototypes are Balanced Units for Efficient and Effective Partially Relevant Video Retrieval
by: Moon, WonJun, et al.
Published: (2025)
by: Moon, WonJun, et al.
Published: (2025)
Context-Aware Temporal Embedding of Objects in Video Data
by: Farhan, Ahnaf, et al.
Published: (2024)
by: Farhan, Ahnaf, et al.
Published: (2024)
Video Representation Learning with Joint-Embedding Predictive Architectures
by: Drozdov, Katrina, et al.
Published: (2024)
by: Drozdov, Katrina, et al.
Published: (2024)
Abnormal Event Detection In Videos Using Deep Embedding
by: Venkatrayappa, Darshan
Published: (2024)
by: Venkatrayappa, Darshan
Published: (2024)
HEART: Hyperspherical Embedding Alignment via Kent-Representation Traversal in Diffusion Models
by: Roy, Arani, et al.
Published: (2026)
by: Roy, Arani, et al.
Published: (2026)
Knowledge-Refined Dual Context-Aware Network for Partially Relevant Video Retrieval
by: Yang, Junkai, et al.
Published: (2026)
by: Yang, Junkai, et al.
Published: (2026)
Uncertainty-Guided Self-Questioning and Answering for Video-Language Alignment
by: Chen, Jin, et al.
Published: (2024)
by: Chen, Jin, et al.
Published: (2024)
BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models
by: Tang, Jianting, et al.
Published: (2025)
by: Tang, Jianting, et al.
Published: (2025)
Embedded Representation Learning Network for Animating Styled Video Portrait
by: Wang, Tianyong, et al.
Published: (2024)
by: Wang, Tianyong, et al.
Published: (2024)
SlotMatch: Distilling Object-Centric Representations for Unsupervised Video Segmentation
by: Grigore, Diana-Nicoleta, et al.
Published: (2025)
by: Grigore, Diana-Nicoleta, et al.
Published: (2025)
Deep Unlearning: Fast and Efficient Gradient-free Approach to Class Forgetting
by: Kodge, Sangamesh, et al.
Published: (2023)
by: Kodge, Sangamesh, et al.
Published: (2023)
Alignment-Guided Score Matching for Text-to-Image Alignment in Diffusion Models
by: Lee, Jaa-Yeon, et al.
Published: (2026)
by: Lee, Jaa-Yeon, et al.
Published: (2026)
ViTALS: Vision Transformer for Action Localization in Surgical Nephrectomy
by: Chandra, Soumyadeep, et al.
Published: (2024)
by: Chandra, Soumyadeep, et al.
Published: (2024)
DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization
by: Nguyen, Ngoc-Son, et al.
Published: (2026)
by: Nguyen, Ngoc-Son, et al.
Published: (2026)
Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding
by: Kim, Namho, et al.
Published: (2025)
by: Kim, Namho, et al.
Published: (2025)
V-LynX: Token Interface Alignment for Video+X LLMs
by: Park, Jungin, et al.
Published: (2026)
by: Park, Jungin, et al.
Published: (2026)
Joint-Task Regularization for Partially Labeled Multi-Task Learning
by: Nishi, Kento, et al.
Published: (2024)
by: Nishi, Kento, et al.
Published: (2024)
REGen: Multimodal Retrieval-Embedded Generation for Long-to-Short Video Editing
by: Xu, Weihan, et al.
Published: (2025)
by: Xu, Weihan, et al.
Published: (2025)
Align Where the Words Look: Cross-Attention-Guided Patch Alignment with Contrastive and Transport Regularization for Bengali Captioning
by: Anonto, Riad Ahmed, et al.
Published: (2025)
by: Anonto, Riad Ahmed, et al.
Published: (2025)
CTFlow: Video-Inspired Latent Flow Matching for 3D CT Synthesis
by: Wang, Jiayi, et al.
Published: (2025)
by: Wang, Jiayi, et al.
Published: (2025)
Em-Garde: A Propose-Match Framework for Proactive Streaming Video Understanding
by: Zheng, Yikai, et al.
Published: (2026)
by: Zheng, Yikai, et al.
Published: (2026)
Similar Items
-
Towards Visual Syntactical Understanding
by: Chowdhury, Sayeed Shafayet, et al.
Published: (2024) -
Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
by: Alwis, Praditha, et al.
Published: (2026) -
Temporal Alignment-Free Video Matching for Few-shot Action Recognition
by: Lee, SuBeen, et al.
Published: (2025) -
Panoptic Diffusion Models: co-generation of images and segmentation maps
by: Long, Yinghan, et al.
Published: (2024) -
Unsupervised Domain Adaptation for Action Recognition via Self-Ensembling and Conditional Embedding Alignment
by: Ghosh, Indrajeet, et al.
Published: (2024)