Enregistré dans:
| Auteurs principaux: | Yang, Jin, Wei, Ping, Li, Huan, Ren, Ziyang |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | https://arxiv.org/abs/2404.09263 |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
VideoLights: Feature Refinement and Cross-Task Alignment Transformer for Joint Video Highlight Detection and Moment Retrieval
par: Paul, Dhiman, et autres
Publié: (2024)
par: Paul, Dhiman, et autres
Publié: (2024)
TR-DETR: Task-Reciprocal Transformer for Joint Moment Retrieval and Highlight Detection
par: Sun, Hao, et autres
Publié: (2024)
par: Sun, Hao, et autres
Publié: (2024)
Watch Video, Catch Keyword: Context-aware Keyword Attention for Moment Retrieval and Highlight Detection
par: Um, Sung Jin, et autres
Publié: (2025)
par: Um, Sung Jin, et autres
Publié: (2025)
MomentSeeker: A Task-Oriented Benchmark For Long-Video Moment Retrieval
par: Yuan, Huaying, et autres
Publié: (2025)
par: Yuan, Huaying, et autres
Publié: (2025)
GPTSee: Enhancing Moment Retrieval and Highlight Detection via Description-Based Similarity Features
par: Sun, Yunzhuo, et autres
Publié: (2024)
par: Sun, Yunzhuo, et autres
Publié: (2024)
Bridging Pixels and Words: Mask-Aware Local Semantic Fusion for Multimodal Media Verification
par: Chen, Zizhao, et autres
Publié: (2026)
par: Chen, Zizhao, et autres
Publié: (2026)
Decouple-Then-Merge: Finetune Diffusion Models as Multi-Task Learning
par: Ma, Qianli, et autres
Publié: (2024)
par: Ma, Qianli, et autres
Publié: (2024)
Joint-Task Regularization for Partially Labeled Multi-Task Learning
par: Nishi, Kento, et autres
Publié: (2024)
par: Nishi, Kento, et autres
Publié: (2024)
Towards Unified Modeling in Federated Multi-Task Learning via Subspace Decoupling
par: Wei, Yipan, et autres
Publié: (2025)
par: Wei, Yipan, et autres
Publié: (2025)
When One Moment Isn't Enough: Multi-Moment Retrieval with Cross-Moment Interactions
par: Cao, Zhuo, et autres
Publié: (2025)
par: Cao, Zhuo, et autres
Publié: (2025)
Two-Stream Interactive Joint Learning of Scene Parsing and Geometric Vision Tasks
par: Tang, Guanfeng, et autres
Publié: (2026)
par: Tang, Guanfeng, et autres
Publié: (2026)
Multi-path Exploration and Feedback Adjustment for Text-to-Image Person Retrieval
par: Kang, Bin, et autres
Publié: (2024)
par: Kang, Bin, et autres
Publié: (2024)
DiffusionVMR: Diffusion Model for Joint Video Moment Retrieval and Highlight Detection
par: Zhao, Henghao, et autres
Publié: (2023)
par: Zhao, Henghao, et autres
Publié: (2023)
Exploring Task-Level Optimal Prompts for Visual In-Context Learning
par: Zhu, Yan, et autres
Publié: (2025)
par: Zhu, Yan, et autres
Publié: (2025)
MedUnifier: Unifying Vision-and-Language Pre-training on Medical Data with Vision Generation Task using Discrete Visual Representations
par: Zhang, Ziyang, et autres
Publié: (2025)
par: Zhang, Ziyang, et autres
Publié: (2025)
GIRL-DETR: Gradient-Isolated Reinforcement Learning for Video Moment Retrieval
par: Zhang, Shihang, et autres
Publié: (2026)
par: Zhang, Shihang, et autres
Publié: (2026)
A Light-Weight Framework for Open-Set Object Detection with Decoupled Feature Alignment in Joint Space
par: He, Yonghao, et autres
Publié: (2024)
par: He, Yonghao, et autres
Publié: (2024)
Stability Plasticity Decoupled Fine-tuning For Few-shot end-to-end Object Detection
par: Yin, Yuantao, et autres
Publié: (2024)
par: Yin, Yuantao, et autres
Publié: (2024)
MomentMix Augmentation with Length-Aware DETR for Temporally Robust Moment Retrieval
par: Park, Seojeong, et autres
Publié: (2024)
par: Park, Seojeong, et autres
Publié: (2024)
Deep Extrinsic Manifold Representation for Vision Tasks
par: Zhang, Tongtong, et autres
Publié: (2024)
par: Zhang, Tongtong, et autres
Publié: (2024)
Transferability-Guided Cross-Domain Cross-Task Transfer Learning
par: Tan, Yang, et autres
Publié: (2022)
par: Tan, Yang, et autres
Publié: (2022)
Understanding Retrieval-Augmented Task Adaptation for Vision-Language Models
par: Ming, Yifei, et autres
Publié: (2024)
par: Ming, Yifei, et autres
Publié: (2024)
General and Task-Oriented Video Segmentation
par: Chen, Mu, et autres
Publié: (2024)
par: Chen, Mu, et autres
Publié: (2024)
Unleash the Potential of CLIP for Video Highlight Detection
par: Han, Donghoon, et autres
Publié: (2024)
par: Han, Donghoon, et autres
Publié: (2024)
Scale Decoupled Distillation
par: Luo, Shicai Wei Chunbo Luo Yang
Publié: (2024)
par: Luo, Shicai Wei Chunbo Luo Yang
Publié: (2024)
AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks
par: Ku, Max, et autres
Publié: (2024)
par: Ku, Max, et autres
Publié: (2024)
NuWa: Deriving Lightweight Task-Specific Vision Transformers for Edge Devices
par: Wei, Ziteng, et autres
Publié: (2025)
par: Wei, Ziteng, et autres
Publié: (2025)
Denoising Task Routing for Diffusion Models
par: Park, Byeongjun, et autres
Publié: (2023)
par: Park, Byeongjun, et autres
Publié: (2023)
Task Me Anything
par: Zhang, Jieyu, et autres
Publié: (2024)
par: Zhang, Jieyu, et autres
Publié: (2024)
Musketeer: Joint Training for Multi-task Vision Language Model with Task Explanation Prompts
par: Zhang, Zhaoyang, et autres
Publié: (2023)
par: Zhang, Zhaoyang, et autres
Publié: (2023)
Task Prototype-Based Knowledge Retrieval for Multi-Task Learning from Partially Annotated Data
par: Oh, Youngmin, et autres
Publié: (2026)
par: Oh, Youngmin, et autres
Publié: (2026)
HotelMatch-LLM: Joint Multi-Task Training of Small and Large Language Models for Efficient Multimodal Hotel Retrieval
par: Askari, Arian, et autres
Publié: (2025)
par: Askari, Arian, et autres
Publié: (2025)
VERIFIED: A Video Corpus Moment Retrieval Benchmark for Fine-Grained Video Understanding
par: Chen, Houlun, et autres
Publié: (2024)
par: Chen, Houlun, et autres
Publié: (2024)
Eureka-Moments in Transformers: Multi-Step Tasks Reveal Softmax Induced Optimization Problems
par: Hoffmann, David T., et autres
Publié: (2023)
par: Hoffmann, David T., et autres
Publié: (2023)
SLVideo: A Sign Language Video Moment Retrieval Framework
par: Martins, Gonçalo Vinagre, et autres
Publié: (2024)
par: Martins, Gonçalo Vinagre, et autres
Publié: (2024)
SMART: Shot-Aware Multimodal Video Moment Retrieval with Audio-Enhanced MLLM
par: Yu, An, et autres
Publié: (2025)
par: Yu, An, et autres
Publié: (2025)
Information-Theoretic Optimization for Task-Adapted Compressed Sensing Magnetic Resonance Imaging
par: Peng, Xinyu, et autres
Publié: (2026)
par: Peng, Xinyu, et autres
Publié: (2026)
Beyond Adapter Retrieval: Latent Geometry-Preserving Composition via Sparse Task Projection
par: Jin, Pengfei, et autres
Publié: (2024)
par: Jin, Pengfei, et autres
Publié: (2024)
MapDream: Task-Driven Map Learning for Vision-Language Navigation
par: Lian, Guoxin, et autres
Publié: (2026)
par: Lian, Guoxin, et autres
Publié: (2026)
Apollo: Unified Multi-Task Audio-Video Joint Generation
par: Wang, Jun, et autres
Publié: (2026)
par: Wang, Jun, et autres
Publié: (2026)
Documents similaires
-
VideoLights: Feature Refinement and Cross-Task Alignment Transformer for Joint Video Highlight Detection and Moment Retrieval
par: Paul, Dhiman, et autres
Publié: (2024) -
TR-DETR: Task-Reciprocal Transformer for Joint Moment Retrieval and Highlight Detection
par: Sun, Hao, et autres
Publié: (2024) -
Watch Video, Catch Keyword: Context-aware Keyword Attention for Moment Retrieval and Highlight Detection
par: Um, Sung Jin, et autres
Publié: (2025) -
MomentSeeker: A Task-Oriented Benchmark For Long-Video Moment Retrieval
par: Yuan, Huaying, et autres
Publié: (2025) -
GPTSee: Enhancing Moment Retrieval and Highlight Detection via Description-Based Similarity Features
par: Sun, Yunzhuo, et autres
Publié: (2024)