Gespeichert in:
| Hauptverfasser: | Yang, Jin, Wei, Ping, Li, Huan, Ren, Ziyang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2404.09263 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VideoLights: Feature Refinement and Cross-Task Alignment Transformer for Joint Video Highlight Detection and Moment Retrieval
von: Paul, Dhiman, et al.
Veröffentlicht: (2024)
von: Paul, Dhiman, et al.
Veröffentlicht: (2024)
TR-DETR: Task-Reciprocal Transformer for Joint Moment Retrieval and Highlight Detection
von: Sun, Hao, et al.
Veröffentlicht: (2024)
von: Sun, Hao, et al.
Veröffentlicht: (2024)
Watch Video, Catch Keyword: Context-aware Keyword Attention for Moment Retrieval and Highlight Detection
von: Um, Sung Jin, et al.
Veröffentlicht: (2025)
von: Um, Sung Jin, et al.
Veröffentlicht: (2025)
MomentSeeker: A Task-Oriented Benchmark For Long-Video Moment Retrieval
von: Yuan, Huaying, et al.
Veröffentlicht: (2025)
von: Yuan, Huaying, et al.
Veröffentlicht: (2025)
GPTSee: Enhancing Moment Retrieval and Highlight Detection via Description-Based Similarity Features
von: Sun, Yunzhuo, et al.
Veröffentlicht: (2024)
von: Sun, Yunzhuo, et al.
Veröffentlicht: (2024)
Bridging Pixels and Words: Mask-Aware Local Semantic Fusion for Multimodal Media Verification
von: Chen, Zizhao, et al.
Veröffentlicht: (2026)
von: Chen, Zizhao, et al.
Veröffentlicht: (2026)
Decouple-Then-Merge: Finetune Diffusion Models as Multi-Task Learning
von: Ma, Qianli, et al.
Veröffentlicht: (2024)
von: Ma, Qianli, et al.
Veröffentlicht: (2024)
Joint-Task Regularization for Partially Labeled Multi-Task Learning
von: Nishi, Kento, et al.
Veröffentlicht: (2024)
von: Nishi, Kento, et al.
Veröffentlicht: (2024)
Towards Unified Modeling in Federated Multi-Task Learning via Subspace Decoupling
von: Wei, Yipan, et al.
Veröffentlicht: (2025)
von: Wei, Yipan, et al.
Veröffentlicht: (2025)
When One Moment Isn't Enough: Multi-Moment Retrieval with Cross-Moment Interactions
von: Cao, Zhuo, et al.
Veröffentlicht: (2025)
von: Cao, Zhuo, et al.
Veröffentlicht: (2025)
Two-Stream Interactive Joint Learning of Scene Parsing and Geometric Vision Tasks
von: Tang, Guanfeng, et al.
Veröffentlicht: (2026)
von: Tang, Guanfeng, et al.
Veröffentlicht: (2026)
Multi-path Exploration and Feedback Adjustment for Text-to-Image Person Retrieval
von: Kang, Bin, et al.
Veröffentlicht: (2024)
von: Kang, Bin, et al.
Veröffentlicht: (2024)
DiffusionVMR: Diffusion Model for Joint Video Moment Retrieval and Highlight Detection
von: Zhao, Henghao, et al.
Veröffentlicht: (2023)
von: Zhao, Henghao, et al.
Veröffentlicht: (2023)
Exploring Task-Level Optimal Prompts for Visual In-Context Learning
von: Zhu, Yan, et al.
Veröffentlicht: (2025)
von: Zhu, Yan, et al.
Veröffentlicht: (2025)
MedUnifier: Unifying Vision-and-Language Pre-training on Medical Data with Vision Generation Task using Discrete Visual Representations
von: Zhang, Ziyang, et al.
Veröffentlicht: (2025)
von: Zhang, Ziyang, et al.
Veröffentlicht: (2025)
GIRL-DETR: Gradient-Isolated Reinforcement Learning for Video Moment Retrieval
von: Zhang, Shihang, et al.
Veröffentlicht: (2026)
von: Zhang, Shihang, et al.
Veröffentlicht: (2026)
A Light-Weight Framework for Open-Set Object Detection with Decoupled Feature Alignment in Joint Space
von: He, Yonghao, et al.
Veröffentlicht: (2024)
von: He, Yonghao, et al.
Veröffentlicht: (2024)
Stability Plasticity Decoupled Fine-tuning For Few-shot end-to-end Object Detection
von: Yin, Yuantao, et al.
Veröffentlicht: (2024)
von: Yin, Yuantao, et al.
Veröffentlicht: (2024)
MomentMix Augmentation with Length-Aware DETR for Temporally Robust Moment Retrieval
von: Park, Seojeong, et al.
Veröffentlicht: (2024)
von: Park, Seojeong, et al.
Veröffentlicht: (2024)
Deep Extrinsic Manifold Representation for Vision Tasks
von: Zhang, Tongtong, et al.
Veröffentlicht: (2024)
von: Zhang, Tongtong, et al.
Veröffentlicht: (2024)
Transferability-Guided Cross-Domain Cross-Task Transfer Learning
von: Tan, Yang, et al.
Veröffentlicht: (2022)
von: Tan, Yang, et al.
Veröffentlicht: (2022)
Understanding Retrieval-Augmented Task Adaptation for Vision-Language Models
von: Ming, Yifei, et al.
Veröffentlicht: (2024)
von: Ming, Yifei, et al.
Veröffentlicht: (2024)
General and Task-Oriented Video Segmentation
von: Chen, Mu, et al.
Veröffentlicht: (2024)
von: Chen, Mu, et al.
Veröffentlicht: (2024)
Unleash the Potential of CLIP for Video Highlight Detection
von: Han, Donghoon, et al.
Veröffentlicht: (2024)
von: Han, Donghoon, et al.
Veröffentlicht: (2024)
Scale Decoupled Distillation
von: Luo, Shicai Wei Chunbo Luo Yang
Veröffentlicht: (2024)
von: Luo, Shicai Wei Chunbo Luo Yang
Veröffentlicht: (2024)
AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks
von: Ku, Max, et al.
Veröffentlicht: (2024)
von: Ku, Max, et al.
Veröffentlicht: (2024)
NuWa: Deriving Lightweight Task-Specific Vision Transformers for Edge Devices
von: Wei, Ziteng, et al.
Veröffentlicht: (2025)
von: Wei, Ziteng, et al.
Veröffentlicht: (2025)
Denoising Task Routing for Diffusion Models
von: Park, Byeongjun, et al.
Veröffentlicht: (2023)
von: Park, Byeongjun, et al.
Veröffentlicht: (2023)
Task Me Anything
von: Zhang, Jieyu, et al.
Veröffentlicht: (2024)
von: Zhang, Jieyu, et al.
Veröffentlicht: (2024)
Musketeer: Joint Training for Multi-task Vision Language Model with Task Explanation Prompts
von: Zhang, Zhaoyang, et al.
Veröffentlicht: (2023)
von: Zhang, Zhaoyang, et al.
Veröffentlicht: (2023)
Task Prototype-Based Knowledge Retrieval for Multi-Task Learning from Partially Annotated Data
von: Oh, Youngmin, et al.
Veröffentlicht: (2026)
von: Oh, Youngmin, et al.
Veröffentlicht: (2026)
HotelMatch-LLM: Joint Multi-Task Training of Small and Large Language Models for Efficient Multimodal Hotel Retrieval
von: Askari, Arian, et al.
Veröffentlicht: (2025)
von: Askari, Arian, et al.
Veröffentlicht: (2025)
VERIFIED: A Video Corpus Moment Retrieval Benchmark for Fine-Grained Video Understanding
von: Chen, Houlun, et al.
Veröffentlicht: (2024)
von: Chen, Houlun, et al.
Veröffentlicht: (2024)
Eureka-Moments in Transformers: Multi-Step Tasks Reveal Softmax Induced Optimization Problems
von: Hoffmann, David T., et al.
Veröffentlicht: (2023)
von: Hoffmann, David T., et al.
Veröffentlicht: (2023)
SLVideo: A Sign Language Video Moment Retrieval Framework
von: Martins, Gonçalo Vinagre, et al.
Veröffentlicht: (2024)
von: Martins, Gonçalo Vinagre, et al.
Veröffentlicht: (2024)
SMART: Shot-Aware Multimodal Video Moment Retrieval with Audio-Enhanced MLLM
von: Yu, An, et al.
Veröffentlicht: (2025)
von: Yu, An, et al.
Veröffentlicht: (2025)
Information-Theoretic Optimization for Task-Adapted Compressed Sensing Magnetic Resonance Imaging
von: Peng, Xinyu, et al.
Veröffentlicht: (2026)
von: Peng, Xinyu, et al.
Veröffentlicht: (2026)
Beyond Adapter Retrieval: Latent Geometry-Preserving Composition via Sparse Task Projection
von: Jin, Pengfei, et al.
Veröffentlicht: (2024)
von: Jin, Pengfei, et al.
Veröffentlicht: (2024)
MapDream: Task-Driven Map Learning for Vision-Language Navigation
von: Lian, Guoxin, et al.
Veröffentlicht: (2026)
von: Lian, Guoxin, et al.
Veröffentlicht: (2026)
Apollo: Unified Multi-Task Audio-Video Joint Generation
von: Wang, Jun, et al.
Veröffentlicht: (2026)
von: Wang, Jun, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
VideoLights: Feature Refinement and Cross-Task Alignment Transformer for Joint Video Highlight Detection and Moment Retrieval
von: Paul, Dhiman, et al.
Veröffentlicht: (2024) -
TR-DETR: Task-Reciprocal Transformer for Joint Moment Retrieval and Highlight Detection
von: Sun, Hao, et al.
Veröffentlicht: (2024) -
Watch Video, Catch Keyword: Context-aware Keyword Attention for Moment Retrieval and Highlight Detection
von: Um, Sung Jin, et al.
Veröffentlicht: (2025) -
MomentSeeker: A Task-Oriented Benchmark For Long-Video Moment Retrieval
von: Yuan, Huaying, et al.
Veröffentlicht: (2025) -
GPTSee: Enhancing Moment Retrieval and Highlight Detection via Description-Based Similarity Features
von: Sun, Yunzhuo, et al.
Veröffentlicht: (2024)