Multimodal LLM-based Query Paraphrasing for Video Search
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Jiaxin, Ngo, Chong-Wah, Chan, Wing-Kwong, Zhong, Sheng-Hua, Wei, Xiong-Yong, Li, Qing |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Interpretable Embedding for Ad-hoc Video Search
by: Wu, Jiaxin, et al.
Published: (2024)
by: Wu, Jiaxin, et al.
Published: (2024)
Robust Relevance Feedback for Interactive Known-Item Video Search
by: Ma, Zhixin, et al.
Published: (2025)
by: Ma, Zhixin, et al.
Published: (2025)
Improving Interpretable Embeddings for Ad-hoc Video Search with Generative Captions and Multi-word Concept Bank
by: Wu, Jiaxin, et al.
Published: (2024)
by: Wu, Jiaxin, et al.
Published: (2024)
Towards Multimodal Emotional Support Conversation Systems
by: Chu, Yuqi, et al.
Published: (2024)
by: Chu, Yuqi, et al.
Published: (2024)
Towards Unbiased Cross-Modal Representation Learning for Food Image-to-Recipe Retrieval
by: Wang, Qing, et al.
Published: (2025)
by: Wang, Qing, et al.
Published: (2025)
Efficient Prompt Tuning for Hierarchical Ingredient Recognition
by: Gui, Yinxuan, et al.
Published: (2025)
by: Gui, Yinxuan, et al.
Published: (2025)
Adaptive Multi-Agent Reasoning for Text-to-Video Retrieval
by: Wu, Jiaxin, et al.
Published: (2025)
by: Wu, Jiaxin, et al.
Published: (2025)
Mitigating Cross-modal Representation Bias for Multicultural Image-to-Recipe Retrieval
by: Wang, Qing, et al.
Published: (2025)
by: Wang, Qing, et al.
Published: (2025)
TOP:A New Target-Audience Oriented Content Paraphrase Task
by: Lin, Boda, et al.
Published: (2024)
by: Lin, Boda, et al.
Published: (2024)
Navigating Weight Prediction with Diet Diary
by: Gui, Yinxuan, et al.
Published: (2024)
by: Gui, Yinxuan, et al.
Published: (2024)
PolySmart @ TRECVid 2024 Video Captioning (VTT)
by: Wu, Jiaxin, et al.
Published: (2024)
by: Wu, Jiaxin, et al.
Published: (2024)
PolySmart @ TRECVid 2024 Medical Video Question Answering
by: Wu, Jiaxin, et al.
Published: (2024)
by: Wu, Jiaxin, et al.
Published: (2024)
Class Agnostic Instance-level Descriptor for Visual Instance Search
by: Sun, Qi-Ying, et al.
Published: (2025)
by: Sun, Qi-Ying, et al.
Published: (2025)
OVFoodSeg: Elevating Open-Vocabulary Food Image Segmentation via Image-Informed Textual Representation
by: Wu, Xiongwei, et al.
Published: (2024)
by: Wu, Xiongwei, et al.
Published: (2024)
Harnessing Multimodal Large Language Models for Personalized Product Search with Query-aware Refinement
by: Zhang, Beibei, et al.
Published: (2025)
by: Zhang, Beibei, et al.
Published: (2025)
QMAVIS: Long Video-Audio Understanding using Fusion of Large Multimodal Models
by: Lin, Zixing, et al.
Published: (2026)
by: Lin, Zixing, et al.
Published: (2026)
Hi3D: Pursuing High-Resolution Image-to-3D Generation with Video Diffusion Models
by: Yang, Haibo, et al.
Published: (2024)
by: Yang, Haibo, et al.
Published: (2024)
PolySmart and VIREO @ TRECVid 2024 Ad-hoc Video Search
by: Wu, Jiaxin, et al.
Published: (2024)
by: Wu, Jiaxin, et al.
Published: (2024)
Where to Focus: Query-Modulated Multimodal Keyframe Selection for Long Video Understanding
by: Wang, Shaoguang, et al.
Published: (2026)
by: Wang, Shaoguang, et al.
Published: (2026)
When Video Coding Meets Multimodal Large Language Models: A Unified Paradigm for Video Coding
by: Zhang, Pingping, et al.
Published: (2024)
by: Zhang, Pingping, et al.
Published: (2024)
Divide and Conquer: Multimodal Video Deepfake Detection via Cross-Modal Fusion and Localization
by: Li, Qingcao, et al.
Published: (2026)
by: Li, Qingcao, et al.
Published: (2026)
Virbo: Multimodal Multilingual Avatar Video Generation in Digital Marketing
by: Zhang, Juan, et al.
Published: (2024)
by: Zhang, Juan, et al.
Published: (2024)
CLIPRerank: An Extremely Simple Method for Improving Ad-hoc Video Search
by: Chen, Aozhu, et al.
Published: (2024)
by: Chen, Aozhu, et al.
Published: (2024)
TheoremExplainAgent: Towards Video-based Multimodal Explanations for LLM Theorem Understanding
by: Ku, Max, et al.
Published: (2025)
by: Ku, Max, et al.
Published: (2025)
Official-NV: An LLM-Generated News Video Dataset for Multimodal Fake News Detection
by: Wang, Yihao, et al.
Published: (2024)
by: Wang, Yihao, et al.
Published: (2024)
FLARE: Full-Modality Long-Video Audiovisual Retrieval Benchmark with User-Simulated Queries
by: You, Qijie, et al.
Published: (2026)
by: You, Qijie, et al.
Published: (2026)
Multimodal Semantic Communication for Generative Audio-Driven Video Conferencing
by: Tong, Haonan, et al.
Published: (2024)
by: Tong, Haonan, et al.
Published: (2024)
QuMATL: Query-based Multi-annotator Tendency Learning
by: Zhang, Liyun, et al.
Published: (2025)
by: Zhang, Liyun, et al.
Published: (2025)
Vidformer: Drop-in Declarative Optimization for Rendering Video-Native Query Results
by: Winecki, Dominik, et al.
Published: (2026)
by: Winecki, Dominik, et al.
Published: (2026)
MaskSearch: Querying Image Masks at Scale
by: He, Dong, et al.
Published: (2023)
by: He, Dong, et al.
Published: (2023)
Beyond Forced Modality Balance: Intrinsic Information Budgets for Multimodal Learning
by: Xiong, Zechang, et al.
Published: (2026)
by: Xiong, Zechang, et al.
Published: (2026)
Multimodal Classification and Out-of-distribution Detection for Multimodal Intent Understanding
by: Zhang, Hanlei, et al.
Published: (2024)
by: Zhang, Hanlei, et al.
Published: (2024)
Multimodal Fake News Video Explanation: Dataset, Analysis and Evaluation
by: Chen, Lizhi, et al.
Published: (2025)
by: Chen, Lizhi, et al.
Published: (2025)
MMC: Iterative Refinement of VLM Reasoning via MCTS-based Multimodal Critique
by: Liu, Shuhang, et al.
Published: (2025)
by: Liu, Shuhang, et al.
Published: (2025)
Multi-modal and Metadata Capture Model for Micro Video Popularity Prediction
by: Lu, Jiacheng, et al.
Published: (2025)
by: Lu, Jiacheng, et al.
Published: (2025)
Demonstration of MaskSearch: Efficiently Querying Image Masks for Machine Learning Workflows
by: Wei, Lindsey Linxi, et al.
Published: (2024)
by: Wei, Lindsey Linxi, et al.
Published: (2024)
Evolutionary Multimodal Reasoning via Hierarchical Semantic Representation for Intent Recognition
by: Zhou, Qianrui, et al.
Published: (2026)
by: Zhou, Qianrui, et al.
Published: (2026)
PRM-BAS: Enhancing Multimodal Reasoning through PRM-guided Beam Annealing Search
by: Hu, Pengfei, et al.
Published: (2025)
by: Hu, Pengfei, et al.
Published: (2025)
MCSC-Bench: Multimodal Context-to-Script Creation for Realistic Video Production
by: Hu, Huanran, et al.
Published: (2026)
by: Hu, Huanran, et al.
Published: (2026)
ELLMPEG: An Edge-based Agentic LLM Video Processing Tool
by: Azimi, Zoha, et al.
Published: (2026)
by: Azimi, Zoha, et al.
Published: (2026)
Similar Items
-
Interpretable Embedding for Ad-hoc Video Search
by: Wu, Jiaxin, et al.
Published: (2024) -
Robust Relevance Feedback for Interactive Known-Item Video Search
by: Ma, Zhixin, et al.
Published: (2025) -
Improving Interpretable Embeddings for Ad-hoc Video Search with Generative Captions and Multi-word Concept Bank
by: Wu, Jiaxin, et al.
Published: (2024) -
Towards Multimodal Emotional Support Conversation Systems
by: Chu, Yuqi, et al.
Published: (2024) -
Towards Unbiased Cross-Modal Representation Learning for Food Image-to-Recipe Retrieval
by: Wang, Qing, et al.
Published: (2025)