Multimodal Query-guided Object Localization
Fuente:
arXiv
Salvato in:
| Autori principali: | Tripathi, Aditay, Dani, Rajath R, Mishra, Anand, Chakraborty, Anirban |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2022
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Sketch-guided Image Inpainting with Partial Discrete Diffusion Process
di: Sharma, Nakul, et al.
Pubblicazione: (2024)
di: Sharma, Nakul, et al.
Pubblicazione: (2024)
Prompt Estimation from Prototypes for Federated Prompt Tuning of Vision Transformers
di: Yashwanth, M, et al.
Pubblicazione: (2025)
di: Yashwanth, M, et al.
Pubblicazione: (2025)
O3SLM: Open Weight, Open Data, and Open Vocabulary Sketch-Language Model
di: Gupta, Rishi, et al.
Pubblicazione: (2025)
di: Gupta, Rishi, et al.
Pubblicazione: (2025)
Aligning Moments in Time using Video Queries
di: Kumar, Yogesh, et al.
Pubblicazione: (2025)
di: Kumar, Yogesh, et al.
Pubblicazione: (2025)
Temporal Object-Aware Vision Transformer for Few-Shot Video Object Detection
di: Kumar, Yogesh, et al.
Pubblicazione: (2025)
di: Kumar, Yogesh, et al.
Pubblicazione: (2025)
Multimodal Object Query Initialization for 3D Object Detection
di: van Geerenstein, Mathijs R., et al.
Pubblicazione: (2023)
di: van Geerenstein, Mathijs R., et al.
Pubblicazione: (2023)
Text-guided Zero-Shot Object Localization
di: Wang, Jingjing, et al.
Pubblicazione: (2024)
di: Wang, Jingjing, et al.
Pubblicazione: (2024)
PRVQL: Progressive Knowledge-guided Refinement for Robust Egocentric Visual Query Localization
di: Fan, Bing, et al.
Pubblicazione: (2025)
di: Fan, Bing, et al.
Pubblicazione: (2025)
Objects in Generated Videos Are Slower Than They Appear: Models Suffer Sub-Earth Gravity and Don't Know Galileo's Principle...for now
di: Thozhiyoor, Varun Varma, et al.
Pubblicazione: (2025)
di: Thozhiyoor, Varun Varma, et al.
Pubblicazione: (2025)
Knowledge-guided Causal Intervention for Weakly-supervised Object Localization
di: Shao, Feifei, et al.
Pubblicazione: (2023)
di: Shao, Feifei, et al.
Pubblicazione: (2023)
Localizing Events in Videos with Multimodal Queries
di: Zhang, Gengyuan, et al.
Pubblicazione: (2024)
di: Zhang, Gengyuan, et al.
Pubblicazione: (2024)
Pose-Transformation and Radial Distance Clustering for Unsupervised Person Re-identification
di: Seth, Siddharth, et al.
Pubblicazione: (2024)
di: Seth, Siddharth, et al.
Pubblicazione: (2024)
Composite Sketch+Text Queries for Retrieving Objects with Elusive Names and Complex Interactions
di: Gatti, Prajwal, et al.
Pubblicazione: (2025)
di: Gatti, Prajwal, et al.
Pubblicazione: (2025)
Online Episodic Memory Visual Query Localization with Egocentric Streaming Object Memory
di: Manigrasso, Zaira, et al.
Pubblicazione: (2024)
di: Manigrasso, Zaira, et al.
Pubblicazione: (2024)
Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models
di: Wu, Tz-Ying, et al.
Pubblicazione: (2025)
di: Wu, Tz-Ying, et al.
Pubblicazione: (2025)
DQ3D: Depth-guided Query for Transformer-Based 3D Object Detection in Traffic Scenarios
di: Wang, Ziyu, et al.
Pubblicazione: (2025)
di: Wang, Ziyu, et al.
Pubblicazione: (2025)
SAM-DAQ: Segment Anything Model with Depth-guided Adaptive Queries for RGB-D Video Salient Object Detection
di: Lin, Jia, et al.
Pubblicazione: (2025)
di: Lin, Jia, et al.
Pubblicazione: (2025)
LoVoRA: Text-guided and Mask-free Video Object Removal and Addition with Learnable Object-aware Localization
di: Xiao, Zhihan, et al.
Pubblicazione: (2025)
di: Xiao, Zhihan, et al.
Pubblicazione: (2025)
LDA-AQU: Adaptive Query-guided Upsampling via Local Deformable Attention
di: Du, Zewen, et al.
Pubblicazione: (2024)
di: Du, Zewen, et al.
Pubblicazione: (2024)
Improving Domain Adaptation Through Class Aware Frequency Transformation
di: Kumar, Vikash, et al.
Pubblicazione: (2024)
di: Kumar, Vikash, et al.
Pubblicazione: (2024)
PatentLMM: Large Multimodal Model for Generating Descriptions for Patent Figures
di: Shukla, Shreya, et al.
Pubblicazione: (2025)
di: Shukla, Shreya, et al.
Pubblicazione: (2025)
Selective Query-guided Debiasing for Video Corpus Moment Retrieval
di: Yoon, Sunjae, et al.
Pubblicazione: (2022)
di: Yoon, Sunjae, et al.
Pubblicazione: (2022)
Query-guided Prototype Evolution Network for Few-Shot Segmentation
di: Cong, Runmin, et al.
Pubblicazione: (2024)
di: Cong, Runmin, et al.
Pubblicazione: (2024)
Video Object Segmentation with Dynamic Query Modulation
di: Zhou, Hantao, et al.
Pubblicazione: (2024)
di: Zhou, Hantao, et al.
Pubblicazione: (2024)
Harnessing Object Grounding for Time-Sensitive Video Understanding
di: Wu, Tz-Ying, et al.
Pubblicazione: (2025)
di: Wu, Tz-Ying, et al.
Pubblicazione: (2025)
Generalized-Scale Object Counting with Gradual Query Aggregation
di: Pelhan, Jer, et al.
Pubblicazione: (2025)
di: Pelhan, Jer, et al.
Pubblicazione: (2025)
Reconstructing Objects along Hand Interaction Timelines in Egocentric Video
di: Zhu, Zhifan, et al.
Pubblicazione: (2025)
di: Zhu, Zhifan, et al.
Pubblicazione: (2025)
Visual Text Matters: Improving Text-KVQA with Visual Text Entity Knowledge-aware Large Multimodal Assistant
di: Penamakuri, Abhirama Subramanyam, et al.
Pubblicazione: (2024)
di: Penamakuri, Abhirama Subramanyam, et al.
Pubblicazione: (2024)
Slot-guided Volumetric Object Radiance Fields
di: Qi, Di, et al.
Pubblicazione: (2024)
di: Qi, Di, et al.
Pubblicazione: (2024)
FedSCAl: Leveraging Server and Client Alignment for Unsupervised Federated Source-Free Domain Adaptation
di: Yashwanth, M, et al.
Pubblicazione: (2025)
di: Yashwanth, M, et al.
Pubblicazione: (2025)
A vision-language model and platform for temporally mapping surgery from video
di: Kiyasseh, Dani
Pubblicazione: (2026)
di: Kiyasseh, Dani
Pubblicazione: (2026)
DQ-DETR: DETR with Dynamic Query for Tiny Object Detection
di: Huang, Yi-Xin, et al.
Pubblicazione: (2024)
di: Huang, Yi-Xin, et al.
Pubblicazione: (2024)
Referring Video Object Segmentation with Cross-Modality Proxy Queries
di: Sun, Baoli, et al.
Pubblicazione: (2025)
di: Sun, Baoli, et al.
Pubblicazione: (2025)
TACO-Net: Topological Signatures Triumph in 3D Object Classification
di: Ghosh, Anirban, et al.
Pubblicazione: (2025)
di: Ghosh, Anirban, et al.
Pubblicazione: (2025)
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization
di: Pham, Tan-Hanh, et al.
Pubblicazione: (2024)
di: Pham, Tan-Hanh, et al.
Pubblicazione: (2024)
Intrinsic-feature-guided 3D Object Detection
di: Zhang, Wanjing, et al.
Pubblicazione: (2025)
di: Zhang, Wanjing, et al.
Pubblicazione: (2025)
Stream and Query-guided Feature Aggregation for Efficient and Effective 3D Occupancy Prediction
di: Moon, Seokha, et al.
Pubblicazione: (2025)
di: Moon, Seokha, et al.
Pubblicazione: (2025)
Graph-Based Multimodal and Multi-view Alignment for Keystep Recognition
di: Romero, Julia Lee, et al.
Pubblicazione: (2025)
di: Romero, Julia Lee, et al.
Pubblicazione: (2025)
Towards Visual Query Localization in the 3D World
di: Peng, Liang, et al.
Pubblicazione: (2026)
di: Peng, Liang, et al.
Pubblicazione: (2026)
Decoupled PROB: Decoupled Query Initialization Tasks and Objectness-Class Learning for Open World Object Detection
di: Inoue, Riku, et al.
Pubblicazione: (2025)
di: Inoue, Riku, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Sketch-guided Image Inpainting with Partial Discrete Diffusion Process
di: Sharma, Nakul, et al.
Pubblicazione: (2024) -
Prompt Estimation from Prototypes for Federated Prompt Tuning of Vision Transformers
di: Yashwanth, M, et al.
Pubblicazione: (2025) -
O3SLM: Open Weight, Open Data, and Open Vocabulary Sketch-Language Model
di: Gupta, Rishi, et al.
Pubblicazione: (2025) -
Aligning Moments in Time using Video Queries
di: Kumar, Yogesh, et al.
Pubblicazione: (2025) -
Temporal Object-Aware Vision Transformer for Few-Shot Video Object Detection
di: Kumar, Yogesh, et al.
Pubblicazione: (2025)