The Solution for the 5th GCAIAC Zero-shot Referring Expression Comprehension Challenge
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Huang, Longfei, Yu, Feng, Guan, Zhihao, Wan, Zhonghua, Yang, Yang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Solution for the sequential task continual learning track of the 2nd Greater Bay Area International Algorithm Competition
von: Pan, Sishun, et al.
Veröffentlicht: (2024)
von: Pan, Sishun, et al.
Veröffentlicht: (2024)
Harlequin: Color-driven Generation of Synthetic Data for Referring Expression Comprehension
von: Parolari, Luca, et al.
Veröffentlicht: (2024)
von: Parolari, Luca, et al.
Veröffentlicht: (2024)
The Solution for the GAIIC2024 RGB-TIR object detection Challenge
von: Wu, Xiangyu, et al.
Veröffentlicht: (2024)
von: Wu, Xiangyu, et al.
Veröffentlicht: (2024)
Interpretable Zero-shot Referring Expression Comprehension with Query-driven Scene Graphs
von: Wu, Yike, et al.
Veröffentlicht: (2026)
von: Wu, Yike, et al.
Veröffentlicht: (2026)
The Solution for the ICCV 2023 Perception Test Challenge 2023 -- Task 6 -- Grounded videoQA
von: Zhang, Hailiang, et al.
Veröffentlicht: (2024)
von: Zhang, Hailiang, et al.
Veröffentlicht: (2024)
Zero-shot Compound Expression Recognition with Visual Language Model at the 6th ABAW Challenge
von: Wang, Jiahe, et al.
Veröffentlicht: (2024)
von: Wang, Jiahe, et al.
Veröffentlicht: (2024)
FineCops-Ref: A new Dataset and Task for Fine-Grained Compositional Referring Expression Comprehension
von: Liu, Junzhuo, et al.
Veröffentlicht: (2024)
von: Liu, Junzhuo, et al.
Veröffentlicht: (2024)
Zero-shot Referring Expression Comprehension via Structural Similarity Between Images and Captions
von: Han, Zeyu, et al.
Veröffentlicht: (2023)
von: Han, Zeyu, et al.
Veröffentlicht: (2023)
Few-shot Adaptation to Distribution Shifts By Mixing Source and Target Embeddings
von: Xue, Yihao, et al.
Veröffentlicht: (2023)
von: Xue, Yihao, et al.
Veröffentlicht: (2023)
The Solution for Language-Enhanced Image New Category Discovery
von: Xu, Haonan, et al.
Veröffentlicht: (2024)
von: Xu, Haonan, et al.
Veröffentlicht: (2024)
ReMeREC: Relation-aware and Multi-entity Referring Expression Comprehension
von: Hu, Yizhi, et al.
Veröffentlicht: (2025)
von: Hu, Yizhi, et al.
Veröffentlicht: (2025)
The Championship-Winning Solution for the 5th CLVISION Challenge 2024
von: Pan, Sishun, et al.
Veröffentlicht: (2024)
von: Pan, Sishun, et al.
Veröffentlicht: (2024)
Image-Caption Encoding for Improving Zero-Shot Generalization
von: Yu, Eric Yang, et al.
Veröffentlicht: (2024)
von: Yu, Eric Yang, et al.
Veröffentlicht: (2024)
Zero-shot Image Editing with Reference Imitation
von: Chen, Xi, et al.
Veröffentlicht: (2024)
von: Chen, Xi, et al.
Veröffentlicht: (2024)
Troika: Multi-Path Cross-Modal Traction for Compositional Zero-Shot Learning
von: Huang, Siteng, et al.
Veröffentlicht: (2023)
von: Huang, Siteng, et al.
Veröffentlicht: (2023)
Referring Expression Generation in Visually Grounded Dialogue with Discourse-aware Comprehension Guiding
von: Willemsen, Bram, et al.
Veröffentlicht: (2024)
von: Willemsen, Bram, et al.
Veröffentlicht: (2024)
MaPPER: Multimodal Prior-guided Parameter Efficient Tuning for Referring Expression Comprehension
von: Liu, Ting, et al.
Veröffentlicht: (2024)
von: Liu, Ting, et al.
Veröffentlicht: (2024)
IFCap: Image-like Retrieval and Frequency-based Entity Filtering for Zero-shot Captioning
von: Lee, Soeun, et al.
Veröffentlicht: (2024)
von: Lee, Soeun, et al.
Veröffentlicht: (2024)
CK-Transformer: Commonsense Knowledge Enhanced Transformers for Referring Expression Comprehension
von: Zhang, Zhi, et al.
Veröffentlicht: (2023)
von: Zhang, Zhi, et al.
Veröffentlicht: (2023)
Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering
von: Beliaev, Mark, et al.
Veröffentlicht: (2025)
von: Beliaev, Mark, et al.
Veröffentlicht: (2025)
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension
von: Wang, Xiyao, et al.
Veröffentlicht: (2024)
von: Wang, Xiyao, et al.
Veröffentlicht: (2024)
Exploiting GPT-4 Vision for Zero-shot Point Cloud Understanding
von: Sun, Qi, et al.
Veröffentlicht: (2024)
von: Sun, Qi, et al.
Veröffentlicht: (2024)
Spatial Transcriptomics Analysis of Zero-shot Gene Expression Prediction
von: Yang, Yan, et al.
Veröffentlicht: (2024)
von: Yang, Yan, et al.
Veröffentlicht: (2024)
CAPability: A Comprehensive Visual Caption Benchmark for Evaluating Both Correctness and Thoroughness
von: Liu, Zhihang, et al.
Veröffentlicht: (2025)
von: Liu, Zhihang, et al.
Veröffentlicht: (2025)
1st Place Solution for 5th LSVOS Challenge: Referring Video Object Segmentation
von: Luo, Zhuoyan, et al.
Veröffentlicht: (2024)
von: Luo, Zhuoyan, et al.
Veröffentlicht: (2024)
DreamLLM: Synergistic Multimodal Comprehension and Creation
von: Dong, Runpei, et al.
Veröffentlicht: (2023)
von: Dong, Runpei, et al.
Veröffentlicht: (2023)
An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM
von: Kim, Wonkyun, et al.
Veröffentlicht: (2024)
von: Kim, Wonkyun, et al.
Veröffentlicht: (2024)
SynC: Synthetic Image Caption Dataset Refinement with One-to-many Mapping for Zero-shot Image Captioning
von: Kim, Si-Woo, et al.
Veröffentlicht: (2025)
von: Kim, Si-Woo, et al.
Veröffentlicht: (2025)
Solution for OOD-CV Workshop SSB Challenge 2024 (Open-Set Recognition Track)
von: Feng, Mingxu, et al.
Veröffentlicht: (2024)
von: Feng, Mingxu, et al.
Veröffentlicht: (2024)
ChartREG++: Towards Benchmarking and Improving Chart Referring Expression Grounding under Diverse referring clues and Multi-Target Referring
von: Niu, Tianhao, et al.
Veröffentlicht: (2026)
von: Niu, Tianhao, et al.
Veröffentlicht: (2026)
The Solution for the CVPR2024 NICE Image Captioning Challenge
von: Huang, Longfei, et al.
Veröffentlicht: (2024)
von: Huang, Longfei, et al.
Veröffentlicht: (2024)
SafaRi:Adaptive Sequence Transformer for Weakly Supervised Referring Expression Segmentation
von: Nag, Sayan, et al.
Veröffentlicht: (2024)
von: Nag, Sayan, et al.
Veröffentlicht: (2024)
Ref-Adv: Exploring MLLM Visual Reasoning in Referring Expression Tasks
von: Dong, Qihua, et al.
Veröffentlicht: (2026)
von: Dong, Qihua, et al.
Veröffentlicht: (2026)
Towards Perceiving Small Visual Details in Zero-shot Visual Question Answering with Multimodal LLMs
von: Zhang, Jiarui, et al.
Veröffentlicht: (2023)
von: Zhang, Jiarui, et al.
Veröffentlicht: (2023)
UniFine: A Unified and Fine-grained Approach for Zero-shot Vision-Language Understanding
von: Wang, Zhecan, et al.
Veröffentlicht: (2023)
von: Wang, Zhecan, et al.
Veröffentlicht: (2023)
AMU-Tuning: Effective Logit Bias for CLIP-based Few-shot Learning
von: Tang, Yuwei, et al.
Veröffentlicht: (2024)
von: Tang, Yuwei, et al.
Veröffentlicht: (2024)
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks
von: Fan, Wan-Cyuan, et al.
Veröffentlicht: (2024)
von: Fan, Wan-Cyuan, et al.
Veröffentlicht: (2024)
Phrase-Instance Alignment for Generalized Referring Segmentation
von: Nguyen, E-Ro, et al.
Veröffentlicht: (2024)
von: Nguyen, E-Ro, et al.
Veröffentlicht: (2024)
Including Facial Expressions in Contextual Embeddings for Sign Language Generation
von: Viegas, Carla, et al.
Veröffentlicht: (2022)
von: Viegas, Carla, et al.
Veröffentlicht: (2022)
Transcrib3D: 3D Referring Expression Resolution through Large Language Models
von: Fang, Jiading, et al.
Veröffentlicht: (2024)
von: Fang, Jiading, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
The Solution for the sequential task continual learning track of the 2nd Greater Bay Area International Algorithm Competition
von: Pan, Sishun, et al.
Veröffentlicht: (2024) -
Harlequin: Color-driven Generation of Synthetic Data for Referring Expression Comprehension
von: Parolari, Luca, et al.
Veröffentlicht: (2024) -
The Solution for the GAIIC2024 RGB-TIR object detection Challenge
von: Wu, Xiangyu, et al.
Veröffentlicht: (2024) -
Interpretable Zero-shot Referring Expression Comprehension with Query-driven Scene Graphs
von: Wu, Yike, et al.
Veröffentlicht: (2026) -
The Solution for the ICCV 2023 Perception Test Challenge 2023 -- Task 6 -- Grounded videoQA
von: Zhang, Hailiang, et al.
Veröffentlicht: (2024)