Saved in:
| Main Authors: | Han, Zeyu, Zhu, Fangrui, Lao, Qianru, Jiang, Huaizu |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2311.17048 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Flexible Visual Relationship Segmentation
by: Zhu, Fangrui, et al.
Published: (2024)
by: Zhu, Fangrui, et al.
Published: (2024)
SGCap: Decoding Semantic Group for Zero-shot Video Captioning
by: Pan, Zeyu, et al.
Published: (2025)
by: Pan, Zeyu, et al.
Published: (2025)
Interpretable Zero-shot Referring Expression Comprehension with Query-driven Scene Graphs
by: Wu, Yike, et al.
Published: (2026)
by: Wu, Yike, et al.
Published: (2026)
Struct2D: A Perception-Guided Framework for Spatial Reasoning in MLLMs
by: Zhu, Fangrui, et al.
Published: (2025)
by: Zhu, Fangrui, et al.
Published: (2025)
The Solution for the 5th GCAIAC Zero-shot Referring Expression Comprehension Challenge
by: Huang, Longfei, et al.
Published: (2024)
by: Huang, Longfei, et al.
Published: (2024)
MeaCap: Memory-Augmented Zero-shot Image Captioning
by: Zeng, Zequn, et al.
Published: (2024)
by: Zeng, Zequn, et al.
Published: (2024)
Rethinking Diffusion for Text-Driven Human Motion Generation: Redundant Representations, Evaluation, and Masked Autoregression
by: Meng, Zichong, et al.
Published: (2024)
by: Meng, Zichong, et al.
Published: (2024)
Absolute Coordinates Make Motion Generation Easy
by: Meng, Zichong, et al.
Published: (2025)
by: Meng, Zichong, et al.
Published: (2025)
EgoReasoner: Learning Egocentric 4D Reasoning via Task-Adaptive Structured Thinking
by: Zhu, Fangrui, et al.
Published: (2026)
by: Zhu, Fangrui, et al.
Published: (2026)
Zero-shot Image Editing with Reference Imitation
by: Chen, Xi, et al.
Published: (2024)
by: Chen, Xi, et al.
Published: (2024)
GREx: Generalized Referring Expression Segmentation, Comprehension, and Generation
by: Ding, Henghui, et al.
Published: (2026)
by: Ding, Henghui, et al.
Published: (2026)
IG Captioner: Information Gain Captioners are Strong Zero-shot Classifiers
by: Yang, Chenglin, et al.
Published: (2023)
by: Yang, Chenglin, et al.
Published: (2023)
DCVNet: Dilated Cost Volume Networks for Fast Optical Flow
by: Jiang, Huaizu, et al.
Published: (2021)
by: Jiang, Huaizu, et al.
Published: (2021)
Zero-Shot Referring Expression Comprehension via Vison-Language True/False Verification
by: Liu, Jeffrey, et al.
Published: (2025)
by: Liu, Jeffrey, et al.
Published: (2025)
Bridge the Gap Between Visual and Linguistic Comprehension for Generalized Zero-shot Semantic Segmentation
by: Guo, Xiaoqing, et al.
Published: (2025)
by: Guo, Xiaoqing, et al.
Published: (2025)
CLIP-SCGI: Synthesized Caption-Guided Inversion for Person Re-Identification
by: Han, Qianru, et al.
Published: (2024)
by: Han, Qianru, et al.
Published: (2024)
A Strong Baseline for Point Cloud Registration via Direct Superpoints Matching
by: Gupta, Aniket, et al.
Published: (2023)
by: Gupta, Aniket, et al.
Published: (2023)
SynC: Synthetic Image Caption Dataset Refinement with One-to-many Mapping for Zero-shot Image Captioning
by: Kim, Si-Woo, et al.
Published: (2025)
by: Kim, Si-Woo, et al.
Published: (2025)
Zero-shot Quantization: A Comprehensive Survey
by: Kim, Minjun, et al.
Published: (2025)
by: Kim, Minjun, et al.
Published: (2025)
Spatial Transcriptomics Analysis of Zero-shot Gene Expression Prediction
by: Yang, Yan, et al.
Published: (2024)
by: Yang, Yan, et al.
Published: (2024)
ODTFormer: Efficient Obstacle Detection and Tracking with Stereo Cameras Based on Transformer
by: Ding, Tianye, et al.
Published: (2024)
by: Ding, Tianye, et al.
Published: (2024)
ZeroIDIR: Zero-Reference Illumination Degradation Image Restoration with Perturbed Consistency Diffusion Models
by: Jiang, Hai, et al.
Published: (2026)
by: Jiang, Hai, et al.
Published: (2026)
Hierarchical Alignment-enhanced Adaptive Grounding Network for Generalized Referring Expression Comprehension
by: Wang, Yaxian, et al.
Published: (2025)
by: Wang, Yaxian, et al.
Published: (2025)
IFCap: Image-like Retrieval and Frequency-based Entity Filtering for Zero-shot Captioning
by: Lee, Soeun, et al.
Published: (2024)
by: Lee, Soeun, et al.
Published: (2024)
ScanFormer: Referring Expression Comprehension by Iteratively Scanning
by: Su, Wei, et al.
Published: (2024)
by: Su, Wei, et al.
Published: (2024)
Beyond Referring Expressions: Scenario Comprehension Visual Grounding
by: He, Ruozhen, et al.
Published: (2026)
by: He, Ruozhen, et al.
Published: (2026)
Referring Expression Comprehension for Small Objects
by: Goto, Kanoko, et al.
Published: (2025)
by: Goto, Kanoko, et al.
Published: (2025)
SV4D 2.0: Enhancing Spatio-Temporal Consistency in Multi-View Video Diffusion for High-Quality 4D Generation
by: Yao, Chun-Han, et al.
Published: (2025)
by: Yao, Chun-Han, et al.
Published: (2025)
SV4D: Dynamic 3D Content Generation with Multi-Frame and Multi-View Consistency
by: Xie, Yiming, et al.
Published: (2024)
by: Xie, Yiming, et al.
Published: (2024)
ZeroComp: Zero-shot Object Compositing from Image Intrinsics via Diffusion
by: Zhang, Zitian, et al.
Published: (2024)
by: Zhang, Zitian, et al.
Published: (2024)
Multi-method Integration with Confidence-based Weighting for Zero-shot Image Classification
by: Yin, Siqi, et al.
Published: (2024)
by: Yin, Siqi, et al.
Published: (2024)
DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding
by: Wu, Hao, et al.
Published: (2024)
by: Wu, Hao, et al.
Published: (2024)
UniCorrn: Unified Correspondence Transformer Across 2D and 3D
by: Goswami, Prajnan, et al.
Published: (2026)
by: Goswami, Prajnan, et al.
Published: (2026)
LLM-Free Image Captioning Evaluation in Reference-Flexible Settings
by: Hirano, Shinnosuke, et al.
Published: (2025)
by: Hirano, Shinnosuke, et al.
Published: (2025)
Knowledge-Enhanced Dual-stream Zero-shot Composed Image Retrieval
by: Suo, Yucheng, et al.
Published: (2024)
by: Suo, Yucheng, et al.
Published: (2024)
Few-shot Learner Parameterization by Diffusion Time-steps
by: Yue, Zhongqi, et al.
Published: (2024)
by: Yue, Zhongqi, et al.
Published: (2024)
Beyond Cosine Similarity: Magnitude-Aware CLIP for No-Reference Image Quality Assessment
by: Liao, Zhicheng, et al.
Published: (2025)
by: Liao, Zhicheng, et al.
Published: (2025)
RefDrone: A Challenging Benchmark for Referring Expression Comprehension in Drone Scenes
by: Sun, Zhichao, et al.
Published: (2025)
by: Sun, Zhichao, et al.
Published: (2025)
MSD-Score: Multi-Scale Distributional Scoring for Reference-Free Image Caption Evaluation
by: Kan, Shichao, et al.
Published: (2026)
by: Kan, Shichao, et al.
Published: (2026)
Negative Entity Suppression for Zero-Shot Captioning with Synthetic Images
by: Lu, Zimao, et al.
Published: (2025)
by: Lu, Zimao, et al.
Published: (2025)
Similar Items
-
Towards Flexible Visual Relationship Segmentation
by: Zhu, Fangrui, et al.
Published: (2024) -
SGCap: Decoding Semantic Group for Zero-shot Video Captioning
by: Pan, Zeyu, et al.
Published: (2025) -
Interpretable Zero-shot Referring Expression Comprehension with Query-driven Scene Graphs
by: Wu, Yike, et al.
Published: (2026) -
Struct2D: A Perception-Guided Framework for Spatial Reasoning in MLLMs
by: Zhu, Fangrui, et al.
Published: (2025) -
The Solution for the 5th GCAIAC Zero-shot Referring Expression Comprehension Challenge
by: Huang, Longfei, et al.
Published: (2024)