Panoptic Captioning: An Equivalence Bridge for Image and Text
Fuente:
arXiv
Saved in:
| Main Authors: | Lin, Kun-Yu, Wang, Hongjun, Ren, Weining, Han, Kai |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Fin3R: Fine-tuning Feed-forward 3D Reconstruction Models via Monocular Knowledge Distillation
by: Ren, Weining, et al.
Published: (2025)
by: Ren, Weining, et al.
Published: (2025)
Panoptic Segmentation of Mammograms with Text-To-Image Diffusion Model
by: Zhao, Kun, et al.
Published: (2024)
by: Zhao, Kun, et al.
Published: (2024)
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation
by: Peng, Cihang, et al.
Published: (2025)
by: Peng, Cihang, et al.
Published: (2025)
Speed3R: Sparse Feed-forward 3D Reconstruction Models
by: Ren, Weining, et al.
Published: (2026)
by: Ren, Weining, et al.
Published: (2026)
Dynamic Prompting of Frozen Text-to-Image Diffusion Models for Panoptic Narrative Grounding
by: Li, Hongyu, et al.
Published: (2024)
by: Li, Hongyu, et al.
Published: (2024)
COCONut-PanCap: Joint Panoptic Segmentation and Grounded Captions for Fine-Grained Understanding and Generation
by: Deng, Xueqing, et al.
Published: (2025)
by: Deng, Xueqing, et al.
Published: (2025)
Text Data-Centric Image Captioning with Interactive Prompts
by: Wang, Yiyu, et al.
Published: (2024)
by: Wang, Yiyu, et al.
Published: (2024)
Is Your Text-to-Image Model Robust to Caption Noise?
by: Yu, Weichen, et al.
Published: (2024)
by: Yu, Weichen, et al.
Published: (2024)
VisualPrompter: Semantic-Aware Prompt Optimization with Visual Feedback for Text-to-Image Synthesis
by: Wu, Shiyu, et al.
Published: (2025)
by: Wu, Shiyu, et al.
Published: (2025)
SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning
by: Zhang, Lin, et al.
Published: (2025)
by: Zhang, Lin, et al.
Published: (2025)
Removing Distributional Discrepancies in Captions Improves Image-Text Alignment
by: Li, Yuheng, et al.
Published: (2024)
by: Li, Yuheng, et al.
Published: (2024)
Improving Text Generation on Images with Synthetic Captions
by: Koh, Jun Young, et al.
Published: (2024)
by: Koh, Jun Young, et al.
Published: (2024)
Harnessing Caption Detailness for Data-Efficient Text-to-Image Generation
by: Wang, Xinran, et al.
Published: (2025)
by: Wang, Xinran, et al.
Published: (2025)
Evaluating Image Caption via Cycle-consistent Text-to-Image Generation
by: Cui, Tianyu, et al.
Published: (2025)
by: Cui, Tianyu, et al.
Published: (2025)
PanopticRecon: Leverage Open-vocabulary Instance Segmentation for Zero-shot Panoptic Reconstruction
by: Yu, Xuan, et al.
Published: (2024)
by: Yu, Xuan, et al.
Published: (2024)
PhysCorr: Dual-Reward DPO for Physics-Constrained Text-to-Video Generation with Automated Preference Selection
by: Wang, Peiyao, et al.
Published: (2025)
by: Wang, Peiyao, et al.
Published: (2025)
PanopticSplatting: End-to-End Panoptic Gaussian Splatting
by: Xie, Yuxuan, et al.
Published: (2025)
by: Xie, Yuxuan, et al.
Published: (2025)
Multi-Modal LLM based Image Captioning in ICT: Bridging the Gap Between General and Industry Domain
by: Chao, Lianying, et al.
Published: (2026)
by: Chao, Lianying, et al.
Published: (2026)
Text-only Synthesis for Image Captioning
by: Zhou, Qing, et al.
Published: (2024)
by: Zhou, Qing, et al.
Published: (2024)
Continual Learning for Image Captioning through Improved Image-Text Alignment
by: Taetz, Bertram, et al.
Published: (2025)
by: Taetz, Bertram, et al.
Published: (2025)
ITIScore: An Image-to-Text-to-Image Rating Framework for the Image Captioning Ability of MLLMs
by: Xu, Zitong, et al.
Published: (2026)
by: Xu, Zitong, et al.
Published: (2026)
PanopticPartFormer++: A Unified and Decoupled View for Panoptic Part Segmentation
by: Li, Xiangtai, et al.
Published: (2023)
by: Li, Xiangtai, et al.
Published: (2023)
Panoptic-FlashOcc: An Efficient Baseline to Marry Semantic Occupancy with Panoptic via Instance Center
by: Yu, Zichen, et al.
Published: (2024)
by: Yu, Zichen, et al.
Published: (2024)
COCO-OLAC: A Benchmark for Occluded Panoptic Segmentation and Image Understanding
by: Wei, Wenbo, et al.
Published: (2024)
by: Wei, Wenbo, et al.
Published: (2024)
TextPSG: Panoptic Scene Graph Generation from Textual Descriptions
by: Zhao, Chengyang, et al.
Published: (2023)
by: Zhao, Chengyang, et al.
Published: (2023)
OwlCap: Harmonizing Motion-Detail for Video Captioning via HMD-270K and Caption Set Equivalence Reward
by: Zhong, Chunlin, et al.
Published: (2025)
by: Zhong, Chunlin, et al.
Published: (2025)
Generating an Image From 1,000 Words: Enhancing Text-to-Image With Structured Captions
by: Gutflaish, Eyal, et al.
Published: (2025)
by: Gutflaish, Eyal, et al.
Published: (2025)
From Images to Detection: Machine Learning for Blood Pattern Classification
by: Li, Yilin, et al.
Published: (2025)
by: Li, Yilin, et al.
Published: (2025)
InstanceBEV: Unifying Instance and BEV Representation for 3D Panoptic Segmentation
by: Li, Feng, et al.
Published: (2025)
by: Li, Feng, et al.
Published: (2025)
VIXEN: Visual Text Comparison Network for Image Difference Captioning
by: Black, Alexander, et al.
Published: (2024)
by: Black, Alexander, et al.
Published: (2024)
Mitigating Objectness Bias and Region-to-Text Misalignment for Open-Vocabulary Panoptic Segmentation
by: Kormushev, Nikolay, et al.
Published: (2026)
by: Kormushev, Nikolay, et al.
Published: (2026)
PanoSSC: Exploring Monocular Panoptic 3D Scene Reconstruction for Autonomous Driving
by: Shi, Yining, et al.
Published: (2024)
by: Shi, Yining, et al.
Published: (2024)
Parrot Captions Teach CLIP to Spot Text
by: Lin, Yiqi, et al.
Published: (2023)
by: Lin, Yiqi, et al.
Published: (2023)
Precision or Recall? An Analysis of Image Captions for Training Text-to-Image Generation Model
by: Cheng, Sheng, et al.
Published: (2024)
by: Cheng, Sheng, et al.
Published: (2024)
PLGS: Robust Panoptic Lifting with 3D Gaussian Splatting
by: Wang, Yu, et al.
Published: (2024)
by: Wang, Yu, et al.
Published: (2024)
Dissecting Out-of-Distribution Detection and Open-Set Recognition: A Critical Analysis of Methods and Benchmarks
by: Wang, Hongjun, et al.
Published: (2024)
by: Wang, Hongjun, et al.
Published: (2024)
HiLo: A Learning Framework for Generalized Category Discovery Robust to Domain Shifts
by: Wang, Hongjun, et al.
Published: (2024)
by: Wang, Hongjun, et al.
Published: (2024)
SPTNet: An Efficient Alternative Framework for Generalized Category Discovery with Spatial Prompt Tuning
by: Wang, Hongjun, et al.
Published: (2024)
by: Wang, Hongjun, et al.
Published: (2024)
CaptionQA: Is Your Caption as Useful as the Image Itself?
by: Yang, Shijia, et al.
Published: (2025)
by: Yang, Shijia, et al.
Published: (2025)
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos
by: Lin, Weifeng, et al.
Published: (2025)
by: Lin, Weifeng, et al.
Published: (2025)
Similar Items
-
Fin3R: Fine-tuning Feed-forward 3D Reconstruction Models via Monocular Knowledge Distillation
by: Ren, Weining, et al.
Published: (2025) -
Panoptic Segmentation of Mammograms with Text-To-Image Diffusion Model
by: Zhao, Kun, et al.
Published: (2024) -
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation
by: Peng, Cihang, et al.
Published: (2025) -
Speed3R: Sparse Feed-forward 3D Reconstruction Models
by: Ren, Weining, et al.
Published: (2026) -
Dynamic Prompting of Frozen Text-to-Image Diffusion Models for Panoptic Narrative Grounding
by: Li, Hongyu, et al.
Published: (2024)