Pix2Cap-COCO: Advancing Visual Comprehension via Pixel-Level Captioning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | You, Zuyao, Wang, Junke, Kong, Lingyu, He, Bo, Wu, Zuxuan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FOCUS: Towards Universal Foreground Segmentation
von: You, Zuyao, et al.
Veröffentlicht: (2025)
von: You, Zuyao, et al.
Veröffentlicht: (2025)
VideoLoom: A Video Large Language Model for Joint Spatial-Temporal Understanding
von: Shi, Jiapeng, et al.
Veröffentlicht: (2026)
von: Shi, Jiapeng, et al.
Veröffentlicht: (2026)
FluxMem: Adaptive Hierarchical Memory for Streaming Video Understanding
von: Xie, Yiweng, et al.
Veröffentlicht: (2026)
von: Xie, Yiweng, et al.
Veröffentlicht: (2026)
Seg-R1: Segmentation Can Be Surprisingly Simple with Reinforcement Learning
von: You, Zuyao, et al.
Veröffentlicht: (2025)
von: You, Zuyao, et al.
Veröffentlicht: (2025)
Pixelis: Reasoning in Pixels, from Seeing to Acting
von: Zhou, Yunpeng
Veröffentlicht: (2026)
von: Zhou, Yunpeng
Veröffentlicht: (2026)
PixLore: A Dataset-driven Approach to Rich Image Captioning
von: Bonilla-Salvador, Diego, et al.
Veröffentlicht: (2023)
von: Bonilla-Salvador, Diego, et al.
Veröffentlicht: (2023)
DiffCap-Bench: A Comprehensive, Challenging, Robust Benchmark for Image Difference Captioning
von: Wei, Yuancheng, et al.
Veröffentlicht: (2026)
von: Wei, Yuancheng, et al.
Veröffentlicht: (2026)
DualCap: Enhancing Lightweight Image Captioning via Dual Retrieval with Similar Scenes Visual Prompts
von: Li, Binbin, et al.
Veröffentlicht: (2025)
von: Li, Binbin, et al.
Veröffentlicht: (2025)
UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual Reasoning
von: Liu, Ye, et al.
Veröffentlicht: (2025)
von: Liu, Ye, et al.
Veröffentlicht: (2025)
Multimodal Crowd Counting with Pix2Pix GANs
von: Khan, Muhammad Asif, et al.
Veröffentlicht: (2024)
von: Khan, Muhammad Asif, et al.
Veröffentlicht: (2024)
Learning Accurate Segmentation Purely from Self-Supervision
von: You, Zuyao, et al.
Veröffentlicht: (2026)
von: You, Zuyao, et al.
Veröffentlicht: (2026)
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption
von: Fan, Tiehan, et al.
Veröffentlicht: (2024)
von: Fan, Tiehan, et al.
Veröffentlicht: (2024)
XMeCap: Meme Caption Generation with Sub-Image Adaptability
von: Chen, Yuyan, et al.
Veröffentlicht: (2024)
von: Chen, Yuyan, et al.
Veröffentlicht: (2024)
CapRL: Stimulating Dense Image Caption Capabilities via Reinforcement Learning
von: Xing, Long, et al.
Veröffentlicht: (2025)
von: Xing, Long, et al.
Veröffentlicht: (2025)
PixelPrune: Pixel-Level Adaptive Visual Token Reduction via Predictive Coding
von: Wang, Nan, et al.
Veröffentlicht: (2026)
von: Wang, Nan, et al.
Veröffentlicht: (2026)
Knowledge Completes the Vision: A Multimodal Entity-aware Retrieval-Augmented Generation Framework for News Image Captioning
von: You, Xiaoxing, et al.
Veröffentlicht: (2025)
von: You, Xiaoxing, et al.
Veröffentlicht: (2025)
ChartCap: Mitigating Hallucination of Dense Chart Captioning
von: Lim, Junyoung, et al.
Veröffentlicht: (2025)
von: Lim, Junyoung, et al.
Veröffentlicht: (2025)
BalCapRL: A Balanced Framework for RL-Based MLLM Image Captioning
von: Ye, Shaokai, et al.
Veröffentlicht: (2026)
von: Ye, Shaokai, et al.
Veröffentlicht: (2026)
CompCap: Improving Multimodal Large Language Models with Composite Captions
von: Chen, Xiaohui, et al.
Veröffentlicht: (2024)
von: Chen, Xiaohui, et al.
Veröffentlicht: (2024)
RoCOCO: Robustness Benchmark of MS-COCO to Stress-test Image-Text Matching Models
von: Park, Seulki, et al.
Veröffentlicht: (2023)
von: Park, Seulki, et al.
Veröffentlicht: (2023)
SRU-Pix2Pix: A Fusion-Driven Generator Network for Medical Image Translation with Few-Shot Learning
von: Qiu, Xihe, et al.
Veröffentlicht: (2026)
von: Qiu, Xihe, et al.
Veröffentlicht: (2026)
Pix2Code: Learning to Compose Neural Visual Concepts as Programs
von: Wüst, Antonia, et al.
Veröffentlicht: (2024)
von: Wüst, Antonia, et al.
Veröffentlicht: (2024)
CapGeo: A Caption-Assisted Approach to Geometric Reasoning
von: Li, Yuying, et al.
Veröffentlicht: (2025)
von: Li, Yuying, et al.
Veröffentlicht: (2025)
PixelArena: A benchmark for Pixel-Precision Visual Intelligence
von: Liang, Feng, et al.
Veröffentlicht: (2025)
von: Liang, Feng, et al.
Veröffentlicht: (2025)
Pix2Gif: Motion-Guided Diffusion for GIF Generation
von: Kandala, Hitesh, et al.
Veröffentlicht: (2024)
von: Kandala, Hitesh, et al.
Veröffentlicht: (2024)
RubiCap: Rubric-Guided Reinforcement Learning for Dense Image Captioning
von: Huang, Tzu-Heng, et al.
Veröffentlicht: (2026)
von: Huang, Tzu-Heng, et al.
Veröffentlicht: (2026)
IBISAgent: Reinforcing Pixel-Level Visual Reasoning in MLLMs for Universal Biomedical Object Referring and Segmentation
von: Jiang, Yankai, et al.
Veröffentlicht: (2026)
von: Jiang, Yankai, et al.
Veröffentlicht: (2026)
ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning
von: Xu, Ziqiang, et al.
Veröffentlicht: (2025)
von: Xu, Ziqiang, et al.
Veröffentlicht: (2025)
Aquila-plus: Prompt-Driven Visual-Language Models for Pixel-Level Remote Sensing Image Understanding
von: Lu, Kaixuan
Veröffentlicht: (2024)
von: Lu, Kaixuan
Veröffentlicht: (2024)
Cut to the Chase: Training-free Multimodal Summarization via Chain-of-Events
von: You, Xiaoxing, et al.
Veröffentlicht: (2026)
von: You, Xiaoxing, et al.
Veröffentlicht: (2026)
LaMP-Cap: Personalized Figure Caption Generation With Multimodal Figure Profiles
von: Ng, Ho Yin 'Sam', et al.
Veröffentlicht: (2025)
von: Ng, Ho Yin 'Sam', et al.
Veröffentlicht: (2025)
Context-Aware Semantic Segmentation: Enhancing Pixel-Level Understanding with Large Language Models for Advanced Vision Applications
von: Rahman, Ben
Veröffentlicht: (2025)
von: Rahman, Ben
Veröffentlicht: (2025)
Target-Dependent Multimodal Sentiment Analysis Via Employing Visual-to Emotional-Caption Translation Network using Visual-Caption Pairs
von: Pandey, Ananya, et al.
Veröffentlicht: (2024)
von: Pandey, Ananya, et al.
Veröffentlicht: (2024)
Modelling Visual Semantics via Image Captioning to extract Enhanced Multi-Level Cross-Modal Semantic Incongruity Representation with Attention for Multimodal Sarcasm Detection
von: Aggarwal, Sajal, et al.
Veröffentlicht: (2024)
von: Aggarwal, Sajal, et al.
Veröffentlicht: (2024)
Synthesize, Diagnose, and Optimize: Towards Fine-Grained Vision-Language Understanding
von: Peng, Wujian, et al.
Veröffentlicht: (2023)
von: Peng, Wujian, et al.
Veröffentlicht: (2023)
Any2Caption:Interpreting Any Condition to Caption for Controllable Video Generation
von: Wu, Shengqiong, et al.
Veröffentlicht: (2025)
von: Wu, Shengqiong, et al.
Veröffentlicht: (2025)
Beyond Pixels: Visual Metaphor Transfer via Schema-Driven Agentic Reasoning
von: Xu, Yu, et al.
Veröffentlicht: (2026)
von: Xu, Yu, et al.
Veröffentlicht: (2026)
From Pixels to Posts: Retrieval-Augmented Fashion Captioning and Hashtag Generation
von: Gondal, Moazzam Umer, et al.
Veröffentlicht: (2025)
von: Gondal, Moazzam Umer, et al.
Veröffentlicht: (2025)
Efficient-LVSM: Faster, Cheaper, and Better Large View Synthesis Model via Decoupled Co-Refinement Attention
von: Jia, Xiaosong, et al.
Veröffentlicht: (2026)
von: Jia, Xiaosong, et al.
Veröffentlicht: (2026)
On Explaining Visual Captioning with Hybrid Markov Logic Networks
von: Shah, Monika, et al.
Veröffentlicht: (2025)
von: Shah, Monika, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
FOCUS: Towards Universal Foreground Segmentation
von: You, Zuyao, et al.
Veröffentlicht: (2025) -
VideoLoom: A Video Large Language Model for Joint Spatial-Temporal Understanding
von: Shi, Jiapeng, et al.
Veröffentlicht: (2026) -
FluxMem: Adaptive Hierarchical Memory for Streaming Video Understanding
von: Xie, Yiweng, et al.
Veröffentlicht: (2026) -
Seg-R1: Segmentation Can Be Surprisingly Simple with Reinforcement Learning
von: You, Zuyao, et al.
Veröffentlicht: (2025) -
Pixelis: Reasoning in Pixels, from Seeing to Acting
von: Zhou, Yunpeng
Veröffentlicht: (2026)