Aligning Forest and Trees in Images & Long Captions for Visually Grounded Understanding
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Woo, Byeongju, Wang, Zilin, Pak, Byeonghyun, Mo, Sangwoo, Yu, Stella X. |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Pixel-level Scene Understanding in One Token: Visual States Need What-is-Where Composition
par: Lee, Seokmin, et autres
Publié: (2026)
par: Lee, Seokmin, et autres
Publié: (2026)
Rethinking FID Through the Geometry of the Reference Dataset
par: Lee, Yunghee, et autres
Publié: (2026)
par: Lee, Yunghee, et autres
Publié: (2026)
Open Ad-hoc Categorization with Contextualized Feature Learning
par: Wang, Zilin, et autres
Publié: (2025)
par: Wang, Zilin, et autres
Publié: (2025)
Textual Query-Driven Mask Transformer for Domain Generalized Segmentation
par: Pak, Byeonghyun, et autres
Publié: (2024)
par: Pak, Byeonghyun, et autres
Publié: (2024)
Learning Hierarchical Image Segmentation For Recognition and By Recognition
par: Ke, Tsung-Wei, et autres
Publié: (2022)
par: Ke, Tsung-Wei, et autres
Publié: (2022)
Paint Outside the Box: Synthesizing and Selecting Training Data for Visual Grounding
par: Du, Zilin, et autres
Publié: (2024)
par: Du, Zilin, et autres
Publié: (2024)
Understanding the Effects of Distractors on Reasoning Vision-Language Models
par: Bae, Jiyun, et autres
Publié: (2025)
par: Bae, Jiyun, et autres
Publié: (2025)
ViPCap: Retrieval Text-Based Visual Prompts for Lightweight Image Captioning
par: Kim, Taewhan, et autres
Publié: (2024)
par: Kim, Taewhan, et autres
Publié: (2024)
Graph-Based Captioning: Enhancing Visual Descriptions by Interconnecting Region Captions
par: Hsieh, Yu-Guan, et autres
Publié: (2024)
par: Hsieh, Yu-Guan, et autres
Publié: (2024)
Align Your Query: Representation Alignment for Multimodality Medical Object Detection
par: Seo, Ara, et autres
Publié: (2025)
par: Seo, Ara, et autres
Publié: (2025)
TALC: Time-Aligned Captions for Multi-Scene Text-to-Video Generation
par: Bansal, Hritik, et autres
Publié: (2024)
par: Bansal, Hritik, et autres
Publié: (2024)
SynC: Synthetic Image Caption Dataset Refinement with One-to-many Mapping for Zero-shot Image Captioning
par: Kim, Si-Woo, et autres
Publié: (2025)
par: Kim, Si-Woo, et autres
Publié: (2025)
Generalizable Geometric Image Caption Synthesis
par: Xin, Yue, et autres
Publié: (2025)
par: Xin, Yue, et autres
Publié: (2025)
Seeing the Trees for the Forest: Rethinking Weakly-Supervised Medical Visual Grounding
par: Huy, Ta Duc, et autres
Publié: (2025)
par: Huy, Ta Duc, et autres
Publié: (2025)
Image Captions are Natural Prompts for Text-to-Image Models
par: Lei, Shiye, et autres
Publié: (2023)
par: Lei, Shiye, et autres
Publié: (2023)
LSPT: Long-term Spatial Prompt Tuning for Visual Representation Learning
par: Mo, Shentong, et autres
Publié: (2024)
par: Mo, Shentong, et autres
Publié: (2024)
Controllable Hybrid Captioner for Improved Long-form Video Understanding
par: Sasse, Kuleen, et autres
Publié: (2025)
par: Sasse, Kuleen, et autres
Publié: (2025)
Sample Selection via Contrastive Fragmentation for Noisy Label Regression
par: Kim, Chris Dongjoo, et autres
Publié: (2025)
par: Kim, Chris Dongjoo, et autres
Publié: (2025)
Emergent Visual Grounding in Large Multimodal Models Without Grounding Supervision
par: Cao, Shengcao, et autres
Publié: (2024)
par: Cao, Shengcao, et autres
Publié: (2024)
Divide, then Ground: Adapting Frame Selection to Query Types for Long-Form Video Understanding
par: Li, Jialuo, et autres
Publié: (2025)
par: Li, Jialuo, et autres
Publié: (2025)
Tiled Prompts: Overcoming Prompt Misguidance in Image and Video Super-Resolution
par: Kim, Bryan Sangwoo, et autres
Publié: (2026)
par: Kim, Bryan Sangwoo, et autres
Publié: (2026)
Extreme Blind Image Restoration via Prompt-Conditioned Information Bottleneck
par: Kim, Hongeun, et autres
Publié: (2025)
par: Kim, Hongeun, et autres
Publié: (2025)
Co-domain Symmetry for Complex-Valued Deep Learning
par: Singhal, Utkarsh, et autres
Publié: (2021)
par: Singhal, Utkarsh, et autres
Publié: (2021)
IFCap: Image-like Retrieval and Frequency-based Entity Filtering for Zero-shot Captioning
par: Lee, Soeun, et autres
Publié: (2024)
par: Lee, Soeun, et autres
Publié: (2024)
Free$^2$Guide: Training-Free Text-to-Video Alignment using Image LVLM
par: Kim, Jaemin, et autres
Publié: (2024)
par: Kim, Jaemin, et autres
Publié: (2024)
LAViTeR: Learning Aligned Visual and Textual Representations Assisted by Image and Caption Generation
par: Hashemi, Mohammad Abuzar, et autres
Publié: (2021)
par: Hashemi, Mohammad Abuzar, et autres
Publié: (2021)
VeCLIP: Improving CLIP Training via Visual-enriched Captions
par: Lai, Zhengfeng, et autres
Publié: (2023)
par: Lai, Zhengfeng, et autres
Publié: (2023)
Curvature-Aware Captioning:Leveraging Geodesic Attention for 3D Scene Understanding
par: He, Ziyao, et autres
Publié: (2026)
par: He, Ziyao, et autres
Publié: (2026)
IG Captioner: Information Gain Captioners are Strong Zero-shot Classifiers
par: Yang, Chenglin, et autres
Publié: (2023)
par: Yang, Chenglin, et autres
Publié: (2023)
Semi-Supervised Image Captioning Considering Wasserstein Graph Matching
par: Yang, Yang
Publié: (2024)
par: Yang, Yang
Publié: (2024)
Towards Understanding Visual Grounding in Visual Language Models
par: Pantazopoulos, Georgios, et autres
Publié: (2025)
par: Pantazopoulos, Georgios, et autres
Publié: (2025)
Detecting and Understanding Hateful Contents in Memes Through Captioning and Visual Question-Answering
par: Anaissi, Ali, et autres
Publié: (2025)
par: Anaissi, Ali, et autres
Publié: (2025)
Hyperdimensional Cross-Modal Alignment of Frozen Language and Image Models for Efficient Image Captioning
par: Dalvi, Abhishek, et autres
Publié: (2026)
par: Dalvi, Abhishek, et autres
Publié: (2026)
LVD-2M: A Long-take Video Dataset with Temporally Dense Captions
par: Xiong, Tianwei, et autres
Publié: (2024)
par: Xiong, Tianwei, et autres
Publié: (2024)
Caption-Driven Explorations: Aligning Image and Text Embeddings through Human-Inspired Foveated Vision
par: Zanca, Dario, et autres
Publié: (2024)
par: Zanca, Dario, et autres
Publié: (2024)
RubiCap: Rubric-Guided Reinforcement Learning for Dense Image Captioning
par: Huang, Tzu-Heng, et autres
Publié: (2026)
par: Huang, Tzu-Heng, et autres
Publié: (2026)
Aligning Audio-Visual Joint Representations with an Agentic Workflow
par: Mo, Shentong, et autres
Publié: (2024)
par: Mo, Shentong, et autres
Publié: (2024)
Pix2Cap-COCO: Advancing Visual Comprehension via Pixel-Level Captioning
par: You, Zuyao, et autres
Publié: (2025)
par: You, Zuyao, et autres
Publié: (2025)
LVRPO: Language-Visual Alignment with GRPO for Multimodal Understanding and Generation
par: Mo, Shentong, et autres
Publié: (2026)
par: Mo, Shentong, et autres
Publié: (2026)
CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects
par: Fiastre, Gabriel, et autres
Publié: (2025)
par: Fiastre, Gabriel, et autres
Publié: (2025)
Documents similaires
-
Pixel-level Scene Understanding in One Token: Visual States Need What-is-Where Composition
par: Lee, Seokmin, et autres
Publié: (2026) -
Rethinking FID Through the Geometry of the Reference Dataset
par: Lee, Yunghee, et autres
Publié: (2026) -
Open Ad-hoc Categorization with Contextualized Feature Learning
par: Wang, Zilin, et autres
Publié: (2025) -
Textual Query-Driven Mask Transformer for Domain Generalized Segmentation
par: Pak, Byeonghyun, et autres
Publié: (2024) -
Learning Hierarchical Image Segmentation For Recognition and By Recognition
par: Ke, Tsung-Wei, et autres
Publié: (2022)