Pink: Unveiling the Power of Referential Comprehension for Multi-modal LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Xuan, Shiyu, Guo, Qingpei, Yang, Ming, Zhang, Shiliang |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FlattenGPT: Depth Compression for Transformer with Layer Flattening
by: Xu, Ruihan, et al.
Published: (2026)
by: Xu, Ruihan, et al.
Published: (2026)
SciVerse: Unveiling the Knowledge Comprehension and Visual Reasoning of LMMs on Multi-modal Scientific Problems
by: Guo, Ziyu, et al.
Published: (2025)
by: Guo, Ziyu, et al.
Published: (2025)
SyCoCa: Symmetrizing Contrastive Captioners with Attentive Masking for Multimodal Alignment
by: Ma, Ziping, et al.
Published: (2024)
by: Ma, Ziping, et al.
Published: (2024)
Multi-modal Generative AI: Multi-modal LLMs, Diffusions, and the Unification
by: Wang, Xin, et al.
Published: (2024)
by: Wang, Xin, et al.
Published: (2024)
Artemis: Towards Referential Understanding in Complex Videos
by: Qiu, Jihao, et al.
Published: (2024)
by: Qiu, Jihao, et al.
Published: (2024)
Recent Advances in Multi-modal 3D Intelligence: A Comprehensive Survey and Evaluation
by: Lei, Yinjie, et al.
Published: (2023)
by: Lei, Yinjie, et al.
Published: (2023)
Decoupled Contrastive Learning for Long-Tailed Recognition
by: Xuan, Shiyu, et al.
Published: (2024)
by: Xuan, Shiyu, et al.
Published: (2024)
Grounding Language in Multi-Perspective Referential Communication
by: Tang, Zineng, et al.
Published: (2024)
by: Tang, Zineng, et al.
Published: (2024)
Referencing Where to Focus: Improving VisualGrounding with Referential Query
by: Wang, Yabing, et al.
Published: (2024)
by: Wang, Yabing, et al.
Published: (2024)
Can Multi-modal (reasoning) LLMs work as deepfake detectors?
by: Ren, Simiao, et al.
Published: (2025)
by: Ren, Simiao, et al.
Published: (2025)
LaViT: Aligning Latent Visual Thoughts for Multi-modal Reasoning
by: Wu, Linquan, et al.
Published: (2026)
by: Wu, Linquan, et al.
Published: (2026)
M2-Encoder: Advancing Bilingual Image-Text Understanding by Large-scale Efficient Pretraining
by: Guo, Qingpei, et al.
Published: (2024)
by: Guo, Qingpei, et al.
Published: (2024)
SimVG: A Simple Framework for Visual Grounding with Decoupled Multi-modal Fusion
by: Dai, Ming, et al.
Published: (2024)
by: Dai, Ming, et al.
Published: (2024)
Unveiling the Power of Self-supervision for Multi-view Multi-human Association and Tracking
by: Feng, Wei, et al.
Published: (2024)
by: Feng, Wei, et al.
Published: (2024)
HOTVCOM: Generating Buzzworthy Comments for Videos
by: Chen, Yuyan, et al.
Published: (2024)
by: Chen, Yuyan, et al.
Published: (2024)
According to Me: Long-Term Personalized Referential Memory QA
by: Mei, Jingbiao, et al.
Published: (2026)
by: Mei, Jingbiao, et al.
Published: (2026)
Incomplete Multi-View Multi-Label Classification via Shared Codebook and Fused-Teacher Self-Distillation
by: Yan, Xu, et al.
Published: (2026)
by: Yan, Xu, et al.
Published: (2026)
EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents
by: Yang, Rui, et al.
Published: (2025)
by: Yang, Rui, et al.
Published: (2025)
Distilling Neuro-Symbolic Programs into 3D Multi-modal LLMs
by: Mo, Wentao, et al.
Published: (2026)
by: Mo, Wentao, et al.
Published: (2026)
Face-Human-Bench: A Comprehensive Benchmark of Face and Human Understanding for Multi-modal Assistants
by: Qin, Lixiong, et al.
Published: (2025)
by: Qin, Lixiong, et al.
Published: (2025)
Awesome Multi-modal Object Tracking
by: Zhang, Chunhui, et al.
Published: (2024)
by: Zhang, Chunhui, et al.
Published: (2024)
PixelGen: Improving Pixel Diffusion with Perceptual Supervision
by: Ma, Zehong, et al.
Published: (2026)
by: Ma, Zehong, et al.
Published: (2026)
Explaining multimodal LLMs via intra-modal token interactions
by: Liang, Jiawei, et al.
Published: (2025)
by: Liang, Jiawei, et al.
Published: (2025)
LLMTrack: Semantic Multi-Object Tracking with Multi-modal Large Language Models
by: Liao, Pan, et al.
Published: (2026)
by: Liao, Pan, et al.
Published: (2026)
Knowledge-enhanced Multi-perspective Video Representation Learning for Scene Recognition
by: Yu, Xuzheng, et al.
Published: (2024)
by: Yu, Xuzheng, et al.
Published: (2024)
LLM-RG: Referential Grounding in Outdoor Scenarios using Large Language Models
by: Saxena, Pranav, et al.
Published: (2025)
by: Saxena, Pranav, et al.
Published: (2025)
Robust Domain Generalization for Multi-modal Object Recognition
by: Qiao, Yuxin, et al.
Published: (2024)
by: Qiao, Yuxin, et al.
Published: (2024)
Large Multi-modal Model Cartographic Map Comprehension for Textual Locality Georeferencing
by: Wijegunarathna, Kalana, et al.
Published: (2025)
by: Wijegunarathna, Kalana, et al.
Published: (2025)
SPARROW: Learning Spatial Precision and Temporal Referential Consistency in Pixel-Grounded Video MLLMs
by: Alansari, Mohamad, et al.
Published: (2026)
by: Alansari, Mohamad, et al.
Published: (2026)
The Labyrinth of Links: Navigating the Associative Maze of Multi-modal LLMs
by: Li, Hong, et al.
Published: (2024)
by: Li, Hong, et al.
Published: (2024)
MulCPred: Learning Multi-modal Concepts for Explainable Pedestrian Action Prediction
by: Feng, Yan, et al.
Published: (2024)
by: Feng, Yan, et al.
Published: (2024)
Generative RLHF-V: Learning Principles from Multi-modal Human Preference
by: Zhou, Jiayi, et al.
Published: (2025)
by: Zhou, Jiayi, et al.
Published: (2025)
Hierarchical Multi-modal Transformer for Cross-modal Long Document Classification
by: Liu, Tengfei, et al.
Published: (2024)
by: Liu, Tengfei, et al.
Published: (2024)
TrajFlow: Multi-modal Motion Prediction via Flow Matching
by: Yan, Qi, et al.
Published: (2025)
by: Yan, Qi, et al.
Published: (2025)
LocLLM: Exploiting Generalizable Human Keypoint Localization via Large Language Model
by: Wang, Dongkai, et al.
Published: (2024)
by: Wang, Dongkai, et al.
Published: (2024)
ImagebindDC: Compressing Multi-modal Data with Imagebind-based Condensation
by: Min, Yue, et al.
Published: (2025)
by: Min, Yue, et al.
Published: (2025)
Knowledge Graph Enhanced Generative Multi-modal Models for Class-Incremental Learning
by: Cao, Xusheng, et al.
Published: (2025)
by: Cao, Xusheng, et al.
Published: (2025)
Modality-Aware and Shift Mixer for Multi-modal Brain Tumor Segmentation
by: Huang, Zhongzhen, et al.
Published: (2024)
by: Huang, Zhongzhen, et al.
Published: (2024)
VideoScaffold: Elastic-Scale Visual Hierarchies for Streaming Video Understanding in MLLMs
by: Zheng, Naishan, et al.
Published: (2025)
by: Zheng, Naishan, et al.
Published: (2025)
Multi-modality Anomaly Segmentation on the Road
by: Gao, Heng, et al.
Published: (2025)
by: Gao, Heng, et al.
Published: (2025)
Similar Items
-
FlattenGPT: Depth Compression for Transformer with Layer Flattening
by: Xu, Ruihan, et al.
Published: (2026) -
SciVerse: Unveiling the Knowledge Comprehension and Visual Reasoning of LMMs on Multi-modal Scientific Problems
by: Guo, Ziyu, et al.
Published: (2025) -
SyCoCa: Symmetrizing Contrastive Captioners with Attentive Masking for Multimodal Alignment
by: Ma, Ziping, et al.
Published: (2024) -
Multi-modal Generative AI: Multi-modal LLMs, Diffusions, and the Unification
by: Wang, Xin, et al.
Published: (2024) -
Artemis: Towards Referential Understanding in Complex Videos
by: Qiu, Jihao, et al.
Published: (2024)