GRAPHGPT-O: Synergistic Multimodal Comprehension and Generation on Graphs
Fuente:
arXiv
Guardado en:
| Autores principales: | Fang, Yi, Jin, Bowen, Shen, Jiacheng, Ding, Sirui, Tan, Qiaoyu, Han, Jiawei |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents
por: Luo, Yaxin, et al.
Publicado: (2025)
por: Luo, Yaxin, et al.
Publicado: (2025)
InstructG2I: Synthesizing Images from Multimodal Attributed Graphs
por: Jin, Bowen, et al.
Publicado: (2024)
por: Jin, Bowen, et al.
Publicado: (2024)
InteractiveVideo: User-Centric Controllable Video Generation with Synergistic Multimodal Instructions
por: Zhang, Yiyuan, et al.
Publicado: (2024)
por: Zhang, Yiyuan, et al.
Publicado: (2024)
When Graph meets Multimodal: Benchmarking and Meditating on Multimodal Attributed Graphs Learning
por: Yan, Hao, et al.
Publicado: (2024)
por: Yan, Hao, et al.
Publicado: (2024)
Enhancing Scene Graph Generation with Hierarchical Relationships and Commonsense Knowledge
por: Jiang, Bowen, et al.
Publicado: (2023)
por: Jiang, Bowen, et al.
Publicado: (2023)
Uncovering and Mitigating Transient Blindness in Multimodal Model Editing
por: Han, Xiaoqi, et al.
Publicado: (2025)
por: Han, Xiaoqi, et al.
Publicado: (2025)
InfMasking: Unleashing Synergistic Information by Contrastive Multimodal Interactions
por: Wen, Liangjian, et al.
Publicado: (2025)
por: Wen, Liangjian, et al.
Publicado: (2025)
DenseMLLM: Standard Multimodal LLMs for Dense Prediction
por: Li, Yi, et al.
Publicado: (2026)
por: Li, Yi, et al.
Publicado: (2026)
FedAFD: Multimodal Federated Learning via Adversarial Fusion and Distillation
por: Tan, Min, et al.
Publicado: (2026)
por: Tan, Min, et al.
Publicado: (2026)
BTC-LLM: Efficient Sub-1-Bit LLM Quantization via Learnable Transformation and Binary Codebook
por: Gu, Hao, et al.
Publicado: (2025)
por: Gu, Hao, et al.
Publicado: (2025)
Token Activation Map to Visually Explain Multimodal LLMs
por: Li, Yi, et al.
Publicado: (2025)
por: Li, Yi, et al.
Publicado: (2025)
MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs
por: Fu, Chaoyou, et al.
Publicado: (2024)
por: Fu, Chaoyou, et al.
Publicado: (2024)
GraphVLM: Benchmarking Vision Language Models for Multimodal Graph Learning
por: Liu, Jiajin, et al.
Publicado: (2026)
por: Liu, Jiajin, et al.
Publicado: (2026)
Are We Making Progress in Multimodal Domain Generalization? A Comprehensive Benchmark Study
por: Dong, Hao, et al.
Publicado: (2026)
por: Dong, Hao, et al.
Publicado: (2026)
V-Zero: Self-Improving Multimodal Reasoning with Zero Annotation
por: Wang, Han, et al.
Publicado: (2026)
por: Wang, Han, et al.
Publicado: (2026)
PAIR-Diffusion: A Comprehensive Multimodal Object-Level Image Editor
por: Goel, Vidit, et al.
Publicado: (2023)
por: Goel, Vidit, et al.
Publicado: (2023)
TRIP-Evaluate: An Open Multimodal Benchmark for Evaluating Large Models in Transportation
por: Gong, Han, et al.
Publicado: (2026)
por: Gong, Han, et al.
Publicado: (2026)
A Comprehensive Survey on Test-Time Adaptation under Distribution Shifts
por: Liang, Jian, et al.
Publicado: (2023)
por: Liang, Jian, et al.
Publicado: (2023)
Woodpecker: Hallucination Correction for Multimodal Large Language Models
por: Yin, Shukang, et al.
Publicado: (2023)
por: Yin, Shukang, et al.
Publicado: (2023)
A Survey on Multimodal Large Language Models
por: Yin, Shukang, et al.
Publicado: (2023)
por: Yin, Shukang, et al.
Publicado: (2023)
A Frustratingly Simple Yet Highly Effective Attack Baseline: Over 90% Success Rate Against the Strong Black-box Models of GPT-4.5/4o/o1
por: Li, Zhaoyi, et al.
Publicado: (2025)
por: Li, Zhaoyi, et al.
Publicado: (2025)
A Comprehensive Survey of Foundation Models in Medicine
por: Khan, Wasif, et al.
Publicado: (2024)
por: Khan, Wasif, et al.
Publicado: (2024)
A Comprehensive Survey of Forgetting in Deep Learning Beyond Continual Learning
por: Wang, Zhenyi, et al.
Publicado: (2023)
por: Wang, Zhenyi, et al.
Publicado: (2023)
Comprehensive Exploration of Synthetic Data Generation: A Survey
por: Bauer, André, et al.
Publicado: (2024)
por: Bauer, André, et al.
Publicado: (2024)
Graph Your Own Prompt
por: Ding, Xi, et al.
Publicado: (2025)
por: Ding, Xi, et al.
Publicado: (2025)
From Masks to Pixels and Meaning: A New Taxonomy, Benchmark, and Metrics for VLM Image Tampering
por: Shang, Xinyi, et al.
Publicado: (2026)
por: Shang, Xinyi, et al.
Publicado: (2026)
HGTDP-DTA: Hybrid Graph-Transformer with Dynamic Prompt for Drug-Target Binding Affinity Prediction
por: Xiao, Xi, et al.
Publicado: (2024)
por: Xiao, Xi, et al.
Publicado: (2024)
UniGLM: Training One Unified Language Model for Text-Attributed Graph Embedding
por: Fang, Yi, et al.
Publicado: (2024)
por: Fang, Yi, et al.
Publicado: (2024)
Enhancing Vision-Language Model Reliability with Uncertainty-Guided Dropout Decoding
por: Fang, Yixiong, et al.
Publicado: (2024)
por: Fang, Yixiong, et al.
Publicado: (2024)
Contrastive Learning Is Spectral Clustering On Similarity Graph
por: Tan, Zhiquan, et al.
Publicado: (2023)
por: Tan, Zhiquan, et al.
Publicado: (2023)
PunchBench: Benchmarking MLLMs in Multimodal Punchline Comprehension
por: Ouyang, Kun, et al.
Publicado: (2024)
por: Ouyang, Kun, et al.
Publicado: (2024)
Masked Generative Extractor for Synergistic Representation and 3D Generation of Point Clouds
por: Zeng, Hongliang, et al.
Publicado: (2024)
por: Zeng, Hongliang, et al.
Publicado: (2024)
MultiTrust: A Comprehensive Benchmark Towards Trustworthy Multimodal Large Language Models
por: Zhang, Yichi, et al.
Publicado: (2024)
por: Zhang, Yichi, et al.
Publicado: (2024)
MMMG: a Comprehensive and Reliable Evaluation Suite for Multitask Multimodal Generation
por: Yao, Jihan, et al.
Publicado: (2025)
por: Yao, Jihan, et al.
Publicado: (2025)
Causal Decoding for Hallucination-Resistant Multimodal Large Language Models
por: Tan, Shiwei, et al.
Publicado: (2026)
por: Tan, Shiwei, et al.
Publicado: (2026)
Can Multimodal Large Language Models Truly Perform Multimodal In-Context Learning?
por: Chen, Shuo, et al.
Publicado: (2023)
por: Chen, Shuo, et al.
Publicado: (2023)
GHOST: Hallucination-Inducing Image Generation for Multimodal LLMs
por: Parast, Aryan Yazdan, et al.
Publicado: (2025)
por: Parast, Aryan Yazdan, et al.
Publicado: (2025)
Rethinking Multi-domain Generalization with A General Learning Objective
por: Tan, Zhaorui, et al.
Publicado: (2024)
por: Tan, Zhaorui, et al.
Publicado: (2024)
Knowledge Graphs Meet Multi-Modal Learning: A Comprehensive Survey
por: Chen, Zhuo, et al.
Publicado: (2024)
por: Chen, Zhuo, et al.
Publicado: (2024)
Constrained Layout Generation with Factor Graphs
por: Dupty, Mohammed Haroon, et al.
Publicado: (2024)
por: Dupty, Mohammed Haroon, et al.
Publicado: (2024)
Ejemplares similares
-
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents
por: Luo, Yaxin, et al.
Publicado: (2025) -
InstructG2I: Synthesizing Images from Multimodal Attributed Graphs
por: Jin, Bowen, et al.
Publicado: (2024) -
InteractiveVideo: User-Centric Controllable Video Generation with Synergistic Multimodal Instructions
por: Zhang, Yiyuan, et al.
Publicado: (2024) -
When Graph meets Multimodal: Benchmarking and Meditating on Multimodal Attributed Graphs Learning
por: Yan, Hao, et al.
Publicado: (2024) -
Enhancing Scene Graph Generation with Hierarchical Relationships and Commonsense Knowledge
por: Jiang, Bowen, et al.
Publicado: (2023)