Towards Understanding Graphical Perception in Large Multimodal Models
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Kai, Yang, Jianwei, Inala, Jeevana Priya, Singh, Chandan, Gao, Jianfeng, Su, Yu, Wang, Chenglong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Rethinking Interpretability in the Era of Large Language Models
by: Singh, Chandan, et al.
Published: (2024)
by: Singh, Chandan, et al.
Published: (2024)
Chat2SVG: Vector Graphics Generation with Large Language Models and Image Diffusion Models
by: Wu, Ronghuan, et al.
Published: (2024)
by: Wu, Ronghuan, et al.
Published: (2024)
VisionGPT-3D: A Generalized Multimodal Agent for Enhanced 3D Vision Understanding
by: Kelly, Chris, et al.
Published: (2024)
by: Kelly, Chris, et al.
Published: (2024)
Is Your World Simulator a Good Story Presenter? A Consecutive Events-Based Benchmark for Future Long Video Generation
by: Wang, Yiping, et al.
Published: (2024)
by: Wang, Yiping, et al.
Published: (2024)
ArchGPT: Understanding the World's Architectures with Large Multimodal Models
by: Wang, Yuze, et al.
Published: (2025)
by: Wang, Yuze, et al.
Published: (2025)
Global Position Aware Group Choreography using Large Language Model
by: Pang, Haozhou, et al.
Published: (2025)
by: Pang, Haozhou, et al.
Published: (2025)
A Study of the Framework and Real-World Applications of Language Embedding for 3D Scene Understanding
by: Zaouali, Mahmoud Chick, et al.
Published: (2025)
by: Zaouali, Mahmoud Chick, et al.
Published: (2025)
Textured Mesh Saliency: Bridging Geometry and Texture for Human Perception in 3D Graphics
by: Zhang, Kaiwei, et al.
Published: (2024)
by: Zhang, Kaiwei, et al.
Published: (2024)
OpenCOLE: Towards Reproducible Automatic Graphic Design Generation
by: Inoue, Naoto, et al.
Published: (2024)
by: Inoue, Naoto, et al.
Published: (2024)
BlenderAlchemy: Editing 3D Graphics with Vision-Language Models
by: Huang, Ian, et al.
Published: (2024)
by: Huang, Ian, et al.
Published: (2024)
Towards Interactive Intelligence for Digital Humans
by: Cai, Yiyi, et al.
Published: (2025)
by: Cai, Yiyi, et al.
Published: (2025)
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation
by: Song, Lin, et al.
Published: (2026)
by: Song, Lin, et al.
Published: (2026)
Is Self-Repair a Silver Bullet for Code Generation?
by: Olausson, Theo X., et al.
Published: (2023)
by: Olausson, Theo X., et al.
Published: (2023)
OmniX: From Unified Panoramic Generation and Perception to Graphics-Ready 3D Scenes
by: Huang, Yukun, et al.
Published: (2025)
by: Huang, Yukun, et al.
Published: (2025)
MeshLLM: Empowering Large Language Models to Progressively Understand and Generate 3D Mesh
by: Fang, Shuangkang, et al.
Published: (2025)
by: Fang, Shuangkang, et al.
Published: (2025)
Towards Understanding Depth Perception in Foveated Rendering
by: Kergaßner, Sophie, et al.
Published: (2025)
by: Kergaßner, Sophie, et al.
Published: (2025)
MeshLRM: Large Reconstruction Model for High-Quality Meshes
by: Wei, Xinyue, et al.
Published: (2024)
by: Wei, Xinyue, et al.
Published: (2024)
Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning
by: Yin, Shaofeng, et al.
Published: (2026)
by: Yin, Shaofeng, et al.
Published: (2026)
Establishing Stochastic Object Models from Noisy Data via Ambient Measurement-Integrated Diffusion
by: Lei, Xiaoning, et al.
Published: (2025)
by: Lei, Xiaoning, et al.
Published: (2025)
Fast Sprite Decomposition from Animated Graphics
by: Suzuki, Tomoyuki, et al.
Published: (2024)
by: Suzuki, Tomoyuki, et al.
Published: (2024)
Multimodal Cinematic Video Synthesis Using Text-to-Image and Audio Generation Models
by: S, Sridhar, et al.
Published: (2025)
by: S, Sridhar, et al.
Published: (2025)
MoVer: Motion Verification for Motion Graphics Animations
by: Ma, Jiaju, et al.
Published: (2025)
by: Ma, Jiaju, et al.
Published: (2025)
MG-Gen: Single Image to Motion Graphics Generation
by: Shirakawa, Takahiro, et al.
Published: (2025)
by: Shirakawa, Takahiro, et al.
Published: (2025)
LayerD: Decomposing Raster Graphic Designs into Layers
by: Suzuki, Tomoyuki, et al.
Published: (2025)
by: Suzuki, Tomoyuki, et al.
Published: (2025)
Bezier Splatting for Fast and Differentiable Vector Graphics Rendering
by: Liu, Xi, et al.
Published: (2025)
by: Liu, Xi, et al.
Published: (2025)
ReverBERT: A State Space Model for Efficient Text-Driven Speech Style Transfer
by: Brown, Michael, et al.
Published: (2025)
by: Brown, Michael, et al.
Published: (2025)
CGVQM+D: Computer Graphics Video Quality Metric and Dataset
by: Jindal, Akshay, et al.
Published: (2025)
by: Jindal, Akshay, et al.
Published: (2025)
DesigNet: Learning to Draw Vector Graphics as Designers Do
by: Guija-Valiente, Tomas, et al.
Published: (2026)
by: Guija-Valiente, Tomas, et al.
Published: (2026)
Can GPTs Evaluate Graphic Design Based on Design Principles?
by: Haraguchi, Daichi, et al.
Published: (2024)
by: Haraguchi, Daichi, et al.
Published: (2024)
FIFA: Unified Faithfulness Evaluation Framework for Text-to-Video and Video-to-Text Generation
by: Jing, Liqiang, et al.
Published: (2025)
by: Jing, Liqiang, et al.
Published: (2025)
Is this chart lying to me? Automating the detection of misleading visualizations
by: Tonglet, Jonathan, et al.
Published: (2025)
by: Tonglet, Jonathan, et al.
Published: (2025)
ComfyGen: Prompt-Adaptive Workflows for Text-to-Image Generation
by: Gal, Rinon, et al.
Published: (2024)
by: Gal, Rinon, et al.
Published: (2024)
TexGS-VolVis: Expressive Scene Editing for Volume Visualization via Textured Gaussian Splatting
by: Tang, Kaiyuan, et al.
Published: (2025)
by: Tang, Kaiyuan, et al.
Published: (2025)
FlairGPT: Repurposing LLMs for Interior Designs
by: Littlefair, Gabrielle, et al.
Published: (2025)
by: Littlefair, Gabrielle, et al.
Published: (2025)
Co-Layout: LLM-driven Co-optimization for Interior Layout
by: Xiang, Chucheng, et al.
Published: (2025)
by: Xiang, Chucheng, et al.
Published: (2025)
T$^3$-S2S: Training-free Triplet Tuning for Sketch to Scene Synthesis in Controllable Concept Art Generation
by: Sun, Zhenhong, et al.
Published: (2024)
by: Sun, Zhenhong, et al.
Published: (2024)
CAP: Evaluation of Persuasive and Creative Image Generation
by: Aghazadeh, Aysan, et al.
Published: (2024)
by: Aghazadeh, Aysan, et al.
Published: (2024)
Optimizing ID Consistency in Multimodal Large Models: Facial Restoration via Alignment, Entanglement, and Disentanglement
by: Dong, Yuran, et al.
Published: (2026)
by: Dong, Yuran, et al.
Published: (2026)
Aligning Human Motion Generation with Human Perceptions
by: Wang, Haoru, et al.
Published: (2024)
by: Wang, Haoru, et al.
Published: (2024)
Mesh-based Gaussian Splatting for Real-time Large-scale Deformation
by: Gao, Lin, et al.
Published: (2024)
by: Gao, Lin, et al.
Published: (2024)
Similar Items
-
Rethinking Interpretability in the Era of Large Language Models
by: Singh, Chandan, et al.
Published: (2024) -
Chat2SVG: Vector Graphics Generation with Large Language Models and Image Diffusion Models
by: Wu, Ronghuan, et al.
Published: (2024) -
VisionGPT-3D: A Generalized Multimodal Agent for Enhanced 3D Vision Understanding
by: Kelly, Chris, et al.
Published: (2024) -
Is Your World Simulator a Good Story Presenter? A Consecutive Events-Based Benchmark for Future Long Video Generation
by: Wang, Yiping, et al.
Published: (2024) -
ArchGPT: Understanding the World's Architectures with Large Multimodal Models
by: Wang, Yuze, et al.
Published: (2025)