Grounding Language in Multi-Perspective Referential Communication
Fuente:
arXiv
Saved in:
| Main Authors: | Tang, Zineng, Mao, Lingjun, Suhr, Alane |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evaluating Model Perception of Color Illusions in Photorealistic Scenes
by: Mao, Lingjun, et al.
Published: (2024)
by: Mao, Lingjun, et al.
Published: (2024)
TULIP: Towards Unified Language-Image Pretraining
by: Tang, Zineng, et al.
Published: (2025)
by: Tang, Zineng, et al.
Published: (2025)
SIMS: Simulating Stylized Human-Scene Interactions with Retrieval-Augmented Script Generation
by: Wang, Wenjia, et al.
Published: (2024)
by: Wang, Wenjia, et al.
Published: (2024)
VisionGPT-3D: A Generalized Multimodal Agent for Enhanced 3D Vision Understanding
by: Kelly, Chris, et al.
Published: (2024)
by: Kelly, Chris, et al.
Published: (2024)
DynamicGTR: Leveraging Graph Topology Representation Preferences to Boost VLM Capabilities on Graph QAs
by: Wei, Yanbin, et al.
Published: (2026)
by: Wei, Yanbin, et al.
Published: (2026)
Towards Understanding Graphical Perception in Large Multimodal Models
by: Zhang, Kai, et al.
Published: (2025)
by: Zhang, Kai, et al.
Published: (2025)
Generative Powers of Ten
by: Wang, Xiaojuan, et al.
Published: (2023)
by: Wang, Xiaojuan, et al.
Published: (2023)
Image Generation Models: A Technical History
by: Shirvani, Rouzbeh
Published: (2026)
by: Shirvani, Rouzbeh
Published: (2026)
Multi-LoRA Composition for Image Generation
by: Zhong, Ming, et al.
Published: (2024)
by: Zhong, Ming, et al.
Published: (2024)
Generating by Understanding: Neural Visual Generation with Logical Symbol Groundings
by: Peng, Yifei, et al.
Published: (2023)
by: Peng, Yifei, et al.
Published: (2023)
GAF: Gaussian Avatar Reconstruction from Monocular Videos via Multi-view Diffusion
by: Tang, Jiapeng, et al.
Published: (2024)
by: Tang, Jiapeng, et al.
Published: (2024)
StyleMotif: Multi-Modal Motion Stylization using Style-Content Cross Fusion
by: Guo, Ziyu, et al.
Published: (2025)
by: Guo, Ziyu, et al.
Published: (2025)
An Image is Worth Multiple Words: Discovering Object Level Concepts using Multi-Concept Prompt Learning
by: Jin, Chen, et al.
Published: (2023)
by: Jin, Chen, et al.
Published: (2023)
LSD-3D: Large-Scale 3D Driving Scene Generation with Geometry Grounding
by: Ost, Julian, et al.
Published: (2025)
by: Ost, Julian, et al.
Published: (2025)
Multimodal Cinematic Video Synthesis Using Text-to-Image and Audio Generation Models
by: S, Sridhar, et al.
Published: (2025)
by: S, Sridhar, et al.
Published: (2025)
MAPWise: Evaluating Vision-Language Models for Advanced Map Queries
by: Mukhopadhyay, Srija, et al.
Published: (2024)
by: Mukhopadhyay, Srija, et al.
Published: (2024)
Feed-Forward 3D Scene Modeling: A Problem-Driven Perspective
by: Wang, Weijie, et al.
Published: (2026)
by: Wang, Weijie, et al.
Published: (2026)
StyleRF-VolVis: Style Transfer of Neural Radiance Fields for Expressive Volume Visualization
by: Tang, Kaiyuan, et al.
Published: (2024)
by: Tang, Kaiyuan, et al.
Published: (2024)
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation
by: Song, Lin, et al.
Published: (2026)
by: Song, Lin, et al.
Published: (2026)
Learning Disentangled Speech- and Expression-Driven Blendshapes for 3D Talking Face Animation
by: Mao, Yuxiang, et al.
Published: (2025)
by: Mao, Yuxiang, et al.
Published: (2025)
Meta-INR: Efficient Encoding of Volumetric Data via Meta-Learning Implicit Neural Representation
by: Yang, Maizhe, et al.
Published: (2025)
by: Yang, Maizhe, et al.
Published: (2025)
Go-SLAM: Grounded Object Segmentation and Localization with Gaussian Splatting SLAM
by: Pham, Phu, et al.
Published: (2024)
by: Pham, Phu, et al.
Published: (2024)
LatentEdit: Adaptive Latent Control for Consistent Semantic Editing
by: Liu, Siyi, et al.
Published: (2025)
by: Liu, Siyi, et al.
Published: (2025)
WorldCraft: Photo-Realistic 3D World Creation and Customization via LLM Agents
by: Liu, Xinhang, et al.
Published: (2025)
by: Liu, Xinhang, et al.
Published: (2025)
VideoPanda: Video Panoramic Diffusion with Multi-view Attention
by: Xie, Kevin, et al.
Published: (2025)
by: Xie, Kevin, et al.
Published: (2025)
MVTN: Learning Multi-View Transformations for 3D Understanding
by: Hamdi, Abdullah, et al.
Published: (2022)
by: Hamdi, Abdullah, et al.
Published: (2022)
WonderZoom: Multi-Scale 3D World Generation
by: Cao, Jin, et al.
Published: (2025)
by: Cao, Jin, et al.
Published: (2025)
Exploring Multi-modal Neural Scene Representations With Applications on Thermal Imaging
by: Özer, Mert, et al.
Published: (2024)
by: Özer, Mert, et al.
Published: (2024)
Few-Shot Multi-Human Neural Rendering Using Geometry Constraints
by: li, Qian, et al.
Published: (2025)
by: li, Qian, et al.
Published: (2025)
MV-S2V: Multi-View Subject-Consistent Video Generation
by: Song, Ziyang, et al.
Published: (2026)
by: Song, Ziyang, et al.
Published: (2026)
Generative Object Insertion in Gaussian Splatting with a Multi-View Diffusion Model
by: Zhong, Hongliang, et al.
Published: (2024)
by: Zhong, Hongliang, et al.
Published: (2024)
Training-Free Text-Guided Color Editing with Multi-Modal Diffusion Transformer
by: Yin, Zixin, et al.
Published: (2025)
by: Yin, Zixin, et al.
Published: (2025)
MultiGen: Level-Design for Editable Multiplayer Worlds in Diffusion Game Engines
by: Po, Ryan, et al.
Published: (2026)
by: Po, Ryan, et al.
Published: (2026)
VideoMV: Consistent Multi-View Generation Based on Large Video Generative Model
by: Zuo, Qi, et al.
Published: (2024)
by: Zuo, Qi, et al.
Published: (2024)
MVImgNet2.0: A Larger-scale Dataset of Multi-view Images
by: Han, Xiaoguang, et al.
Published: (2024)
by: Han, Xiaoguang, et al.
Published: (2024)
DreamDrive: Generative 4D Scene Modeling from Street View Images
by: Mao, Jiageng, et al.
Published: (2024)
by: Mao, Jiageng, et al.
Published: (2024)
Gesture2Text: A Generalizable Decoder for Word-Gesture Keyboards in XR Through Trajectory Coarse Discretization and Pre-training
by: Shen, Junxiao, et al.
Published: (2024)
by: Shen, Junxiao, et al.
Published: (2024)
Stylus: Automatic Adapter Selection for Diffusion Models
by: Luo, Michael, et al.
Published: (2024)
by: Luo, Michael, et al.
Published: (2024)
Text-guided Controllable Mesh Refinement for Interactive 3D Modeling
by: Chen, Yun-Chun, et al.
Published: (2024)
by: Chen, Yun-Chun, et al.
Published: (2024)
Unbounded: A Generative Infinite Game of Character Life Simulation
by: Li, Jialu, et al.
Published: (2024)
by: Li, Jialu, et al.
Published: (2024)
Similar Items
-
Evaluating Model Perception of Color Illusions in Photorealistic Scenes
by: Mao, Lingjun, et al.
Published: (2024) -
TULIP: Towards Unified Language-Image Pretraining
by: Tang, Zineng, et al.
Published: (2025) -
SIMS: Simulating Stylized Human-Scene Interactions with Retrieval-Augmented Script Generation
by: Wang, Wenjia, et al.
Published: (2024) -
VisionGPT-3D: A Generalized Multimodal Agent for Enhanced 3D Vision Understanding
by: Kelly, Chris, et al.
Published: (2024) -
DynamicGTR: Leveraging Graph Topology Representation Preferences to Boost VLM Capabilities on Graph QAs
by: Wei, Yanbin, et al.
Published: (2026)