Communicating about Space: Language-Mediated Spatial Integration Across Partial Views
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sikarwar, Ankur, Mishra, Debangan, Nikhil, Sudarshan, Kumaraguru, Ponnurangam, Agrawal, Aishwarya |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
How and What to Imagine? Visual Thinking in Unified Multimodal Models for Cross-View Spatial Reasoning
von: Yang, Qian, et al.
Veröffentlicht: (2026)
von: Yang, Qian, et al.
Veröffentlicht: (2026)
SPIRIT: Short-term Prediction of solar IRradIance for zero-shot Transfer learning using Foundation Models
von: Mishra, Aditya, et al.
Veröffentlicht: (2025)
von: Mishra, Aditya, et al.
Veröffentlicht: (2025)
RiT: Vanilla Diffusion Transformers Suffice in Representation Space
von: Zhang, Le, et al.
Veröffentlicht: (2026)
von: Zhang, Le, et al.
Veröffentlicht: (2026)
What if I ask in \textit{alia lingua}? Measuring Functional Similarity Across Languages
von: Mishra, Debangan, et al.
Veröffentlicht: (2025)
von: Mishra, Debangan, et al.
Veröffentlicht: (2025)
Assessing and Learning Alignment of Unimodal Vision and Language Models
von: Zhang, Le, et al.
Veröffentlicht: (2024)
von: Zhang, Le, et al.
Veröffentlicht: (2024)
The Promise of RL for Autoregressive Image Editing
von: Ahmadi, Saba, et al.
Veröffentlicht: (2025)
von: Ahmadi, Saba, et al.
Veröffentlicht: (2025)
Freeze and Reveal: Exposing Modality Bias in Vision-Language Models
von: Kavuri, Vivek Hruday, et al.
Veröffentlicht: (2025)
von: Kavuri, Vivek Hruday, et al.
Veröffentlicht: (2025)
Contrasting Intra-Modal and Ranking Cross-Modal Hard Negatives to Enhance Visio-Linguistic Compositional Understanding
von: Zhang, Le, et al.
Veröffentlicht: (2023)
von: Zhang, Le, et al.
Veröffentlicht: (2023)
Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes
von: Feng, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Feng, Zhiyuan, et al.
Veröffentlicht: (2025)
Just KIDDIN: Knowledge Infusion and Distillation for Detection of INdecent Memes
von: Garg, Rahul, et al.
Veröffentlicht: (2024)
von: Garg, Rahul, et al.
Veröffentlicht: (2024)
Decompose and Compare Consistency: Measuring VLMs' Answer Reliability via Task-Decomposition Consistency Comparison
von: Yang, Qian, et al.
Veröffentlicht: (2024)
von: Yang, Qian, et al.
Veröffentlicht: (2024)
Investigating Prompting Techniques for Zero- and Few-Shot Visual Question Answering
von: Awal, Rabiul, et al.
Veröffentlicht: (2023)
von: Awal, Rabiul, et al.
Veröffentlicht: (2023)
Random Representations Outperform Online Continually Learned Representations
von: Prabhu, Ameya, et al.
Veröffentlicht: (2024)
von: Prabhu, Ameya, et al.
Veröffentlicht: (2024)
Corrective Machine Unlearning
von: Goel, Shashwat, et al.
Veröffentlicht: (2024)
von: Goel, Shashwat, et al.
Veröffentlicht: (2024)
Improving Automatic VQA Evaluation Using Large Language Models
von: Mañas, Oscar, et al.
Veröffentlicht: (2023)
von: Mañas, Oscar, et al.
Veröffentlicht: (2023)
Towards Infusing Auxiliary Knowledge for Distracted Driver Detection
von: Balappanawar, Ishwar B, et al.
Veröffentlicht: (2024)
von: Balappanawar, Ishwar B, et al.
Veröffentlicht: (2024)
Learning to See the Elephant in the Room: Self-Supervised Context Reasoning in Humans and AI
von: Liu, Xiao, et al.
Veröffentlicht: (2022)
von: Liu, Xiao, et al.
Veröffentlicht: (2022)
An Examination of the Robustness of Reference-Free Image Captioning Evaluation Metrics
von: Ahmadi, Saba, et al.
Veröffentlicht: (2023)
von: Ahmadi, Saba, et al.
Veröffentlicht: (2023)
ParGo: Bridging Vision-Language with Partial and Global Views
von: Wang, An-Lan, et al.
Veröffentlicht: (2024)
von: Wang, An-Lan, et al.
Veröffentlicht: (2024)
Transfer Learning-Based CNN Models for Plant Species Identification Using Leaf Venation Patterns
von: Bharadwaj, Bandita, et al.
Veröffentlicht: (2025)
von: Bharadwaj, Bandita, et al.
Veröffentlicht: (2025)
LABELING COPILOT: A Deep Research Agent for Automated Data Curation in Computer Vision
von: Ganguly, Debargha, et al.
Veröffentlicht: (2025)
von: Ganguly, Debargha, et al.
Veröffentlicht: (2025)
AdaSCALE: Adaptive Scaling for OOD Detection
von: Regmi, Sudarshan
Veröffentlicht: (2025)
von: Regmi, Sudarshan
Veröffentlicht: (2025)
Partial-View Object View Synthesis via Filtered Inversion
von: Sun, Fan-Yun, et al.
Veröffentlicht: (2023)
von: Sun, Fan-Yun, et al.
Veröffentlicht: (2023)
Discovering Failure Modes in Vision-Language Models using RL
von: Jain, Kanishk, et al.
Veröffentlicht: (2026)
von: Jain, Kanishk, et al.
Veröffentlicht: (2026)
Can Machines Imitate Humans? Integrative Turing-like tests for Language and Vision Demonstrate a Narrowing Gap
von: Zhang, Mengmi, et al.
Veröffentlicht: (2022)
von: Zhang, Mengmi, et al.
Veröffentlicht: (2022)
Spatial Calibration of Diffuse LiDARs
von: Behari, Nikhil, et al.
Veröffentlicht: (2026)
von: Behari, Nikhil, et al.
Veröffentlicht: (2026)
Spatial Reasoning with Vision-Language Models in Ego-Centric Multi-View Scenes
von: Gholami, Mohsen, et al.
Veröffentlicht: (2025)
von: Gholami, Mohsen, et al.
Veröffentlicht: (2025)
VisMin: Visual Minimal-Change Understanding
von: Awal, Rabiul, et al.
Veröffentlicht: (2024)
von: Awal, Rabiul, et al.
Veröffentlicht: (2024)
Can Vision-Language Models See Squares? Text-Recognition Mediates Spatial Reasoning Across Three Model Families
von: Levental, Yuval
Veröffentlicht: (2026)
von: Levental, Yuval
Veröffentlicht: (2026)
STUPD: A Synthetic Dataset for Spatial and Temporal Relation Reasoning
von: Agrawal, Palaash, et al.
Veröffentlicht: (2023)
von: Agrawal, Palaash, et al.
Veröffentlicht: (2023)
Learning Multi-View Spatial Reasoning from Cross-View Relations
von: Jeong, Suchae, et al.
Veröffentlicht: (2026)
von: Jeong, Suchae, et al.
Veröffentlicht: (2026)
3ViewSense: Spatial and Mental Perspective Reasoning from Orthographic Views in Vision-Language Models
von: Zhan, Shaoxiong, et al.
Veröffentlicht: (2026)
von: Zhan, Shaoxiong, et al.
Veröffentlicht: (2026)
Image-based Outlier Synthesis With Training Data
von: Regmi, Sudarshan
Veröffentlicht: (2024)
von: Regmi, Sudarshan
Veröffentlicht: (2024)
Fisheye Camera and Ultrasonic Sensor Fusion For Near-Field Obstacle Perception in Bird's-Eye-View
von: Das, Arindam, et al.
Veröffentlicht: (2024)
von: Das, Arindam, et al.
Veröffentlicht: (2024)
Self-Supervised Partial Cycle-Consistency for Multi-View Matching
von: Taggenbrock, Fedor, et al.
Veröffentlicht: (2025)
von: Taggenbrock, Fedor, et al.
Veröffentlicht: (2025)
Evidential Deep Partial Multi-View Classification With Discount Fusion
von: Huang, Haojian, et al.
Veröffentlicht: (2024)
von: Huang, Haojian, et al.
Veröffentlicht: (2024)
Task-Augmented Cross-View Imputation Network for Partial Multi-View Incomplete Multi-Label Classification
von: Zhao, Lian, et al.
Veröffentlicht: (2024)
von: Zhao, Lian, et al.
Veröffentlicht: (2024)
Self-Supervised Spatial Correspondence Across Modalities
von: Shrivastava, Ayush, et al.
Veröffentlicht: (2025)
von: Shrivastava, Ayush, et al.
Veröffentlicht: (2025)
Are VLMs Lost Between Sky and Space? LinkS$^2$Bench for UAV-Satellite Dynamic Cross-View Spatial Intelligence
von: Liu, Dian, et al.
Veröffentlicht: (2026)
von: Liu, Dian, et al.
Veröffentlicht: (2026)
Analyzing the Sensitivity of Vision Language Models in Visual Question Answering
von: Shah, Monika, et al.
Veröffentlicht: (2025)
von: Shah, Monika, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
How and What to Imagine? Visual Thinking in Unified Multimodal Models for Cross-View Spatial Reasoning
von: Yang, Qian, et al.
Veröffentlicht: (2026) -
SPIRIT: Short-term Prediction of solar IRradIance for zero-shot Transfer learning using Foundation Models
von: Mishra, Aditya, et al.
Veröffentlicht: (2025) -
RiT: Vanilla Diffusion Transformers Suffice in Representation Space
von: Zhang, Le, et al.
Veröffentlicht: (2026) -
What if I ask in \textit{alia lingua}? Measuring Functional Similarity Across Languages
von: Mishra, Debangan, et al.
Veröffentlicht: (2025) -
Assessing and Learning Alignment of Unimodal Vision and Language Models
von: Zhang, Le, et al.
Veröffentlicht: (2024)