Seeing the Unseen: Visual Common Sense for Semantic Placement
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ramrakhya, Ram, Kembhavi, Aniruddha, Batra, Dhruv, Kira, Zsolt, Zeng, Kuo-Hao, Weihs, Luca |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PoliFormer: Scaling On-Policy RL with Transformers Results in Masterful Navigators
von: Zeng, Kuo-Hao, et al.
Veröffentlicht: (2024)
von: Zeng, Kuo-Hao, et al.
Veröffentlicht: (2024)
Iterated Learning Improves Compositionality in Large Vision-Language Models
von: Zheng, Chenhao, et al.
Veröffentlicht: (2024)
von: Zheng, Chenhao, et al.
Veröffentlicht: (2024)
The One RING: a Robotic Indoor Navigation Generalist
von: Eftekhar, Ainaz, et al.
Veröffentlicht: (2024)
von: Eftekhar, Ainaz, et al.
Veröffentlicht: (2024)
Selective Visual Representations Improve Convergence and Generalization for Embodied AI
von: Eftekhar, Ainaz, et al.
Veröffentlicht: (2023)
von: Eftekhar, Ainaz, et al.
Veröffentlicht: (2023)
SAGE: Sink-Aware Grounded Decoding for Multimodal Hallucination Mitigation
von: Shukla, Tripti, et al.
Veröffentlicht: (2026)
von: Shukla, Tripti, et al.
Veröffentlicht: (2026)
Scaling Text-Rich Image Understanding via Code-Guided Synthetic Multimodal Data Generation
von: Yang, Yue, et al.
Veröffentlicht: (2025)
von: Yang, Yue, et al.
Veröffentlicht: (2025)
Unseen Visual Anomaly Generation
von: Sun, Han, et al.
Veröffentlicht: (2024)
von: Sun, Han, et al.
Veröffentlicht: (2024)
Generate Any Scene: Scene Graph Driven Data Synthesis for Visual Generation Training
von: Gao, Ziqi, et al.
Veröffentlicht: (2024)
von: Gao, Ziqi, et al.
Veröffentlicht: (2024)
Seeing the Unseen in Low-light Spike Streams
von: Hu, Liwen, et al.
Veröffentlicht: (2025)
von: Hu, Liwen, et al.
Veröffentlicht: (2025)
Imagining the Unseen: Generative Location Modeling for Object Placement
von: Yun, Jooyeol, et al.
Veröffentlicht: (2024)
von: Yun, Jooyeol, et al.
Veröffentlicht: (2024)
SPOC: Imitating Shortest Paths in Simulation Enables Effective Navigation and Manipulation in the Real World
von: Ehsani, Kiana, et al.
Veröffentlicht: (2023)
von: Ehsani, Kiana, et al.
Veröffentlicht: (2023)
FLaRe: Achieving Masterful and Adaptive Robot Policies with Large-Scale Reinforcement Learning Fine-Tuning
von: Hu, Jiaheng, et al.
Veröffentlicht: (2024)
von: Hu, Jiaheng, et al.
Veröffentlicht: (2024)
SeeU: Seeing the Unseen World via 4D Dynamics-aware Generation
von: Yuan, Yu, et al.
Veröffentlicht: (2025)
von: Yuan, Yu, et al.
Veröffentlicht: (2025)
Pre-trained Text-to-Image Diffusion Models Are Versatile Representation Learners for Control
von: Gupta, Gunshi, et al.
Veröffentlicht: (2024)
von: Gupta, Gunshi, et al.
Veröffentlicht: (2024)
FirePlace: Geometric Refinements of LLM Common Sense Reasoning for 3D Object Placement
von: Huang, Ian, et al.
Veröffentlicht: (2025)
von: Huang, Ian, et al.
Veröffentlicht: (2025)
Contextual Self-paced Learning for Weakly Supervised Spatio-Temporal Video Grounding
von: Kumar, Akash, et al.
Veröffentlicht: (2025)
von: Kumar, Akash, et al.
Veröffentlicht: (2025)
Seeing the Unseen: Zooming in the Dark with Event Cameras
von: Kai, Dachun, et al.
Veröffentlicht: (2026)
von: Kai, Dachun, et al.
Veröffentlicht: (2026)
Seeing the Unseen: A Frequency Prompt Guided Transformer for Image Restoration
von: Zhou, Shihao, et al.
Veröffentlicht: (2024)
von: Zhou, Shihao, et al.
Veröffentlicht: (2024)
MARCO: Navigating the Unseen Space of Semantic Correspondence
von: Cuttano, Claudia, et al.
Veröffentlicht: (2026)
von: Cuttano, Claudia, et al.
Veröffentlicht: (2026)
Exposing and Addressing Cross-Task Inconsistency in Unified Vision-Language Models
von: Maharana, Adyasha, et al.
Veröffentlicht: (2023)
von: Maharana, Adyasha, et al.
Veröffentlicht: (2023)
Rethinking Weight Decay for Robust Fine-Tuning of Foundation Models
von: Tian, Junjiao, et al.
Veröffentlicht: (2024)
von: Tian, Junjiao, et al.
Veröffentlicht: (2024)
Uncertainty-based Detection of Adversarial Attacks in Semantic Segmentation
von: Maag, Kira, et al.
Veröffentlicht: (2023)
von: Maag, Kira, et al.
Veröffentlicht: (2023)
FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering
von: Huang, Chengyue, et al.
Veröffentlicht: (2025)
von: Huang, Chengyue, et al.
Veröffentlicht: (2025)
The Geometry of Robustness: Optimizing Loss Landscape Curvature and Feature Manifold Alignment for Robust Finetuning of Vision-Language Models
von: Chopra, Shivang, et al.
Veröffentlicht: (2026)
von: Chopra, Shivang, et al.
Veröffentlicht: (2026)
Preserving Identity with Variational Score for General-purpose 3D Editing
von: Le, Duong H., et al.
Veröffentlicht: (2024)
von: Le, Duong H., et al.
Veröffentlicht: (2024)
Uncertainty-weighted Loss Functions for Improved Adversarial Attacks on Semantic Segmentation
von: Maag, Kira, et al.
Veröffentlicht: (2023)
von: Maag, Kira, et al.
Veröffentlicht: (2023)
CodeNav: Beyond tool-use to using real-world codebases with LLM agents
von: Gupta, Tanmay, et al.
Veröffentlicht: (2024)
von: Gupta, Tanmay, et al.
Veröffentlicht: (2024)
False Negative Reduction in Semantic Segmentation under Domain Shift using Depth Estimation
von: Maag, Kira, et al.
Veröffentlicht: (2022)
von: Maag, Kira, et al.
Veröffentlicht: (2022)
EmbodiedSplat: Personalized Real-to-Sim-to-Real Navigation with Gaussian Splats from a Mobile Device
von: Chhablani, Gunjan, et al.
Veröffentlicht: (2025)
von: Chhablani, Gunjan, et al.
Veröffentlicht: (2025)
Generative Video Diffusion for Unseen Novel Semantic Video Moment Retrieval
von: Luo, Dezhao, et al.
Veröffentlicht: (2024)
von: Luo, Dezhao, et al.
Veröffentlicht: (2024)
Prompt the Unseen: Evaluating Visual-Language Alignment Beyond Supervision
von: Jung, Raehyuk, et al.
Veröffentlicht: (2025)
von: Jung, Raehyuk, et al.
Veröffentlicht: (2025)
Can Graphs Help Vision SSMs See Better?
von: Parikh, Dhruv, et al.
Veröffentlicht: (2026)
von: Parikh, Dhruv, et al.
Veröffentlicht: (2026)
Diffuse, Attend, and Segment: Unsupervised Zero-Shot Segmentation using Stable Diffusion
von: Tian, Junjiao, et al.
Veröffentlicht: (2023)
von: Tian, Junjiao, et al.
Veröffentlicht: (2023)
Seeing the Unseen: Towards Zero-Shot Inspection for Wind Turbine Blades using Knowledge-Augmented Vision Language Models
von: Zhang, Yang, et al.
Veröffentlicht: (2025)
von: Zhang, Yang, et al.
Veröffentlicht: (2025)
Seeing the Unseen: Mask-Driven Positional Encoding and Strip-Convolution Context Modeling for Cross-View Object Geo-Localization
von: Hu, Shuhan, et al.
Veröffentlicht: (2025)
von: Hu, Shuhan, et al.
Veröffentlicht: (2025)
FindingDory: A Benchmark to Evaluate Memory in Embodied Agents
von: Yadav, Karmesh, et al.
Veröffentlicht: (2025)
von: Yadav, Karmesh, et al.
Veröffentlicht: (2025)
When Domain Generalization meets Generalized Category Discovery: An Adaptive Task-Arithmetic Driven Approach
von: Rathore, Vaibhav, et al.
Veröffentlicht: (2025)
von: Rathore, Vaibhav, et al.
Veröffentlicht: (2025)
EscherNet++: Simultaneous Amodal Completion and Scalable View Synthesis through Masked Fine-Tuning and Enhanced Feed-Forward 3D Reconstruction
von: Zhang, Xinan, et al.
Veröffentlicht: (2025)
von: Zhang, Xinan, et al.
Veröffentlicht: (2025)
Holodeck: Language Guided Generation of 3D Embodied AI Environments
von: Yang, Yue, et al.
Veröffentlicht: (2023)
von: Yang, Yue, et al.
Veröffentlicht: (2023)
SAT: Dynamic Spatial Aptitude Training for Multimodal Language Models
von: Ray, Arijit, et al.
Veröffentlicht: (2024)
von: Ray, Arijit, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
PoliFormer: Scaling On-Policy RL with Transformers Results in Masterful Navigators
von: Zeng, Kuo-Hao, et al.
Veröffentlicht: (2024) -
Iterated Learning Improves Compositionality in Large Vision-Language Models
von: Zheng, Chenhao, et al.
Veröffentlicht: (2024) -
The One RING: a Robotic Indoor Navigation Generalist
von: Eftekhar, Ainaz, et al.
Veröffentlicht: (2024) -
Selective Visual Representations Improve Convergence and Generalization for Embodied AI
von: Eftekhar, Ainaz, et al.
Veröffentlicht: (2023) -
SAGE: Sink-Aware Grounded Decoding for Multimodal Hallucination Mitigation
von: Shukla, Tripti, et al.
Veröffentlicht: (2026)