Understanding Museum Exhibits using Vision-Language Reasoning
Fuente:
arXiv
Guardado en:
| Autores principales: | Balauca, Ada-Astrid, Garai, Sanjana, Balauca, Stefan, Shetty, Rasesh Udayakumar, Agrawal, Naitik, Shah, Dhwanil Subhashbhai, Fu, Yuqian, Wang, Xi, Toutanova, Kristina, Paudel, Danda Pani, Van Gool, Luc |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Taming CLIP for Fine-grained and Structured Visual Understanding of Museum Exhibits
por: Balauca, Ada-Astrid, et al.
Publicado: (2024)
por: Balauca, Ada-Astrid, et al.
Publicado: (2024)
Practical Hybrid Quantum Language Models with Observable Readout on Real Hardware
por: Balauca, Stefan, et al.
Publicado: (2025)
por: Balauca, Stefan, et al.
Publicado: (2025)
Vision encoders should be image size agnostic and task driven
por: Prisadnikov, Nedyalko, et al.
Publicado: (2025)
por: Prisadnikov, Nedyalko, et al.
Publicado: (2025)
Self-supervised pretraining for an iterative image size agnostic vision transformer
por: Prisadnikov, Nedyalko, et al.
Publicado: (2026)
por: Prisadnikov, Nedyalko, et al.
Publicado: (2026)
Lego: Learning to Disentangle and Invert Personalized Concepts Beyond Object Appearance in Text-to-Image Diffusion Models
por: Motamed, Saman, et al.
Publicado: (2023)
por: Motamed, Saman, et al.
Publicado: (2023)
EvenNICER-SLAM: Event-based Neural Implicit Encoding SLAM
por: Chen, Shi, et al.
Publicado: (2024)
por: Chen, Shi, et al.
Publicado: (2024)
Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025
por: Fu, Yuqian, et al.
Publicado: (2025)
por: Fu, Yuqian, et al.
Publicado: (2025)
Implicit-Zoo: A Large-Scale Dataset of Neural Implicit Functions for 2D Images and 3D Scenes
por: Ma, Qi, et al.
Publicado: (2024)
por: Ma, Qi, et al.
Publicado: (2024)
Continuous Pose for Monocular Cameras in Neural Implicit Representation
por: Ma, Qi, et al.
Publicado: (2023)
por: Ma, Qi, et al.
Publicado: (2023)
From Synchrony to Sequence: Exo-to-Ego Generation via Interpolation
por: Mahdi, Mohammad, et al.
Publicado: (2026)
por: Mahdi, Mohammad, et al.
Publicado: (2026)
Exo2EgoSyn: Unlocking Foundation Video Generation Models for Exocentric-to-Egocentric Video Synthesis
por: Mahdi, Mohammad, et al.
Publicado: (2025)
por: Mahdi, Mohammad, et al.
Publicado: (2025)
Inferring Compositional 4D Scenes without Ever Seeing One
por: Gokmen, Ahmet Berke, et al.
Publicado: (2025)
por: Gokmen, Ahmet Berke, et al.
Publicado: (2025)
A Simple and Generalist Approach for Panoptic Segmentation
por: Prisadnikov, Nedyalko, et al.
Publicado: (2024)
por: Prisadnikov, Nedyalko, et al.
Publicado: (2024)
RICO: Two Realistic Benchmarks and an In-Depth Analysis for Incremental Learning in Object Detection
por: Neuwirth-Trapp, Matthias, et al.
Publicado: (2025)
por: Neuwirth-Trapp, Matthias, et al.
Publicado: (2025)
Incremental Object Detection with Prompt-based Methods
por: Neuwirth-Trapp, Matthias, et al.
Publicado: (2025)
por: Neuwirth-Trapp, Matthias, et al.
Publicado: (2025)
Occam's LGS: An Efficient Approach for Language Gaussian Splatting
por: Cheng, Jiahuan, et al.
Publicado: (2024)
por: Cheng, Jiahuan, et al.
Publicado: (2024)
Exploration-Driven Generative Interactive Environments
por: Savov, Nedko, et al.
Publicado: (2025)
por: Savov, Nedko, et al.
Publicado: (2025)
StateSpaceDiffuser: Bringing Long Context to Diffusion World Models
por: Savov, Nedko, et al.
Publicado: (2025)
por: Savov, Nedko, et al.
Publicado: (2025)
SeasonScapes: Learning Large-scale Re-lightable 3D Landscapes with Seasonal Variation from Sparse Webcams
por: Kleger, Timo, et al.
Publicado: (2026)
por: Kleger, Timo, et al.
Publicado: (2026)
Ternary-Type Opacity and Hybrid Odometry for RGB NeRF-SLAM
por: Lin, Junru, et al.
Publicado: (2023)
por: Lin, Junru, et al.
Publicado: (2023)
GaussianVLM: Scene-centric 3D Vision-Language Models using Language-aligned Gaussian Splats for Embodied Reasoning and Beyond
por: Halacheva, Anna-Maria, et al.
Publicado: (2025)
por: Halacheva, Anna-Maria, et al.
Publicado: (2025)
LangHOPS: Language Grounded Hierarchical Open-Vocabulary Part Segmentation
por: Miao, Yang, et al.
Publicado: (2025)
por: Miao, Yang, et al.
Publicado: (2025)
ReVLA: Reverting Visual Domain Limitation of Robotic Foundation Models
por: Dey, Sombit, et al.
Publicado: (2024)
por: Dey, Sombit, et al.
Publicado: (2024)
Locate Anything on Earth: Advancing Open-Vocabulary Object Detection for Remote Sensing Community
por: Pan, Jiancheng, et al.
Publicado: (2024)
por: Pan, Jiancheng, et al.
Publicado: (2024)
Accelerating Vision Foundation Models with Drop-in Depthwise Convolution
por: Scribano, Carmelo, et al.
Publicado: (2026)
por: Scribano, Carmelo, et al.
Publicado: (2026)
FireScope: Wildfire Risk Raster Prediction with a Chain-of-Thought Oracle
por: Markov, Mario, et al.
Publicado: (2025)
por: Markov, Mario, et al.
Publicado: (2025)
B-GRTO: Bootstrapped Group Relative Tool Optimization for Referring Segmentation
por: Markov, Mario, et al.
Publicado: (2026)
por: Markov, Mario, et al.
Publicado: (2026)
EgoSpot:Egocentric Multimodal Control for Hands-Free Mobile Manipulation
por: Zhang, Ganlin, et al.
Publicado: (2023)
por: Zhang, Ganlin, et al.
Publicado: (2023)
Articulate3D: Holistic Understanding of 3D Scenes as Universal Scene Description
por: Halacheva, Anna-Maria, et al.
Publicado: (2024)
por: Halacheva, Anna-Maria, et al.
Publicado: (2024)
Partial CLIP is Enough: Chimera-Seg for Zero-shot Semantic Segmentation
por: Chen, Jialei, et al.
Publicado: (2025)
por: Chen, Jialei, et al.
Publicado: (2025)
Rethinking Global Context in Crowd Counting
por: Sun, Guolei, et al.
Publicado: (2021)
por: Sun, Guolei, et al.
Publicado: (2021)
Reducing Unimodal Bias in Multi-Modal Semantic Segmentation with Multi-Scale Functional Entropy Regularization
por: Zheng, Xu, et al.
Publicado: (2025)
por: Zheng, Xu, et al.
Publicado: (2025)
BiXFormer: A Robust Framework for Maximizing Modality Effectiveness in Multi-Modal Semantic Segmentation
por: Chen, Jialei, et al.
Publicado: (2025)
por: Chen, Jialei, et al.
Publicado: (2025)
From Scan to Action: Leveraging Realistic Scans for Embodied Scene Understanding
por: Halacheva, Anna-Maria, et al.
Publicado: (2025)
por: Halacheva, Anna-Maria, et al.
Publicado: (2025)
Learning Generative Interactive Environments By Trained Agent Exploration
por: Kazemi, Naser, et al.
Publicado: (2024)
por: Kazemi, Naser, et al.
Publicado: (2024)
Autonomous Vehicle Controllers From End-to-End Differentiable Simulation
por: Nachkov, Asen, et al.
Publicado: (2024)
por: Nachkov, Asen, et al.
Publicado: (2024)
EgoCross: Benchmarking Multimodal Large Language Models for Cross-Domain Egocentric Video Question Answering
por: Li, Yanjun, et al.
Publicado: (2025)
por: Li, Yanjun, et al.
Publicado: (2025)
OSMDA: OpenStreetMap-based Domain Adaptation for Remote Sensing VLMs
por: Ailuro, Stefan Maria, et al.
Publicado: (2026)
por: Ailuro, Stefan Maria, et al.
Publicado: (2026)
ObjectRelator: Enabling Cross-View Object Relation Understanding Across Ego-Centric and Exo-Centric Perspectives
por: Fu, Yuqian, et al.
Publicado: (2024)
por: Fu, Yuqian, et al.
Publicado: (2024)
Efficient Degradation-agnostic Image Restoration via Channel-Wise Functional Decomposition and Manifold Regularization
por: Ren, Bin, et al.
Publicado: (2025)
por: Ren, Bin, et al.
Publicado: (2025)
Ejemplares similares
-
Taming CLIP for Fine-grained and Structured Visual Understanding of Museum Exhibits
por: Balauca, Ada-Astrid, et al.
Publicado: (2024) -
Practical Hybrid Quantum Language Models with Observable Readout on Real Hardware
por: Balauca, Stefan, et al.
Publicado: (2025) -
Vision encoders should be image size agnostic and task driven
por: Prisadnikov, Nedyalko, et al.
Publicado: (2025) -
Self-supervised pretraining for an iterative image size agnostic vision transformer
por: Prisadnikov, Nedyalko, et al.
Publicado: (2026) -
Lego: Learning to Disentangle and Invert Personalized Concepts Beyond Object Appearance in Text-to-Image Diffusion Models
por: Motamed, Saman, et al.
Publicado: (2023)