Gespeichert in:
| Hauptverfasser: | Tang, Xuejiao, Zhang, Wenbin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2022
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2204.08027 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Curvature-Aware Captioning:Leveraging Geodesic Attention for 3D Scene Understanding
von: He, Ziyao, et al.
Veröffentlicht: (2026)
von: He, Ziyao, et al.
Veröffentlicht: (2026)
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding
von: Zeng, Nianbo, et al.
Veröffentlicht: (2025)
von: Zeng, Nianbo, et al.
Veröffentlicht: (2025)
Lifting Unlabeled Internet-level Data for 3D Scene Understanding
von: Chen, Yixin, et al.
Veröffentlicht: (2026)
von: Chen, Yixin, et al.
Veröffentlicht: (2026)
HAMF: A Hybrid Attention-Mamba Framework for Joint Scene Context Understanding and Future Motion Representation Learning
von: Mei, Xiaodong, et al.
Veröffentlicht: (2025)
von: Mei, Xiaodong, et al.
Veröffentlicht: (2025)
TopoLogic: An Interpretable Pipeline for Lane Topology Reasoning on Driving Scenes
von: Fu, Yanping, et al.
Veröffentlicht: (2024)
von: Fu, Yanping, et al.
Veröffentlicht: (2024)
Towards Holistic Surgical Scene Understanding
von: Valderrama, Natalia, et al.
Veröffentlicht: (2022)
von: Valderrama, Natalia, et al.
Veröffentlicht: (2022)
Fine-Grained Scene Graph Generation via Sample-Level Bias Prediction
von: Li, Yansheng, et al.
Veröffentlicht: (2024)
von: Li, Yansheng, et al.
Veröffentlicht: (2024)
Heat Diffusion Models -- Interpixel Attention Mechanism
von: Zhang, Pengfei, et al.
Veröffentlicht: (2025)
von: Zhang, Pengfei, et al.
Veröffentlicht: (2025)
RieMind: Geometry-Grounded Spatial Agent for Scene Understanding
von: Ropero, Fernando, et al.
Veröffentlicht: (2026)
von: Ropero, Fernando, et al.
Veröffentlicht: (2026)
Toward Robust Multimodal Learning using Multimodal Foundational Models
von: Zhao, Xianbing, et al.
Veröffentlicht: (2024)
von: Zhao, Xianbing, et al.
Veröffentlicht: (2024)
3D-RFT: Reinforcement Fine-Tuning for Video-based 3D Scene Understanding
von: Linghu, Xiongkun, et al.
Veröffentlicht: (2026)
von: Linghu, Xiongkun, et al.
Veröffentlicht: (2026)
Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding
von: Ma, Ke, et al.
Veröffentlicht: (2026)
von: Ma, Ke, et al.
Veröffentlicht: (2026)
Text-Scene: A Scene-to-Language Parsing Framework for 3D Scene Understanding
von: Li, Haoyuan, et al.
Veröffentlicht: (2025)
von: Li, Haoyuan, et al.
Veröffentlicht: (2025)
Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind
von: Li, Qingmei, et al.
Veröffentlicht: (2025)
von: Li, Qingmei, et al.
Veröffentlicht: (2025)
Evaluating Compositional Scene Understanding in Multimodal Generative Models
von: Fu, Shuhao, et al.
Veröffentlicht: (2025)
von: Fu, Shuhao, et al.
Veröffentlicht: (2025)
Solving Scene Understanding for Autonomous Navigation in Unstructured Environments
von: Renji, Naveen Mathews, et al.
Veröffentlicht: (2025)
von: Renji, Naveen Mathews, et al.
Veröffentlicht: (2025)
DenseScan: Advancing 3D Scene Understanding with 2D Dense Annotation
von: Wang, Zirui, et al.
Veröffentlicht: (2025)
von: Wang, Zirui, et al.
Veröffentlicht: (2025)
MASR: Self-Reflective Reasoning through Multimodal Hierarchical Attention Focusing for Agent-based Video Understanding
von: Cao, Shiwen, et al.
Veröffentlicht: (2025)
von: Cao, Shiwen, et al.
Veröffentlicht: (2025)
MTMamba: Enhancing Multi-Task Dense Scene Understanding by Mamba-Based Decoders
von: Lin, Baijiong, et al.
Veröffentlicht: (2024)
von: Lin, Baijiong, et al.
Veröffentlicht: (2024)
Trustworthy Automated Driving through Qualitative Scene Understanding and Explanations
von: Belmecheri, Nassim, et al.
Veröffentlicht: (2024)
von: Belmecheri, Nassim, et al.
Veröffentlicht: (2024)
HexPlane Representation for 3D Semantic Scene Understanding
von: Chen, Zeren, et al.
Veröffentlicht: (2025)
von: Chen, Zeren, et al.
Veröffentlicht: (2025)
PerspectiveNet: Multi-View Perception for Dynamic Scene Understanding
von: Nguyen, Vinh
Veröffentlicht: (2024)
von: Nguyen, Vinh
Veröffentlicht: (2024)
DOCTR: Disentangled Object-Centric Transformer for Point Scene Understanding
von: Yu, Xiaoxuan, et al.
Veröffentlicht: (2024)
von: Yu, Xiaoxuan, et al.
Veröffentlicht: (2024)
Masked Scene Modeling: Narrowing the Gap Between Supervised and Self-Supervised Learning in 3D Scene Understanding
von: Hermosilla, Pedro, et al.
Veröffentlicht: (2025)
von: Hermosilla, Pedro, et al.
Veröffentlicht: (2025)
Seeing is Understanding: Unlocking Causal Attention into Modality-Mutual Attention for Multimodal LLMs
von: Wang, Wei-Yao, et al.
Veröffentlicht: (2025)
von: Wang, Wei-Yao, et al.
Veröffentlicht: (2025)
Learning to Look: Cognitive Attention Alignment with Vision-Language Models
von: Yang, Ryan L., et al.
Veröffentlicht: (2025)
von: Yang, Ryan L., et al.
Veröffentlicht: (2025)
An Efficient Aerial Image Detection with Variable Receptive Fields
von: Wenbin, Liu
Veröffentlicht: (2025)
von: Wenbin, Liu
Veröffentlicht: (2025)
RE-VLM: Event-Augmented Vision-Language Model for Scene Understanding
von: Liu, Hanqing, et al.
Veröffentlicht: (2026)
von: Liu, Hanqing, et al.
Veröffentlicht: (2026)
Cognitive Mismatch in Multimodal Large Language Models for Discrete Symbol Understanding
von: Li, Yinghui, et al.
Veröffentlicht: (2026)
von: Li, Yinghui, et al.
Veröffentlicht: (2026)
Argus: Leveraging Multiview Images for Improved 3-D Scene Understanding With Large Language Models
von: Xu, Yifan, et al.
Veröffentlicht: (2025)
von: Xu, Yifan, et al.
Veröffentlicht: (2025)
Enhancing Human-Centered Dynamic Scene Understanding via Multiple LLMs Collaborated Reasoning
von: Zhang, Hang, et al.
Veröffentlicht: (2024)
von: Zhang, Hang, et al.
Veröffentlicht: (2024)
Towards Unified Surgical Scene Understanding:Bridging Reasoning and Grounding via MLLMs
von: Huang, Jincai, et al.
Veröffentlicht: (2026)
von: Huang, Jincai, et al.
Veröffentlicht: (2026)
Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models
von: Mohamud, Safaa Abdullahi Moallim, et al.
Veröffentlicht: (2025)
von: Mohamud, Safaa Abdullahi Moallim, et al.
Veröffentlicht: (2025)
Towards Comprehensive Scene Understanding: Integrating First and Third-Person Views for LVLMs
von: Lee, Insu, et al.
Veröffentlicht: (2025)
von: Lee, Insu, et al.
Veröffentlicht: (2025)
Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning
von: Fu, Rao, et al.
Veröffentlicht: (2024)
von: Fu, Rao, et al.
Veröffentlicht: (2024)
CogME: A Cognition-Inspired Multi-Dimensional Evaluation Metric for Story Understanding
von: Shin, Minjung, et al.
Veröffentlicht: (2021)
von: Shin, Minjung, et al.
Veröffentlicht: (2021)
Attention over Scene Graphs: Indoor Scene Representations Toward CSAI Classification
von: Barros, Artur, et al.
Veröffentlicht: (2025)
von: Barros, Artur, et al.
Veröffentlicht: (2025)
EventSTR: A Benchmark Dataset and Baselines for Event Stream based Scene Text Recognition
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
NuScenes-SpatialQA: A Spatial Understanding and Reasoning Benchmark for Vision-Language Models in Autonomous Driving
von: Tian, Kexin, et al.
Veröffentlicht: (2025)
von: Tian, Kexin, et al.
Veröffentlicht: (2025)
Attention at Rest Stays at Rest: Breaking Visual Inertia for Cognitive Hallucination Mitigation
von: Gong, Boyang, et al.
Veröffentlicht: (2026)
von: Gong, Boyang, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Curvature-Aware Captioning:Leveraging Geodesic Attention for 3D Scene Understanding
von: He, Ziyao, et al.
Veröffentlicht: (2026) -
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding
von: Zeng, Nianbo, et al.
Veröffentlicht: (2025) -
Lifting Unlabeled Internet-level Data for 3D Scene Understanding
von: Chen, Yixin, et al.
Veröffentlicht: (2026) -
HAMF: A Hybrid Attention-Mamba Framework for Joint Scene Context Understanding and Future Motion Representation Learning
von: Mei, Xiaodong, et al.
Veröffentlicht: (2025) -
TopoLogic: An Interpretable Pipeline for Lane Topology Reasoning on Driving Scenes
von: Fu, Yanping, et al.
Veröffentlicht: (2024)