Enregistré dans:
| Auteurs principaux: | Tang, Xuejiao, Zhang, Wenbin |
|---|---|
| Format: | Preprint |
| Publié: |
2022
|
| Sujets: | |
| Accès en ligne: | https://arxiv.org/abs/2204.08027 |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Curvature-Aware Captioning:Leveraging Geodesic Attention for 3D Scene Understanding
par: He, Ziyao, et autres
Publié: (2026)
par: He, Ziyao, et autres
Publié: (2026)
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding
par: Zeng, Nianbo, et autres
Publié: (2025)
par: Zeng, Nianbo, et autres
Publié: (2025)
Lifting Unlabeled Internet-level Data for 3D Scene Understanding
par: Chen, Yixin, et autres
Publié: (2026)
par: Chen, Yixin, et autres
Publié: (2026)
HAMF: A Hybrid Attention-Mamba Framework for Joint Scene Context Understanding and Future Motion Representation Learning
par: Mei, Xiaodong, et autres
Publié: (2025)
par: Mei, Xiaodong, et autres
Publié: (2025)
TopoLogic: An Interpretable Pipeline for Lane Topology Reasoning on Driving Scenes
par: Fu, Yanping, et autres
Publié: (2024)
par: Fu, Yanping, et autres
Publié: (2024)
Towards Holistic Surgical Scene Understanding
par: Valderrama, Natalia, et autres
Publié: (2022)
par: Valderrama, Natalia, et autres
Publié: (2022)
Fine-Grained Scene Graph Generation via Sample-Level Bias Prediction
par: Li, Yansheng, et autres
Publié: (2024)
par: Li, Yansheng, et autres
Publié: (2024)
Heat Diffusion Models -- Interpixel Attention Mechanism
par: Zhang, Pengfei, et autres
Publié: (2025)
par: Zhang, Pengfei, et autres
Publié: (2025)
RieMind: Geometry-Grounded Spatial Agent for Scene Understanding
par: Ropero, Fernando, et autres
Publié: (2026)
par: Ropero, Fernando, et autres
Publié: (2026)
Toward Robust Multimodal Learning using Multimodal Foundational Models
par: Zhao, Xianbing, et autres
Publié: (2024)
par: Zhao, Xianbing, et autres
Publié: (2024)
3D-RFT: Reinforcement Fine-Tuning for Video-based 3D Scene Understanding
par: Linghu, Xiongkun, et autres
Publié: (2026)
par: Linghu, Xiongkun, et autres
Publié: (2026)
Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding
par: Ma, Ke, et autres
Publié: (2026)
par: Ma, Ke, et autres
Publié: (2026)
Text-Scene: A Scene-to-Language Parsing Framework for 3D Scene Understanding
par: Li, Haoyuan, et autres
Publié: (2025)
par: Li, Haoyuan, et autres
Publié: (2025)
Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind
par: Li, Qingmei, et autres
Publié: (2025)
par: Li, Qingmei, et autres
Publié: (2025)
Evaluating Compositional Scene Understanding in Multimodal Generative Models
par: Fu, Shuhao, et autres
Publié: (2025)
par: Fu, Shuhao, et autres
Publié: (2025)
Solving Scene Understanding for Autonomous Navigation in Unstructured Environments
par: Renji, Naveen Mathews, et autres
Publié: (2025)
par: Renji, Naveen Mathews, et autres
Publié: (2025)
DenseScan: Advancing 3D Scene Understanding with 2D Dense Annotation
par: Wang, Zirui, et autres
Publié: (2025)
par: Wang, Zirui, et autres
Publié: (2025)
MASR: Self-Reflective Reasoning through Multimodal Hierarchical Attention Focusing for Agent-based Video Understanding
par: Cao, Shiwen, et autres
Publié: (2025)
par: Cao, Shiwen, et autres
Publié: (2025)
MTMamba: Enhancing Multi-Task Dense Scene Understanding by Mamba-Based Decoders
par: Lin, Baijiong, et autres
Publié: (2024)
par: Lin, Baijiong, et autres
Publié: (2024)
Trustworthy Automated Driving through Qualitative Scene Understanding and Explanations
par: Belmecheri, Nassim, et autres
Publié: (2024)
par: Belmecheri, Nassim, et autres
Publié: (2024)
HexPlane Representation for 3D Semantic Scene Understanding
par: Chen, Zeren, et autres
Publié: (2025)
par: Chen, Zeren, et autres
Publié: (2025)
PerspectiveNet: Multi-View Perception for Dynamic Scene Understanding
par: Nguyen, Vinh
Publié: (2024)
par: Nguyen, Vinh
Publié: (2024)
DOCTR: Disentangled Object-Centric Transformer for Point Scene Understanding
par: Yu, Xiaoxuan, et autres
Publié: (2024)
par: Yu, Xiaoxuan, et autres
Publié: (2024)
Masked Scene Modeling: Narrowing the Gap Between Supervised and Self-Supervised Learning in 3D Scene Understanding
par: Hermosilla, Pedro, et autres
Publié: (2025)
par: Hermosilla, Pedro, et autres
Publié: (2025)
Seeing is Understanding: Unlocking Causal Attention into Modality-Mutual Attention for Multimodal LLMs
par: Wang, Wei-Yao, et autres
Publié: (2025)
par: Wang, Wei-Yao, et autres
Publié: (2025)
Learning to Look: Cognitive Attention Alignment with Vision-Language Models
par: Yang, Ryan L., et autres
Publié: (2025)
par: Yang, Ryan L., et autres
Publié: (2025)
An Efficient Aerial Image Detection with Variable Receptive Fields
par: Wenbin, Liu
Publié: (2025)
par: Wenbin, Liu
Publié: (2025)
RE-VLM: Event-Augmented Vision-Language Model for Scene Understanding
par: Liu, Hanqing, et autres
Publié: (2026)
par: Liu, Hanqing, et autres
Publié: (2026)
Cognitive Mismatch in Multimodal Large Language Models for Discrete Symbol Understanding
par: Li, Yinghui, et autres
Publié: (2026)
par: Li, Yinghui, et autres
Publié: (2026)
Argus: Leveraging Multiview Images for Improved 3-D Scene Understanding With Large Language Models
par: Xu, Yifan, et autres
Publié: (2025)
par: Xu, Yifan, et autres
Publié: (2025)
Enhancing Human-Centered Dynamic Scene Understanding via Multiple LLMs Collaborated Reasoning
par: Zhang, Hang, et autres
Publié: (2024)
par: Zhang, Hang, et autres
Publié: (2024)
Towards Unified Surgical Scene Understanding:Bridging Reasoning and Grounding via MLLMs
par: Huang, Jincai, et autres
Publié: (2026)
par: Huang, Jincai, et autres
Publié: (2026)
Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models
par: Mohamud, Safaa Abdullahi Moallim, et autres
Publié: (2025)
par: Mohamud, Safaa Abdullahi Moallim, et autres
Publié: (2025)
Towards Comprehensive Scene Understanding: Integrating First and Third-Person Views for LVLMs
par: Lee, Insu, et autres
Publié: (2025)
par: Lee, Insu, et autres
Publié: (2025)
Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning
par: Fu, Rao, et autres
Publié: (2024)
par: Fu, Rao, et autres
Publié: (2024)
CogME: A Cognition-Inspired Multi-Dimensional Evaluation Metric for Story Understanding
par: Shin, Minjung, et autres
Publié: (2021)
par: Shin, Minjung, et autres
Publié: (2021)
Attention over Scene Graphs: Indoor Scene Representations Toward CSAI Classification
par: Barros, Artur, et autres
Publié: (2025)
par: Barros, Artur, et autres
Publié: (2025)
EventSTR: A Benchmark Dataset and Baselines for Event Stream based Scene Text Recognition
par: Wang, Xiao, et autres
Publié: (2025)
par: Wang, Xiao, et autres
Publié: (2025)
NuScenes-SpatialQA: A Spatial Understanding and Reasoning Benchmark for Vision-Language Models in Autonomous Driving
par: Tian, Kexin, et autres
Publié: (2025)
par: Tian, Kexin, et autres
Publié: (2025)
Attention at Rest Stays at Rest: Breaking Visual Inertia for Cognitive Hallucination Mitigation
par: Gong, Boyang, et autres
Publié: (2026)
par: Gong, Boyang, et autres
Publié: (2026)
Documents similaires
-
Curvature-Aware Captioning:Leveraging Geodesic Attention for 3D Scene Understanding
par: He, Ziyao, et autres
Publié: (2026) -
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding
par: Zeng, Nianbo, et autres
Publié: (2025) -
Lifting Unlabeled Internet-level Data for 3D Scene Understanding
par: Chen, Yixin, et autres
Publié: (2026) -
HAMF: A Hybrid Attention-Mamba Framework for Joint Scene Context Understanding and Future Motion Representation Learning
par: Mei, Xiaodong, et autres
Publié: (2025) -
TopoLogic: An Interpretable Pipeline for Lane Topology Reasoning on Driving Scenes
par: Fu, Yanping, et autres
Publié: (2024)