More than the Sum: Panorama-Language Models for Adverse Omni-Scenes
Fuente:
arXiv
Guardado en:
| Autores principales: | Fan, Weijia, Liu, Ruiping, Wei, Jiale, Chen, Yufan, Zheng, Junwei, Zeng, Zichao, Zhang, Jiaming, Li, Qiufu, Shen, Linlin, Stiefelhagen, Rainer |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
SGR3 Model: Scene Graph Retrieval-Reasoning Model in 3D
por: Wang, Zirui, et al.
Publicado: (2026)
por: Wang, Zirui, et al.
Publicado: (2026)
Scene-agnostic Pose Regression for Visual Localization
por: Zheng, Junwei, et al.
Publicado: (2025)
por: Zheng, Junwei, et al.
Publicado: (2025)
DriveXQA: Cross-modal Visual Question Answering for Adverse Driving Scene Understanding
por: Tao, Mingzhe, et al.
Publicado: (2026)
por: Tao, Mingzhe, et al.
Publicado: (2026)
OneBEV: Using One Panoramic Image for Bird's-Eye-View Semantic Mapping
por: Wei, Jiale, et al.
Publicado: (2024)
por: Wei, Jiale, et al.
Publicado: (2024)
Faster or Stronger: Towards Flexible Visual Place Recognition via Weighted Aggregation and Token Pruning
por: Zeng, Zichao, et al.
Publicado: (2026)
por: Zeng, Zichao, et al.
Publicado: (2026)
Comb, Prune, Distill: Towards Unified Pruning for Vision Model Compression
por: Schmitt, Jonas, et al.
Publicado: (2024)
por: Schmitt, Jonas, et al.
Publicado: (2024)
RHO: Robust Holistic OSM-Based Metric Cross-View Geo-Localization
por: Zheng, Junwei, et al.
Publicado: (2026)
por: Zheng, Junwei, et al.
Publicado: (2026)
RoDLA: Benchmarking the Robustness of Document Layout Analysis Models
por: Chen, Yufan, et al.
Publicado: (2024)
por: Chen, Yufan, et al.
Publicado: (2024)
HybriDLA: Hybrid Generation for Document Layout Analysis
por: Chen, Yufan, et al.
Publicado: (2025)
por: Chen, Yufan, et al.
Publicado: (2025)
Graph-based Document Structure Analysis
por: Chen, Yufan, et al.
Publicado: (2025)
por: Chen, Yufan, et al.
Publicado: (2025)
Deformable Mamba for Wide Field of View Segmentation
por: Hu, Jie, et al.
Publicado: (2024)
por: Hu, Jie, et al.
Publicado: (2024)
Fourier Prompt Tuning for Modality-Incomplete Scene Segmentation
por: Liu, Ruiping, et al.
Publicado: (2024)
por: Liu, Ruiping, et al.
Publicado: (2024)
Open Panoramic Segmentation
por: Zheng, Junwei, et al.
Publicado: (2024)
por: Zheng, Junwei, et al.
Publicado: (2024)
Situat3DChange: Situated 3D Change Understanding Dataset for Multimodal Large Language Model
por: Liu, Ruiping, et al.
Publicado: (2025)
por: Liu, Ruiping, et al.
Publicado: (2025)
@Bench: Benchmarking Vision-Language Models for Human-centered Assistive Technology
por: Jiang, Xin, et al.
Publicado: (2024)
por: Jiang, Xin, et al.
Publicado: (2024)
Towards Multi-Source Domain Generalization for Sleep Staging with Noisy Labels
por: Wang, Kening, et al.
Publicado: (2026)
por: Wang, Kening, et al.
Publicado: (2026)
MICA: Multi-Agent Industrial Coordination Assistant
por: Wen, Di, et al.
Publicado: (2025)
por: Wen, Di, et al.
Publicado: (2025)
Elevating Skeleton-Based Action Recognition with Efficient Multi-Modality Self-Supervision
por: Wei, Yiping, et al.
Publicado: (2023)
por: Wei, Yiping, et al.
Publicado: (2023)
EPL: Empirical Prototype Learning for Deep Face Recognition
por: Fan, Weijia, et al.
Publicado: (2024)
por: Fan, Weijia, et al.
Publicado: (2024)
EgoExoMem: Cross-View Memory Reasoning over Synchronized Egocentric and Exocentric Videos
por: Liu, Ruiping, et al.
Publicado: (2026)
por: Liu, Ruiping, et al.
Publicado: (2026)
CHAOS: Chart Analysis with Outlier Samples
por: Moured, Omar, et al.
Publicado: (2025)
por: Moured, Omar, et al.
Publicado: (2025)
Skeleton-Based Human Action Recognition with Noisy Labels
por: Xu, Yi, et al.
Publicado: (2024)
por: Xu, Yi, et al.
Publicado: (2024)
DAP-MAE: Domain-Adaptive Point Cloud Masked Autoencoder for Effective Cross-Domain Learning
por: Gao, Ziqi, et al.
Publicado: (2025)
por: Gao, Ziqi, et al.
Publicado: (2025)
Exploring Self-supervised Skeleton-based Action Recognition in Occluded Environments
por: Chen, Yifei, et al.
Publicado: (2023)
por: Chen, Yifei, et al.
Publicado: (2023)
MateRobot: Material Recognition in Wearable Robotics for People with Visual Impairments
por: Zheng, Junwei, et al.
Publicado: (2023)
por: Zheng, Junwei, et al.
Publicado: (2023)
Exploring Video-Based Driver Activity Recognition under Noisy Labels
por: Fan, Linjuan, et al.
Publicado: (2025)
por: Fan, Linjuan, et al.
Publicado: (2025)
What if? Emulative Simulation with World Models for Situated Reasoning
por: Liu, Ruiping, et al.
Publicado: (2026)
por: Liu, Ruiping, et al.
Publicado: (2026)
Rethinking Video Human-Object Interaction: Set Prediction over Time for Unified Detection and Anticipation
por: Luo, Yuanhao, et al.
Publicado: (2026)
por: Luo, Yuanhao, et al.
Publicado: (2026)
SFDLA: Source-Free Document Layout Analysis
por: Tewes, Sebastian, et al.
Publicado: (2025)
por: Tewes, Sebastian, et al.
Publicado: (2025)
RefChartQA: Grounding Visual Answer on Chart Images through Instruction Tuning
por: Vogel, Alexander, et al.
Publicado: (2025)
por: Vogel, Alexander, et al.
Publicado: (2025)
BCE3S: Binary Cross-Entropy Based Tripartite Synergistic Learning for Long-tailed Recognition
por: Fan, Weijia, et al.
Publicado: (2025)
por: Fan, Weijia, et al.
Publicado: (2025)
ObjectFinder: An Open-Vocabulary Assistive System for Interactive Object Search by Blind People
por: Liu, Ruiping, et al.
Publicado: (2024)
por: Liu, Ruiping, et al.
Publicado: (2024)
Referring Atomic Video Action Recognition
por: Peng, Kunyu, et al.
Publicado: (2024)
por: Peng, Kunyu, et al.
Publicado: (2024)
IMPACT-CYCLE: A Contract-Based Multi-Agent System for Claim-Level Supervisory Correction of Long-Video Semantic Memory
por: Kong, Weitong, et al.
Publicado: (2026)
por: Kong, Weitong, et al.
Publicado: (2026)
OAFuser: Towards Omni-Aperture Fusion for Light Field Semantic Segmentation
por: Teng, Fei, et al.
Publicado: (2023)
por: Teng, Fei, et al.
Publicado: (2023)
Distilled Transformers with Locally Enhanced Global Representations for Face Forgery Detection
por: Zhang, Yaning, et al.
Publicado: (2024)
por: Zhang, Yaning, et al.
Publicado: (2024)
IMPACT-Scribe: Interactive Temporal Action Segmentation with Boundary Scribbles and Query Planning
por: Yin, Qian, et al.
Publicado: (2026)
por: Yin, Qian, et al.
Publicado: (2026)
Exploring Single Domain Generalization of LiDAR-based Semantic Segmentation under Imperfect Labels
por: Kong, Weitong, et al.
Publicado: (2025)
por: Kong, Weitong, et al.
Publicado: (2025)
Go Beyond Earth: Understanding Human Actions and Scenes in Microgravity Environments
por: Wen, Di, et al.
Publicado: (2025)
por: Wen, Di, et al.
Publicado: (2025)
PSGS: Text-driven Panorama Sliding Scene Generation via Gaussian Splatting
por: Zhang, Xin, et al.
Publicado: (2026)
por: Zhang, Xin, et al.
Publicado: (2026)
Ejemplares similares
-
SGR3 Model: Scene Graph Retrieval-Reasoning Model in 3D
por: Wang, Zirui, et al.
Publicado: (2026) -
Scene-agnostic Pose Regression for Visual Localization
por: Zheng, Junwei, et al.
Publicado: (2025) -
DriveXQA: Cross-modal Visual Question Answering for Adverse Driving Scene Understanding
por: Tao, Mingzhe, et al.
Publicado: (2026) -
OneBEV: Using One Panoramic Image for Bird's-Eye-View Semantic Mapping
por: Wei, Jiale, et al.
Publicado: (2024) -
Faster or Stronger: Towards Flexible Visual Place Recognition via Weighted Aggregation and Token Pruning
por: Zeng, Zichao, et al.
Publicado: (2026)