Crab: A Unified Audio-Visual Scene Understanding Model with Explicit Cooperation
Fuente:
arXiv
Saved in:
| Main Authors: | Du, Henghui, Li, Guangyao, Zhou, Chang, Zhang, Chunjie, Zhao, Alan, Hu, Di |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Crab$^{+}$: A Scalable and Unified Audio-Visual Scene Understanding Model with Explicit Cooperation
by: Cai, Dongnuan, et al.
Published: (2026)
by: Cai, Dongnuan, et al.
Published: (2026)
Boosting Audio Visual Question Answering via Key Semantic-Aware Cues
by: Li, Guangyao, et al.
Published: (2024)
by: Li, Guangyao, et al.
Published: (2024)
AV-Unified: A Unified Framework for Audio-visual Scene Understanding
by: Li, Guangyao, et al.
Published: (2026)
by: Li, Guangyao, et al.
Published: (2026)
Video Detective: Seek Critical Clues Recurrently to Answer Question from Long Videos
by: Du, Henghui, et al.
Published: (2025)
by: Du, Henghui, et al.
Published: (2025)
Ref-AVS: Refer and Segment Objects in Audio-Visual Scenes
by: Wang, Yaoting, et al.
Published: (2024)
by: Wang, Yaoting, et al.
Published: (2024)
APPO: Attention-guided Perception Policy Optimization for Video Reasoning
by: Du, Henghui, et al.
Published: (2026)
by: Du, Henghui, et al.
Published: (2026)
Mettle: Meta-Token Learning for Memory-Efficient Audio-Visual Adaptation
by: Zhou, Jinxing, et al.
Published: (2025)
by: Zhou, Jinxing, et al.
Published: (2025)
Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation
by: Ying, Kaining, et al.
Published: (2025)
by: Ying, Kaining, et al.
Published: (2025)
Understanding Robustness of Visual State Space Models for Image Classification
by: Du, Chengbin, et al.
Published: (2024)
by: Du, Chengbin, et al.
Published: (2024)
Unveiling and Mitigating Bias in Audio Visual Segmentation
by: Sun, Peiwen, et al.
Published: (2024)
by: Sun, Peiwen, et al.
Published: (2024)
Object-aware Sound Source Localization via Audio-Visual Scene Understanding
by: Um, Sung Jin, et al.
Published: (2025)
by: Um, Sung Jin, et al.
Published: (2025)
HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation
by: Zhou, Xin, et al.
Published: (2026)
by: Zhou, Xin, et al.
Published: (2026)
UFVideo: Towards Unified Fine-Grained Video Cooperative Understanding with Large Language Models
by: Pan, Hewen, et al.
Published: (2025)
by: Pan, Hewen, et al.
Published: (2025)
EgoAVU: Egocentric Audio-Visual Understanding
by: Seth, Ashish, et al.
Published: (2026)
by: Seth, Ashish, et al.
Published: (2026)
Explicit Relational Reasoning Network for Scene Text Detection
by: Su, Yuchen, et al.
Published: (2024)
by: Su, Yuchen, et al.
Published: (2024)
HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and Generation
by: Zhou, Xin, et al.
Published: (2025)
by: Zhou, Xin, et al.
Published: (2025)
A Unified Framework for 3D Scene Understanding
by: Xu, Wei, et al.
Published: (2024)
by: Xu, Wei, et al.
Published: (2024)
Skywork UniPic: Unified Autoregressive Modeling for Visual Understanding and Generation
by: Wang, Peiyu, et al.
Published: (2025)
by: Wang, Peiyu, et al.
Published: (2025)
Adaptive Visual Scene Understanding: Incremental Scene Graph Generation
by: Khandelwal, Naitik, et al.
Published: (2023)
by: Khandelwal, Naitik, et al.
Published: (2023)
SceneParser: Hierarchical Scene Parsing for Visual Semantics Understanding
by: Xu, Pengxin, et al.
Published: (2026)
by: Xu, Pengxin, et al.
Published: (2026)
UniScene: Unified Occupancy-centric Driving Scene Generation
by: Li, Bohan, et al.
Published: (2024)
by: Li, Bohan, et al.
Published: (2024)
Benchmarking Multimodal Mathematical Reasoning with Explicit Visual Dependency
by: Wang, Zhikai, et al.
Published: (2025)
by: Wang, Zhikai, et al.
Published: (2025)
Prompting Segmentation with Sound Is Generalizable Audio-Visual Source Localizer
by: Wang, Yaoting, et al.
Published: (2023)
by: Wang, Yaoting, et al.
Published: (2023)
SceneDesigner: Controllable Multi-Object Image Generation with 9-DoF Pose Manipulation
by: Qin, Zhenyuan, et al.
Published: (2025)
by: Qin, Zhenyuan, et al.
Published: (2025)
LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning
by: Dai, Yifan, et al.
Published: (2026)
by: Dai, Yifan, et al.
Published: (2026)
CommonScenes: Generating Commonsense 3D Indoor Scenes with Scene Graph Diffusion
by: Zhai, Guangyao, et al.
Published: (2023)
by: Zhai, Guangyao, et al.
Published: (2023)
InteractiveOmni: A Unified Omni-modal Model for Audio-Visual Multi-turn Dialogue
by: Tong, Wenwen, et al.
Published: (2025)
by: Tong, Wenwen, et al.
Published: (2025)
GeoGaussian: Geometry-aware Gaussian Splatting for Scene Rendering
by: Li, Yanyan, et al.
Published: (2024)
by: Li, Yanyan, et al.
Published: (2024)
TUMTraffic-VideoQA: A Benchmark for Unified Spatio-Temporal Video Understanding in Traffic Scenes
by: Zhou, Xingcheng, et al.
Published: (2025)
by: Zhou, Xingcheng, et al.
Published: (2025)
Unified 3D Scene Understanding Through Physical World Modeling
by: Lee, Wanhee, et al.
Published: (2026)
by: Lee, Wanhee, et al.
Published: (2026)
SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes
by: Wang, Yuji, et al.
Published: (2025)
by: Wang, Yuji, et al.
Published: (2025)
Compositional Scene Understanding through Inverse Generative Modeling
by: Wang, Yanbo, et al.
Published: (2025)
by: Wang, Yanbo, et al.
Published: (2025)
Uni-Sign: Toward Unified Sign Language Understanding at Scale
by: Li, Zecheng, et al.
Published: (2025)
by: Li, Zecheng, et al.
Published: (2025)
MUSE: Multi-Subject Unified Synthesis via Explicit Layout Semantic Expansion
by: Peng, Fei, et al.
Published: (2025)
by: Peng, Fei, et al.
Published: (2025)
Explicit Visual Prompts for Visual Object Tracking
by: Shi, Liangtao, et al.
Published: (2024)
by: Shi, Liangtao, et al.
Published: (2024)
UniModel: A Visual-Only Framework for Unified Multimodal Understanding and Generation
by: Zhang, Chi, et al.
Published: (2025)
by: Zhang, Chi, et al.
Published: (2025)
SAVEn-Vid: Synergistic Audio-Visual Integration for Enhanced Understanding in Long Video Context
by: Li, Jungang, et al.
Published: (2024)
by: Li, Jungang, et al.
Published: (2024)
Steering Visual Generation in Unified Multimodal Models with Understanding Supervision
by: Liu, Zeyu, et al.
Published: (2026)
by: Liu, Zeyu, et al.
Published: (2026)
Orient Anything V2: Unifying Orientation and Rotation Understanding
by: Wang, Zehan, et al.
Published: (2026)
by: Wang, Zehan, et al.
Published: (2026)
MGNiceNet: Unified Monocular Geometric Scene Understanding
by: Schön, Markus, et al.
Published: (2024)
by: Schön, Markus, et al.
Published: (2024)
Similar Items
-
Crab$^{+}$: A Scalable and Unified Audio-Visual Scene Understanding Model with Explicit Cooperation
by: Cai, Dongnuan, et al.
Published: (2026) -
Boosting Audio Visual Question Answering via Key Semantic-Aware Cues
by: Li, Guangyao, et al.
Published: (2024) -
AV-Unified: A Unified Framework for Audio-visual Scene Understanding
by: Li, Guangyao, et al.
Published: (2026) -
Video Detective: Seek Critical Clues Recurrently to Answer Question from Long Videos
by: Du, Henghui, et al.
Published: (2025) -
Ref-AVS: Refer and Segment Objects in Audio-Visual Scenes
by: Wang, Yaoting, et al.
Published: (2024)