SAW-Bench: Learning Situated Awareness in the Real World
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Chuhan, Han, Rilyn, Hsu, Joy, Liang, Yongyuan, Dhawan, Rajiv, Wu, Jiajun, Yang, Ming-Hsuan, Wang, Xin Eric |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
From Programs to Poses: Factored Real-World Scene Generation via Learned Program Libraries
von: Hsu, Joy, et al.
Veröffentlicht: (2025)
von: Hsu, Joy, et al.
Veröffentlicht: (2025)
CueBench: Advancing Unified Understanding of Context-Aware Video Anomalies in Real-World
von: Yu, Yating, et al.
Veröffentlicht: (2025)
von: Yu, Yating, et al.
Veröffentlicht: (2025)
Predicate Hierarchies Improve Few-Shot State Classification
von: Jin, Emily, et al.
Veröffentlicht: (2025)
von: Jin, Emily, et al.
Veröffentlicht: (2025)
Dynamic Context-Aware Scene Reasoning Using Vision-Language Alignment in Zero-Shot Real-World Scenarios
von: Rajiv, Manjunath Prasad Holenarasipura, et al.
Veröffentlicht: (2025)
von: Rajiv, Manjunath Prasad Holenarasipura, et al.
Veröffentlicht: (2025)
Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces
von: Yang, Jihan, et al.
Veröffentlicht: (2024)
von: Yang, Jihan, et al.
Veröffentlicht: (2024)
STAR: A Benchmark for Situated Reasoning in Real-World Videos
von: Wu, Bo, et al.
Veröffentlicht: (2024)
von: Wu, Bo, et al.
Veröffentlicht: (2024)
Vision-Language Integration for Zero-Shot Scene Understanding in Real-World Environments
von: Rajiv, Manjunath Prasad Holenarasipura, et al.
Veröffentlicht: (2025)
von: Rajiv, Manjunath Prasad Holenarasipura, et al.
Veröffentlicht: (2025)
Naturally Supervised 3D Visual Grounding with Language-Regularized Concept Learners
von: Feng, Chun, et al.
Veröffentlicht: (2024)
von: Feng, Chun, et al.
Veröffentlicht: (2024)
SOK-Bench: A Situated Video Reasoning Benchmark with Aligned Open-World Knowledge
von: Wang, Andong, et al.
Veröffentlicht: (2024)
von: Wang, Andong, et al.
Veröffentlicht: (2024)
A Tool Bottleneck Framework for Clinically-Informed and Interpretable Medical Image Understanding
von: Liu, Christina, et al.
Veröffentlicht: (2025)
von: Liu, Christina, et al.
Veröffentlicht: (2025)
Real-Time Assessment of Bystander Situation Awareness in Drone-Assisted First Aid
von: Chang, Shen, et al.
Veröffentlicht: (2025)
von: Chang, Shen, et al.
Veröffentlicht: (2025)
Learning Position-Aware Implicit Neural Network for Real-World Face Inpainting
von: Zhao, Bo, et al.
Veröffentlicht: (2024)
von: Zhao, Bo, et al.
Veröffentlicht: (2024)
FaceCam: Portrait Video Camera Control via Scale-Aware Conditioning
von: Lyu, Weijie, et al.
Veröffentlicht: (2026)
von: Lyu, Weijie, et al.
Veröffentlicht: (2026)
What if? Emulative Simulation with World Models for Situated Reasoning
von: Liu, Ruiping, et al.
Veröffentlicht: (2026)
von: Liu, Ruiping, et al.
Veröffentlicht: (2026)
SmokeBench: A Real-World Dataset for Surveillance Image Desmoking in Early-Stage Fire Scenes
von: Jin, Wenzhuo, et al.
Veröffentlicht: (2025)
von: Jin, Wenzhuo, et al.
Veröffentlicht: (2025)
Single Document Image Highlight Removal via A Large-Scale Real-World Dataset and A Location-Aware Network
von: Pan, Lu, et al.
Veröffentlicht: (2025)
von: Pan, Lu, et al.
Veröffentlicht: (2025)
Towards Real-World Document Parsing via Realistic Scene Synthesis and Document-Aware Training
von: Li, Gengluo, et al.
Veröffentlicht: (2026)
von: Li, Gengluo, et al.
Veröffentlicht: (2026)
MEGA-Bench: Scaling Multimodal Evaluation to over 500 Real-World Tasks
von: Chen, Jiacheng, et al.
Veröffentlicht: (2024)
von: Chen, Jiacheng, et al.
Veröffentlicht: (2024)
Distilling Spectral Graph for Object-Context Aware Open-Vocabulary Semantic Segmentation
von: Kim, Chanyoung, et al.
Veröffentlicht: (2024)
von: Kim, Chanyoung, et al.
Veröffentlicht: (2024)
Real-time Ship Recognition and Georeferencing for the Improvement of Maritime Situational Awareness
von: Perez, Borja Carrillo
Veröffentlicht: (2024)
von: Perez, Borja Carrillo
Veröffentlicht: (2024)
SciMDR: Advancing Scientific Multimodal Document Reasoning
von: Chen, Ziyu, et al.
Veröffentlicht: (2026)
von: Chen, Ziyu, et al.
Veröffentlicht: (2026)
SeeSR: Towards Semantics-Aware Real-World Image Super-Resolution
von: Wu, Rongyuan, et al.
Veröffentlicht: (2023)
von: Wu, Rongyuan, et al.
Veröffentlicht: (2023)
KRIS-Bench: Benchmarking Next-Level Intelligent Image Editing Models
von: Wu, Yongliang, et al.
Veröffentlicht: (2025)
von: Wu, Yongliang, et al.
Veröffentlicht: (2025)
Telling Left from Right: Identifying Geometry-Aware Semantic Correspondence
von: Zhang, Junyi, et al.
Veröffentlicht: (2023)
von: Zhang, Junyi, et al.
Veröffentlicht: (2023)
DeeDSR: Towards Real-World Image Super-Resolution via Degradation-Aware Stable Diffusion
von: Bi, Chunyang, et al.
Veröffentlicht: (2024)
von: Bi, Chunyang, et al.
Veröffentlicht: (2024)
Empowering Large Language Models with 3D Situation Awareness
von: Yuan, Zhihao, et al.
Veröffentlicht: (2025)
von: Yuan, Zhihao, et al.
Veröffentlicht: (2025)
Stanford-ORB: A Real-World 3D Object Inverse Rendering Benchmark
von: Kuang, Zhengfei, et al.
Veröffentlicht: (2023)
von: Kuang, Zhengfei, et al.
Veröffentlicht: (2023)
DACESR: Degradation-Aware Conditional Embedding for Real-World Image Super-Resolution
von: Lei, Xiaoyan, et al.
Veröffentlicht: (2026)
von: Lei, Xiaoyan, et al.
Veröffentlicht: (2026)
WeatherBench: A Real-World Benchmark Dataset for All-in-One Adverse Weather Image Restoration
von: Guan, Qiyuan, et al.
Veröffentlicht: (2025)
von: Guan, Qiyuan, et al.
Veröffentlicht: (2025)
DiffST: Spatiotemporal-Aware Diffusion for Real-World Space-Time Video Super-Resolution
von: Chen, Zheng, et al.
Veröffentlicht: (2026)
von: Chen, Zheng, et al.
Veröffentlicht: (2026)
Cascade Prompt Learning for Vision-Language Model Adaptation
von: Wu, Ge, et al.
Veröffentlicht: (2024)
von: Wu, Ge, et al.
Veröffentlicht: (2024)
Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding
von: Wu, Hang, et al.
Veröffentlicht: (2026)
von: Wu, Hang, et al.
Veröffentlicht: (2026)
"PhyWorldBench": A Comprehensive Evaluation of Physical Realism in Text-to-Video Models
von: Gu, Jing, et al.
Veröffentlicht: (2025)
von: Gu, Jing, et al.
Veröffentlicht: (2025)
SAW: Toward a Surgical Action World Model via Controllable and Scalable Video Generation
von: Rapuri, Sampath, et al.
Veröffentlicht: (2026)
von: Rapuri, Sampath, et al.
Veröffentlicht: (2026)
UnsafeBench: Benchmarking Image Safety Classifiers on Real-World and AI-Generated Images
von: Qu, Yiting, et al.
Veröffentlicht: (2024)
von: Qu, Yiting, et al.
Veröffentlicht: (2024)
Degradation-Aware and Structure-Preserving Diffusion for Real-World Image Super-Resolution
von: Ji, Yang, et al.
Veröffentlicht: (2026)
von: Ji, Yang, et al.
Veröffentlicht: (2026)
VA-AR: Learning Velocity-Aware Action Representations with Mixture of Window Attention
von: Wei, Jiangning, et al.
Veröffentlicht: (2025)
von: Wei, Jiangning, et al.
Veröffentlicht: (2025)
Composable Part-Based Manipulation
von: Liu, Weiyu, et al.
Veröffentlicht: (2024)
von: Liu, Weiyu, et al.
Veröffentlicht: (2024)
TEXTS-Diff: TEXTS-Aware Diffusion Model for Real-World Text Image Super-Resolution
von: He, Haodong, et al.
Veröffentlicht: (2026)
von: He, Haodong, et al.
Veröffentlicht: (2026)
Lemon: A Unified and Scalable 3D Multimodal Model for Universal Spatial Understanding
von: Liang, Yongyuan, et al.
Veröffentlicht: (2025)
von: Liang, Yongyuan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
From Programs to Poses: Factored Real-World Scene Generation via Learned Program Libraries
von: Hsu, Joy, et al.
Veröffentlicht: (2025) -
CueBench: Advancing Unified Understanding of Context-Aware Video Anomalies in Real-World
von: Yu, Yating, et al.
Veröffentlicht: (2025) -
Predicate Hierarchies Improve Few-Shot State Classification
von: Jin, Emily, et al.
Veröffentlicht: (2025) -
Dynamic Context-Aware Scene Reasoning Using Vision-Language Alignment in Zero-Shot Real-World Scenarios
von: Rajiv, Manjunath Prasad Holenarasipura, et al.
Veröffentlicht: (2025) -
Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces
von: Yang, Jihan, et al.
Veröffentlicht: (2024)