Beyond a Single Frame: Multi-Frame Spatially Grounded Reasoning Across Volumetric MRI
Fuente:
arXiv
Saved in:
| Main Authors: | Moukheiber, Lama, Yeung, Caleb M., Xue, Haotian, Helbling, Alec, Zhao, Zelin, Chen, Yongxin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ACWM-Phys: Investigating Generalized Physical Interaction in Action-Conditioned Video World Models
by: Xue, Haotian, et al.
Published: (2026)
by: Xue, Haotian, et al.
Published: (2026)
Looking Beyond What You See: An Empirical Analysis on Subgroup Intersectional Fairness for Multi-label Chest X-ray Classification Using Social Determinants of Racial Health Inequities
by: Moukheiber, Dana, et al.
Published: (2024)
by: Moukheiber, Dana, et al.
Published: (2024)
Laplacian Multi-scale Flow Matching for Generative Modeling
by: Zhao, Zelin, et al.
Published: (2026)
by: Zhao, Zelin, et al.
Published: (2026)
Out of Sight, Not Out of Context? Egocentric Spatial Reasoning in VLMs Across Disjoint Frames
by: Ravi, Sahithya, et al.
Published: (2025)
by: Ravi, Sahithya, et al.
Published: (2025)
Beyond the Last Frame: Process-aware Evaluation for Generative Video Reasoning
by: Li, Yifan, et al.
Published: (2025)
by: Li, Yifan, et al.
Published: (2025)
Point-It-Out: Benchmarking Embodied Reasoning for Vision Language Models in Multi-Stage Visual Grounding
by: Xue, Haotian, et al.
Published: (2025)
by: Xue, Haotian, et al.
Published: (2025)
DengueNet: Dengue Prediction using Spatiotemporal Satellite Imagery for Resource-Limited Countries
by: Kuo, Kuan-Ting, et al.
Published: (2024)
by: Kuo, Kuan-Ting, et al.
Published: (2024)
Pixel is a Barrier: Diffusion Models Are More Adversarially Robust Than We Think
by: Xue, Haotian, et al.
Published: (2024)
by: Xue, Haotian, et al.
Published: (2024)
Diffusion Policy Attacker: Crafting Adversarial Attacks for Diffusion-based Policies
by: Chen, Yipu, et al.
Published: (2024)
by: Chen, Yipu, et al.
Published: (2024)
End-to-End Temporal Action Detection with 1B Parameters Across 1000 Frames
by: Liu, Shuming, et al.
Published: (2023)
by: Liu, Shuming, et al.
Published: (2023)
A Single-Frame and Multi-Frame Cascaded Image Super-Resolution Method
by: Sun, Jing, et al.
Published: (2024)
by: Sun, Jing, et al.
Published: (2024)
ClickDiffusion: Harnessing LLMs for Interactive Precise Image Editing
by: Helbling, Alec, et al.
Published: (2024)
by: Helbling, Alec, et al.
Published: (2024)
Chain-of-Frames: Advancing Video Understanding in Multimodal LLMs via Frame-Aware Reasoning
by: Ghazanfari, Sara, et al.
Published: (2025)
by: Ghazanfari, Sara, et al.
Published: (2025)
Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Models
by: Xu, Runsen, et al.
Published: (2025)
by: Xu, Runsen, et al.
Published: (2025)
Scene Summarization: Clustering Scene Videos into Spatially Diverse Frames
by: Chen, Chao, et al.
Published: (2023)
by: Chen, Chao, et al.
Published: (2023)
REM: Evaluating LLM Embodied Spatial Reasoning through Multi-Frame Trajectories
by: Thompson, Jacob, et al.
Published: (2025)
by: Thompson, Jacob, et al.
Published: (2025)
M${^2}$Depth: Self-supervised Two-Frame Multi-camera Metric Depth Estimation
by: Zou, Yingshuang, et al.
Published: (2024)
by: Zou, Yingshuang, et al.
Published: (2024)
FrameMind: Frame-Interleaved Video Reasoning via Reinforcement Learning
by: Ge, Haonan, et al.
Published: (2025)
by: Ge, Haonan, et al.
Published: (2025)
Semantic Frame Interpolation
by: Hong, Yijia, et al.
Published: (2025)
by: Hong, Yijia, et al.
Published: (2025)
Progress-Aware Video Frame Captioning
by: Xue, Zihui, et al.
Published: (2024)
by: Xue, Zihui, et al.
Published: (2024)
FrameThinker: Learning to Think with Long Videos via Multi-Turn Frame Spotlighting
by: He, Zefeng, et al.
Published: (2025)
by: He, Zefeng, et al.
Published: (2025)
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs
by: Zhang, Shaojie, et al.
Published: (2025)
by: Zhang, Shaojie, et al.
Published: (2025)
Beyond the Frame: Single and mutilple video summarization method with user-defined length
by: Kalkhorani, Vahid Ahmadi, et al.
Published: (2023)
by: Kalkhorani, Vahid Ahmadi, et al.
Published: (2023)
Zero-Shot Video Translation and Editing with Frame Spatial-Temporal Correspondence
by: Yang, Shuai, et al.
Published: (2025)
by: Yang, Shuai, et al.
Published: (2025)
RefDrop: Controllable Consistency in Image or Video Generation via Reference Feature Guidance
by: Fan, Jiaojiao, et al.
Published: (2024)
by: Fan, Jiaojiao, et al.
Published: (2024)
Understanding Multimodal Complementarity for Single-Frame Action Anticipation
by: Benavent-Lledo, Manuel, et al.
Published: (2026)
by: Benavent-Lledo, Manuel, et al.
Published: (2026)
Spatial-Temporal Graph Enhanced DETR Towards Multi-Frame 3D Object Detection
by: Zhang, Yifan, et al.
Published: (2023)
by: Zhang, Yifan, et al.
Published: (2023)
A.I.R.: Enabling Adaptive, Iterative, and Reasoning-based Frame Selection For Video Question Answering
by: Zou, Yuanhao, et al.
Published: (2025)
by: Zou, Yuanhao, et al.
Published: (2025)
Frame by Familiar Frame: Understanding Replication in Video Diffusion Models
by: Rahman, Aimon, et al.
Published: (2024)
by: Rahman, Aimon, et al.
Published: (2024)
Tinted Frames: Question Framing Blinds Vision-Language Models
by: Fan, Wan-Cyuan, et al.
Published: (2026)
by: Fan, Wan-Cyuan, et al.
Published: (2026)
GridVAD: Open-Set Video Anomaly Detection via Spatial Reasoning over Stratified Frame Grids
by: Eltahir, Mohamed, et al.
Published: (2026)
by: Eltahir, Mohamed, et al.
Published: (2026)
Poseidon: A ViT-based Architecture for Multi-Frame Pose Estimation with Adaptive Frame Weighting and Multi-Scale Feature Fusion
by: Pace, Cesare Davide, et al.
Published: (2025)
by: Pace, Cesare Davide, et al.
Published: (2025)
Framer: Interactive Frame Interpolation
by: Wang, Wen, et al.
Published: (2024)
by: Wang, Wen, et al.
Published: (2024)
Motion-Aware Video Frame Interpolation
by: Han, Pengfei, et al.
Published: (2024)
by: Han, Pengfei, et al.
Published: (2024)
Beyond Independent Frames: Latent Attention Masked Autoencoders for Multi-View Echocardiography
by: Böhi, Simon, et al.
Published: (2026)
by: Böhi, Simon, et al.
Published: (2026)
Gaze Beyond the Frame: Forecasting Egocentric 3D Visual Span
by: Yun, Heeseung, et al.
Published: (2025)
by: Yun, Heeseung, et al.
Published: (2025)
Beyond the Frame: Generating 360 Panoramic Videos from Perspective Videos
by: Luo, Rundong, et al.
Published: (2025)
by: Luo, Rundong, et al.
Published: (2025)
Video Finetuning Improves Reasoning Between Frames
by: Yang, Ruiqi, et al.
Published: (2025)
by: Yang, Ruiqi, et al.
Published: (2025)
DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding
by: Zhang, Peng, et al.
Published: (2026)
by: Zhang, Peng, et al.
Published: (2026)
Enhancing Video Inpainting with Aligned Frame Interval Guidance
by: Xie, Ming, et al.
Published: (2025)
by: Xie, Ming, et al.
Published: (2025)
Similar Items
-
ACWM-Phys: Investigating Generalized Physical Interaction in Action-Conditioned Video World Models
by: Xue, Haotian, et al.
Published: (2026) -
Looking Beyond What You See: An Empirical Analysis on Subgroup Intersectional Fairness for Multi-label Chest X-ray Classification Using Social Determinants of Racial Health Inequities
by: Moukheiber, Dana, et al.
Published: (2024) -
Laplacian Multi-scale Flow Matching for Generative Modeling
by: Zhao, Zelin, et al.
Published: (2026) -
Out of Sight, Not Out of Context? Egocentric Spatial Reasoning in VLMs Across Disjoint Frames
by: Ravi, Sahithya, et al.
Published: (2025) -
Beyond the Last Frame: Process-aware Evaluation for Generative Video Reasoning
by: Li, Yifan, et al.
Published: (2025)