SIFThinker: Spatially-Aware Image Focus for Visual Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Zhangquan, Zhao, Ruihui, Luo, Chuwei, Sun, Mingze, Yu, Xinlei, Kang, Yangyang, Huang, Ruqi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Think with 3D: Geometric Imagination Grounded Spatial Reasoning from Limited Views
von: Chen, Zhangquan, et al.
Veröffentlicht: (2025)
von: Chen, Zhangquan, et al.
Veröffentlicht: (2025)
4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding
von: Chen, Zhangquan, et al.
Veröffentlicht: (2026)
von: Chen, Zhangquan, et al.
Veröffentlicht: (2026)
VisRL: Intention-Driven Visual Perception via Reinforced Reasoning
von: Chen, Zhangquan, et al.
Veröffentlicht: (2025)
von: Chen, Zhangquan, et al.
Veröffentlicht: (2025)
OmniVideo-R1: Reinforcing Audio-visual Reasoning with Query Intention and Modality Attention
von: Chen, Zhangquan, et al.
Veröffentlicht: (2026)
von: Chen, Zhangquan, et al.
Veröffentlicht: (2026)
GeoHeight-Bench: Towards Height-Aware Multimodal Reasoning in Remote Sensing
von: Hu, Xuran, et al.
Veröffentlicht: (2026)
von: Hu, Xuran, et al.
Veröffentlicht: (2026)
Visual Enhanced Depth Scaling for Multimodal Latent Reasoning
von: Han, Yudong, et al.
Veröffentlicht: (2026)
von: Han, Yudong, et al.
Veröffentlicht: (2026)
U-Net-Like Spiking Neural Networks for Single Image Dehazing
von: Li, Huibin, et al.
Veröffentlicht: (2025)
von: Li, Huibin, et al.
Veröffentlicht: (2025)
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding
von: Han, Yudong, et al.
Veröffentlicht: (2024)
von: Han, Yudong, et al.
Veröffentlicht: (2024)
Context-Aware Indoor Point Cloud Object Generation through User Instructions
von: Luo, Yiyang, et al.
Veröffentlicht: (2023)
von: Luo, Yiyang, et al.
Veröffentlicht: (2023)
FocusedAD: Character-centric Movie Audio Description
von: Ye, Xiaojun, et al.
Veröffentlicht: (2025)
von: Ye, Xiaojun, et al.
Veröffentlicht: (2025)
One Leaf Reveals the Season: Occlusion-Based Contrastive Learning with Semantic-Aware Views for Efficient Visual Representation
von: Yang, Xiaoyu, et al.
Veröffentlicht: (2024)
von: Yang, Xiaoyu, et al.
Veröffentlicht: (2024)
FoR-Net: Learning to Focus on Hard Regions for Efficient Semantic Segmentation
von: Chan, Sheng-Wei, et al.
Veröffentlicht: (2026)
von: Chan, Sheng-Wei, et al.
Veröffentlicht: (2026)
Foreground Focus: Enhancing Coherence and Fidelity in Camouflaged Image Generation
von: Chen, Pei-Chi, et al.
Veröffentlicht: (2025)
von: Chen, Pei-Chi, et al.
Veröffentlicht: (2025)
An Evaluation of a Visual Question Answering Strategy for Zero-shot Facial Expression Recognition in Still Images
von: Castrillón-Santana, Modesto, et al.
Veröffentlicht: (2025)
von: Castrillón-Santana, Modesto, et al.
Veröffentlicht: (2025)
Look, Listen, and Answer: Overcoming Biases for Audio-Visual Question Answering
von: Ma, Jie, et al.
Veröffentlicht: (2024)
von: Ma, Jie, et al.
Veröffentlicht: (2024)
Beyond Localization: A Comprehensive Diagnosis of Perspective-Conditioned Spatial Reasoning in MLLMs from Omnidirectional Images
von: Chen, Yuangong, et al.
Veröffentlicht: (2026)
von: Chen, Yuangong, et al.
Veröffentlicht: (2026)
Vectra: A New Metric, Dataset, and Model for Visual Quality Assessment in E-Commerce In-Image Machine Translation
von: Wu, Qingyu, et al.
Veröffentlicht: (2026)
von: Wu, Qingyu, et al.
Veröffentlicht: (2026)
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
von: Raoufi, Behnam, et al.
Veröffentlicht: (2025)
von: Raoufi, Behnam, et al.
Veröffentlicht: (2025)
Textual and Visual Guided Task Adaptation for Source-Free Cross-Domain Few-Shot Segmentation
von: Liu, Jianming, et al.
Veröffentlicht: (2025)
von: Liu, Jianming, et al.
Veröffentlicht: (2025)
MSTA3D: Multi-scale Twin-attention for 3D Instance Segmentation
von: Tran, Duc Dang Trung, et al.
Veröffentlicht: (2024)
von: Tran, Duc Dang Trung, et al.
Veröffentlicht: (2024)
FUSE-Flow: Scalable Real-Time Multi-View Point Cloud Reconstruction Using Confidence
von: Sun, Chentian
Veröffentlicht: (2026)
von: Sun, Chentian
Veröffentlicht: (2026)
GMAC: Global Multi-View Constraint for Automatic Multi-Camera Extrinsic Calibration
von: Sun, Chentian
Veröffentlicht: (2026)
von: Sun, Chentian
Veröffentlicht: (2026)
SPARK: Scalable Real-Time Point Cloud Aggregation with Multi-View Self-Calibration
von: Sun, Chentian
Veröffentlicht: (2026)
von: Sun, Chentian
Veröffentlicht: (2026)
Instruction-based Image Editing with Planning, Reasoning, and Generation
von: Ji, Liya, et al.
Veröffentlicht: (2026)
von: Ji, Liya, et al.
Veröffentlicht: (2026)
Symmetry Awareness Encoded Deep Learning Framework for Brain Imaging Analysis
von: Ma, Yang, et al.
Veröffentlicht: (2024)
von: Ma, Yang, et al.
Veröffentlicht: (2024)
LEGO: LoRA-Enabled Generator-Oriented Framework for Synthetic Image Detection
von: Xiao, Yutong, et al.
Veröffentlicht: (2026)
von: Xiao, Yutong, et al.
Veröffentlicht: (2026)
Multi-modal Sensor Fusion for Auto Driving Perception: A Survey
von: Huang, Keli, et al.
Veröffentlicht: (2022)
von: Huang, Keli, et al.
Veröffentlicht: (2022)
Perceptual Flow Network for Visually Grounded Reasoning
von: Li, Yangfu, et al.
Veröffentlicht: (2026)
von: Li, Yangfu, et al.
Veröffentlicht: (2026)
Vi-SAFE: A Spatial-Temporal Framework for Efficient Violence Detection in Public Surveillance
von: Chang, Ligang, et al.
Veröffentlicht: (2025)
von: Chang, Ligang, et al.
Veröffentlicht: (2025)
Privacy-Preserving Structureless Visual Localization via Image Obfuscation
von: Panek, Vojtech, et al.
Veröffentlicht: (2026)
von: Panek, Vojtech, et al.
Veröffentlicht: (2026)
Topology-Aware Latent Diffusion for 3D Shape Generation
von: Hu, Jiangbei, et al.
Veröffentlicht: (2024)
von: Hu, Jiangbei, et al.
Veröffentlicht: (2024)
A Recipe for Geometry-Aware 3D Mesh Transformers
von: Farazi, Mohammad, et al.
Veröffentlicht: (2024)
von: Farazi, Mohammad, et al.
Veröffentlicht: (2024)
SCA-Net: Spatial-Contextual Aggregation Network for Enhanced Small Building and Road Change Detection
von: Gholibeigi, Emad, et al.
Veröffentlicht: (2026)
von: Gholibeigi, Emad, et al.
Veröffentlicht: (2026)
Multi-Scale Spatial-Temporal Self-Attention Graph Convolutional Networks for Skeleton-based Action Recognition
von: Nakamura, Ikuo
Veröffentlicht: (2024)
von: Nakamura, Ikuo
Veröffentlicht: (2024)
Semantic Prioritization in Visual Counterfactual Explanations with Weighted Segmentation and Auto-Adaptive Region Selection
von: Zhang, Lintong, et al.
Veröffentlicht: (2025)
von: Zhang, Lintong, et al.
Veröffentlicht: (2025)
Single-Shot Metric Depth from Focused Plenoptic Cameras
von: Lasheras-Hernandez, Blanca, et al.
Veröffentlicht: (2024)
von: Lasheras-Hernandez, Blanca, et al.
Veröffentlicht: (2024)
GroundCap: A Visually Grounded Image Captioning Dataset
von: Oliveira, Daniel A. P., et al.
Veröffentlicht: (2025)
von: Oliveira, Daniel A. P., et al.
Veröffentlicht: (2025)
VidNum-1.4K: A Comprehensive Benchmark for Video-based Numerical Reasoning
von: Cui, Shaoyang, et al.
Veröffentlicht: (2026)
von: Cui, Shaoyang, et al.
Veröffentlicht: (2026)
Video Event Reasoning and Prediction by Fusing World Knowledge from LLMs with Vision Foundation Models
von: Dubois, L'ea, et al.
Veröffentlicht: (2025)
von: Dubois, L'ea, et al.
Veröffentlicht: (2025)
Robust Visual Question Answering: Datasets, Methods, and Future Challenges
von: Ma, Jie, et al.
Veröffentlicht: (2023)
von: Ma, Jie, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Think with 3D: Geometric Imagination Grounded Spatial Reasoning from Limited Views
von: Chen, Zhangquan, et al.
Veröffentlicht: (2025) -
4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding
von: Chen, Zhangquan, et al.
Veröffentlicht: (2026) -
VisRL: Intention-Driven Visual Perception via Reinforced Reasoning
von: Chen, Zhangquan, et al.
Veröffentlicht: (2025) -
OmniVideo-R1: Reinforcing Audio-visual Reasoning with Query Intention and Modality Attention
von: Chen, Zhangquan, et al.
Veröffentlicht: (2026) -
GeoHeight-Bench: Towards Height-Aware Multimodal Reasoning in Remote Sensing
von: Hu, Xuran, et al.
Veröffentlicht: (2026)