Improved Visual-Spatial Reasoning via R1-Zero-Like Training
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liao, Zhenyi, Xie, Qingsong, Zhang, Yanhao, Kong, Zijian, Lu, Haonan, Yang, Zhenyu, Deng, Zhijie |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TLCM: Training-efficient Latent Consistency Model for Image Generation with 2-8 Steps
von: Xie, Qingsong, et al.
Veröffentlicht: (2024)
von: Xie, Qingsong, et al.
Veröffentlicht: (2024)
LOVECon: Text-driven Training-Free Long Video Editing with ControlNet
von: Liao, Zhenyi, et al.
Veröffentlicht: (2023)
von: Liao, Zhenyi, et al.
Veröffentlicht: (2023)
FaceScore: Benchmarking and Enhancing Face Quality in Human Generation
von: Liao, Zhenyi, et al.
Veröffentlicht: (2024)
von: Liao, Zhenyi, et al.
Veröffentlicht: (2024)
Layton: Latent Consistency Tokenizer for 1024-pixel Image Reconstruction and Generation by 256 Tokens
von: Xie, Qingsong, et al.
Veröffentlicht: (2025)
von: Xie, Qingsong, et al.
Veröffentlicht: (2025)
Advancing Text-to-3D Generation with Linearized Lookahead Variational Score Distillation
von: Lei, Yu, et al.
Veröffentlicht: (2025)
von: Lei, Yu, et al.
Veröffentlicht: (2025)
MergeTok: Unified Continuous and Discrete Visual Tokenization via Token Merging
von: Zhang, Luyuan, et al.
Veröffentlicht: (2026)
von: Zhang, Luyuan, et al.
Veröffentlicht: (2026)
Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM
von: Liu, Peng, et al.
Veröffentlicht: (2025)
von: Liu, Peng, et al.
Veröffentlicht: (2025)
H2VU-Benchmark: A Comprehensive Benchmark for Hierarchical Holistic Video Understanding
von: Wu, Qi, et al.
Veröffentlicht: (2025)
von: Wu, Qi, et al.
Veröffentlicht: (2025)
pySpatial: Generating 3D Visual Programs for Zero-Shot Spatial Reasoning
von: Luo, Zhanpeng, et al.
Veröffentlicht: (2026)
von: Luo, Zhanpeng, et al.
Veröffentlicht: (2026)
OASIS: On-Demand Hierarchical Event Memory for Streaming Video Reasoning
von: Liang, Zhijia, et al.
Veröffentlicht: (2026)
von: Liang, Zhijia, et al.
Veröffentlicht: (2026)
DaMo: Data Mixing Optimizer in Fine-tuning Multimodal LLMs for Mobile Phone Agents
von: Shi, Kai, et al.
Veröffentlicht: (2025)
von: Shi, Kai, et al.
Veröffentlicht: (2025)
SCott: Accelerating Diffusion Models with Stochastic Consistency Distillation
von: Liu, Hongjian, et al.
Veröffentlicht: (2024)
von: Liu, Hongjian, et al.
Veröffentlicht: (2024)
SPAZER: Spatial-Semantic Progressive Reasoning Agent for Zero-shot 3D Visual Grounding
von: Jin, Zhao, et al.
Veröffentlicht: (2025)
von: Jin, Zhao, et al.
Veröffentlicht: (2025)
ViGoR-Bench: How Far Are Visual Generative Models From Zero-Shot Visual Reasoners?
von: Han, Haonan, et al.
Veröffentlicht: (2026)
von: Han, Haonan, et al.
Veröffentlicht: (2026)
TTSA3R: Training-Free Temporal-Spatial Adaptive Persistent State for Streaming 3D Reconstruction
von: Zheng, Zhijie, et al.
Veröffentlicht: (2026)
von: Zheng, Zhijie, et al.
Veröffentlicht: (2026)
UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding
von: Xu, Chenkai, et al.
Veröffentlicht: (2025)
von: Xu, Chenkai, et al.
Veröffentlicht: (2025)
PEA-Diffusion: Parameter-Efficient Adapter with Knowledge Distillation in non-English Text-to-Image Generation
von: Ma, Jian, et al.
Veröffentlicht: (2023)
von: Ma, Jian, et al.
Veröffentlicht: (2023)
LAPTOP-Diff: Layer Pruning and Normalized Distillation for Compressing Diffusion Models
von: Zhang, Dingkun, et al.
Veröffentlicht: (2024)
von: Zhang, Dingkun, et al.
Veröffentlicht: (2024)
ViCToR: Improving Visual Comprehension via Token Reconstruction for Pretraining LMMs
von: Xie, Yin, et al.
Veröffentlicht: (2024)
von: Xie, Yin, et al.
Veröffentlicht: (2024)
Refining CLIP's Spatial Awareness: A Visual-Centric Perspective
von: Qiu, Congpei, et al.
Veröffentlicht: (2025)
von: Qiu, Congpei, et al.
Veröffentlicht: (2025)
NoisyRollout: Reinforcing Visual Reasoning with Data Augmentation
von: Liu, Xiangyan, et al.
Veröffentlicht: (2025)
von: Liu, Xiangyan, et al.
Veröffentlicht: (2025)
Learning to Localize Objects Improves Spatial Reasoning in Visual-LLMs
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2024)
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2024)
VideoINSTA: Zero-shot Long Video Understanding via Informative Spatial-Temporal Reasoning with LLMs
von: Liao, Ruotong, et al.
Veröffentlicht: (2024)
von: Liao, Ruotong, et al.
Veröffentlicht: (2024)
X-OmniClaw Technical Report: A Unified Mobile Agent for Multimodal Understanding and Interaction
von: Ren, Xiaoming, et al.
Veröffentlicht: (2026)
von: Ren, Xiaoming, et al.
Veröffentlicht: (2026)
GlyphDraw2: Automatic Generation of Complex Glyph Posters with Diffusion Models and Large Language Models
von: Ma, Jian, et al.
Veröffentlicht: (2024)
von: Ma, Jian, et al.
Veröffentlicht: (2024)
GeoR-Bench: Evaluating Geoscience Visual Reasoning
von: Zheng, Yushuo, et al.
Veröffentlicht: (2026)
von: Zheng, Yushuo, et al.
Veröffentlicht: (2026)
PainterNet: Adaptive Image Inpainting with Actual-Token Attention and Diverse Mask Control
von: Wang, Ruichen, et al.
Veröffentlicht: (2024)
von: Wang, Ruichen, et al.
Veröffentlicht: (2024)
Training-free Boost for Open-Vocabulary Object Detection with Confidence Aggregation
von: Zheng, Yanhao, et al.
Veröffentlicht: (2024)
von: Zheng, Yanhao, et al.
Veröffentlicht: (2024)
DMTrack: Spatio-Temporal Multimodal Tracking via Dual-Adapter
von: Li, Weihong, et al.
Veröffentlicht: (2025)
von: Li, Weihong, et al.
Veröffentlicht: (2025)
GUI-G1: Understanding R1-Zero-Like Training for Visual Grounding in GUI Agents
von: Zhou, Yuqi, et al.
Veröffentlicht: (2025)
von: Zhou, Yuqi, et al.
Veröffentlicht: (2025)
OpenSeg-R: Improving Open-Vocabulary Segmentation via Step-by-Step Visual Reasoning
von: Han, Zongyan, et al.
Veröffentlicht: (2025)
von: Han, Zongyan, et al.
Veröffentlicht: (2025)
SpatialImaginer: Towards Adaptive Visual Imagination for Spatial Reasoning
von: Li, Yian, et al.
Veröffentlicht: (2026)
von: Li, Yian, et al.
Veröffentlicht: (2026)
SpatialReward: Bridging the Perception Gap in Online RL for Image Editing via Explicit Spatial Reasoning
von: Long, Yancheng, et al.
Veröffentlicht: (2026)
von: Long, Yancheng, et al.
Veröffentlicht: (2026)
X2I: Seamless Integration of Multimodal Understanding into Diffusion Transformer via Attention Distillation
von: Ma, Jian, et al.
Veröffentlicht: (2025)
von: Ma, Jian, et al.
Veröffentlicht: (2025)
ZEBRA: Towards Zero-Shot Cross-Subject Generalization for Universal Brain Visual Decoding
von: Wang, Haonan, et al.
Veröffentlicht: (2025)
von: Wang, Haonan, et al.
Veröffentlicht: (2025)
Geometrically-Constrained Agent for Spatial Reasoning
von: Chen, Zeren, et al.
Veröffentlicht: (2025)
von: Chen, Zeren, et al.
Veröffentlicht: (2025)
MetaSpatial: Reinforcing 3D Spatial Reasoning in VLMs for the Metaverse
von: Pan, Zhenyu, et al.
Veröffentlicht: (2025)
von: Pan, Zhenyu, et al.
Veröffentlicht: (2025)
Visual Jigsaw Post-Training Improves MLLMs
von: Wu, Penghao, et al.
Veröffentlicht: (2025)
von: Wu, Penghao, et al.
Veröffentlicht: (2025)
Improving Generalization in Visual Reasoning via Self-Ensemble
von: Nguyen, Tien-Huy, et al.
Veröffentlicht: (2024)
von: Nguyen, Tien-Huy, et al.
Veröffentlicht: (2024)
Learning from Fine-Grained Visual Discrepancies: Mitigating Multimodal Hallucinations via In-Context Visual Contrastive Optimization
von: Deng, Haolin, et al.
Veröffentlicht: (2026)
von: Deng, Haolin, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
TLCM: Training-efficient Latent Consistency Model for Image Generation with 2-8 Steps
von: Xie, Qingsong, et al.
Veröffentlicht: (2024) -
LOVECon: Text-driven Training-Free Long Video Editing with ControlNet
von: Liao, Zhenyi, et al.
Veröffentlicht: (2023) -
FaceScore: Benchmarking and Enhancing Face Quality in Human Generation
von: Liao, Zhenyi, et al.
Veröffentlicht: (2024) -
Layton: Latent Consistency Tokenizer for 1024-pixel Image Reconstruction and Generation by 256 Tokens
von: Xie, Qingsong, et al.
Veröffentlicht: (2025) -
Advancing Text-to-3D Generation with Linearized Lookahead Variational Score Distillation
von: Lei, Yu, et al.
Veröffentlicht: (2025)