Saved in:
| Main Authors: | Zhang, Mingyu, Zhuo, Lifeng, Tan, Tianxi, Xie, Guocan, Nie, Xian, Li, Yan, Zhao, Renjie, He, Zizhu, Wang, Ziyu, Cai, Jiting, Li, Yong-Lu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2511.15407 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Take A Step Back: Rethinking the Two Stages in Visual Reasoning
by: Zhang, Mingyu, et al.
Published: (2024)
by: Zhang, Mingyu, et al.
Published: (2024)
Architect: Generating Vivid and Interactive 3D Scenes with Hierarchical 2D Inpainting
by: Wang, Yian, et al.
Published: (2024)
by: Wang, Yian, et al.
Published: (2024)
FilmSceneDesigner: Chaining Set Design for Procedural Film Scene Generation
by: Xie, Zhifeng, et al.
Published: (2025)
by: Xie, Zhifeng, et al.
Published: (2025)
LSVG: Language-Guided Scene Graphs with 2D-Assisted Multi-Modal Encoding for 3D Visual Grounding
by: Xiao, Feng, et al.
Published: (2025)
by: Xiao, Feng, et al.
Published: (2025)
IPR-NeRF: Ownership Verification meets Neural Radiance Field
by: Ong, Win Kent, et al.
Published: (2024)
by: Ong, Win Kent, et al.
Published: (2024)
The non-overlapping statistical approximation to overlapping group lasso
by: Qi, Mingyu, et al.
Published: (2022)
by: Qi, Mingyu, et al.
Published: (2022)
Light Interaction: Training-Free Inference Acceleration for Interactive Video World Models
by: Lu, Jiacheng, et al.
Published: (2026)
by: Lu, Jiacheng, et al.
Published: (2026)
RD-ViT: Recurrent-Depth Vision Transformer for Semantic Segmentation with Reduced Data Dependence Extending the Recurrent-Depth Transformer Architecture to Dense Prediction
by: He, Renjie
Published: (2026)
by: He, Renjie
Published: (2026)
FocusTrack: One-Stage Focus-and-Suppress Framework for 3D Point Cloud Object Tracking
by: Zhou, Sifan, et al.
Published: (2026)
by: Zhou, Sifan, et al.
Published: (2026)
Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning
by: Liang, Dayong, et al.
Published: (2025)
by: Liang, Dayong, et al.
Published: (2025)
Dancing with Still Images: Video Distillation via Static-Dynamic Disentanglement
by: Wang, Ziyu, et al.
Published: (2023)
by: Wang, Ziyu, et al.
Published: (2023)
EchoVideo: Identity-Preserving Human Video Generation by Multimodal Feature Fusion
by: Wei, Jiangchuan, et al.
Published: (2025)
by: Wei, Jiangchuan, et al.
Published: (2025)
Look Less, Reason More: Rollout-Guided Adaptive Pixel-Space Reasoning
by: Li, Xuchen, et al.
Published: (2025)
by: Li, Xuchen, et al.
Published: (2025)
Efficient Temporal Sentence Grounding in Videos with Multi-Teacher Knowledge Distillation
by: Liang, Renjie, et al.
Published: (2023)
by: Liang, Renjie, et al.
Published: (2023)
Thinking Diffusion: Penalize and Guide Visual-Grounded Reasoning in Diffusion Multimodal Language Models
by: Kim, Keuntae, et al.
Published: (2026)
by: Kim, Keuntae, et al.
Published: (2026)
ReasonDrive: Efficient Visual Question Answering for Autonomous Vehicles with Reasoning-Enhanced Small Vision-Language Models
by: Chahe, Amirhosein, et al.
Published: (2025)
by: Chahe, Amirhosein, et al.
Published: (2025)
CounterCurate: Enhancing Physical and Semantic Visio-Linguistic Compositional Reasoning via Counterfactual Examples
by: Zhang, Jianrui, et al.
Published: (2024)
by: Zhang, Jianrui, et al.
Published: (2024)
Tracking with Human-Intent Reasoning
by: Zhu, Jiawen, et al.
Published: (2023)
by: Zhu, Jiawen, et al.
Published: (2023)
Reason2Drive: Towards Interpretable and Chain-based Reasoning for Autonomous Driving
by: Nie, Ming, et al.
Published: (2023)
by: Nie, Ming, et al.
Published: (2023)
Semantic Visual Anomaly Detection and Reasoning in AI-Generated Images
by: Tan, Chuangchuang, et al.
Published: (2025)
by: Tan, Chuangchuang, et al.
Published: (2025)
Leveraging Adaptive Implicit Representation Mapping for Ultra High-Resolution Image Segmentation
by: Zhao, Ziyu, et al.
Published: (2024)
by: Zhao, Ziyu, et al.
Published: (2024)
RESBev: Making BEV Perception More Robust
by: Zhuo, Lifeng, et al.
Published: (2026)
by: Zhuo, Lifeng, et al.
Published: (2026)
Compositional Physical Reasoning of Objects and Events from Videos
by: Chen, Zhenfang, et al.
Published: (2024)
by: Chen, Zhenfang, et al.
Published: (2024)
I-PHYRE: Interactive Physical Reasoning
by: Li, Shiqian, et al.
Published: (2023)
by: Li, Shiqian, et al.
Published: (2023)
Diverse Generation while Maintaining Semantic Coordination: A Diffusion-Based Data Augmentation Method for Object Detection
by: Nie, Sen, et al.
Published: (2024)
by: Nie, Sen, et al.
Published: (2024)
Evo-0: Vision-Language-Action Model with Implicit Spatial Understanding
by: Lin, Tao, et al.
Published: (2025)
by: Lin, Tao, et al.
Published: (2025)
Semantic Event Graphs for Long-Form Video Question Answering
by: Dixit, Aradhya, et al.
Published: (2026)
by: Dixit, Aradhya, et al.
Published: (2026)
LuciBot: Automated Robot Policy Learning from Generated Videos
by: Qiu, Xiaowen, et al.
Published: (2025)
by: Qiu, Xiaowen, et al.
Published: (2025)
Towards Secure and Usable 3D Assets: A Novel Framework for Automatic Visible Watermarking
by: Singh, Gursimran, et al.
Published: (2024)
by: Singh, Gursimran, et al.
Published: (2024)
Dance of Fireworks: An Interactive Broadcast Gymnastics Training System Based on Pose Estimation
by: Chen, Haotian, et al.
Published: (2025)
by: Chen, Haotian, et al.
Published: (2025)
MicroVQA++: High-Quality Microscopy Reasoning Dataset with Weakly Supervised Graphs for Multimodal Large Language Model
by: Li, Manyu, et al.
Published: (2025)
by: Li, Manyu, et al.
Published: (2025)
FMLGS: Fast Multilevel Language Embedded Gaussians for Part-level Interactive Agents
by: Tan, Xin, et al.
Published: (2025)
by: Tan, Xin, et al.
Published: (2025)
IntentionVLA: Generalizable and Efficient Embodied Intention Reasoning for Human-Robot Interaction
by: Chen, Yandu, et al.
Published: (2025)
by: Chen, Yandu, et al.
Published: (2025)
Exploring Interactive Semantic Alignment for Efficient HOI Detection with Vision-language Model
by: Dong, Jihao, et al.
Published: (2024)
by: Dong, Jihao, et al.
Published: (2024)
S1-VL: Scientific Multimodal Reasoning Model with Thinking-with-Images
by: Li, Qingxiao, et al.
Published: (2026)
by: Li, Qingxiao, et al.
Published: (2026)
Offline-Poly: A Polyhedral Framework For Offline 3D Multi-Object Tracking
by: Li, Xiaoyu, et al.
Published: (2026)
by: Li, Xiaoyu, et al.
Published: (2026)
Reinforce to Learn, Elect to Reason: A Dual Paradigm for Video Reasoning
by: Yang, Songyuan, et al.
Published: (2026)
by: Yang, Songyuan, et al.
Published: (2026)
SpecAware: A Spectral-Content Aware Foundation Model for Unifying Multi-Sensor Learning in Hyperspectral Remote Sensing Mapping
by: Ji, Renjie, et al.
Published: (2025)
by: Ji, Renjie, et al.
Published: (2025)
EditThinker: Unlocking Iterative Reasoning for Any Image Editor
by: Li, Hongyu, et al.
Published: (2025)
by: Li, Hongyu, et al.
Published: (2025)
CompTrack: Information Bottleneck-Guided Low-Rank Dynamic Token Compression for Point Cloud Tracking
by: Zhou, Sifan, et al.
Published: (2025)
by: Zhou, Sifan, et al.
Published: (2025)
Similar Items
-
Take A Step Back: Rethinking the Two Stages in Visual Reasoning
by: Zhang, Mingyu, et al.
Published: (2024) -
Architect: Generating Vivid and Interactive 3D Scenes with Hierarchical 2D Inpainting
by: Wang, Yian, et al.
Published: (2024) -
FilmSceneDesigner: Chaining Set Design for Procedural Film Scene Generation
by: Xie, Zhifeng, et al.
Published: (2025) -
LSVG: Language-Guided Scene Graphs with 2D-Assisted Multi-Modal Encoding for 3D Visual Grounding
by: Xiao, Feng, et al.
Published: (2025) -
IPR-NeRF: Ownership Verification meets Neural Radiance Field
by: Ong, Win Kent, et al.
Published: (2024)