Saved in:
| Main Authors: | Xuan, Wenjie, Zhang, Jing, Liu, Juhua, Du, Bo, Tao, Dacheng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2506.20983 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learn to Think: Improving Multimodal Reasoning through Vision-Aware Self-Improvement Training
by: Zhong, Qihuang, et al.
Published: (2026)
by: Zhong, Qihuang, et al.
Published: (2026)
GoMatching++: Parameter- and Data-Efficient Arbitrary-Shaped Video Text Spotting and Benchmarking
by: He, Haibin, et al.
Published: (2025)
by: He, Haibin, et al.
Published: (2025)
GoMatching: A Simple Baseline for Video Text Spotting via Long and Short Term Matching
by: He, Haibin, et al.
Published: (2024)
by: He, Haibin, et al.
Published: (2024)
When ControlNet Meets Inexplicit Masks: A Case Study of ControlNet on its Contour-following Ability
by: Xuan, Wenjie, et al.
Published: (2024)
by: Xuan, Wenjie, et al.
Published: (2024)
DeepSolo++: Let Transformer Decoder with Explicit Points Solo for Multilingual Text Spotting
by: Ye, Maoyuan, et al.
Published: (2023)
by: Ye, Maoyuan, et al.
Published: (2023)
Hi-SAM: Marrying Segment Anything Model for Hierarchical Text Segmentation
by: Ye, Maoyuan, et al.
Published: (2024)
by: Ye, Maoyuan, et al.
Published: (2024)
RFL-CDNet: Towards Accurate Change Detection via Richer Feature Learning
by: Gan, Yuhang, et al.
Published: (2024)
by: Gan, Yuhang, et al.
Published: (2024)
Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues?
by: He, Haibin, et al.
Published: (2025)
by: He, Haibin, et al.
Published: (2025)
VTAgent: Agentic Keyframe Anchoring for Evidence-Aware Video TextVQA
by: He, Haibin, et al.
Published: (2026)
by: He, Haibin, et al.
Published: (2026)
ET-SAM: Efficient Point Prompt Prediction in SAM for Unified Scene Text Detection and Layout Analysis
by: Zhang, Xike, et al.
Published: (2026)
by: Zhang, Xike, et al.
Published: (2026)
SFA: Scan, Focus, and Amplify toward Guidance-aware Answering for Video TextVQA
by: He, Haibin, et al.
Published: (2025)
by: He, Haibin, et al.
Published: (2025)
Detect Changes like Humans: Incorporating Semantic Priors for Improved Change Detection
by: Gan, Yuhang, et al.
Published: (2024)
by: Gan, Yuhang, et al.
Published: (2024)
LogicOCR: Do Your Large Multimodal Models Excel at Logical Reasoning on Text-Rich Images?
by: Ye, Maoyuan, et al.
Published: (2025)
by: Ye, Maoyuan, et al.
Published: (2025)
PoseBench: Benchmarking the Robustness of Pose Estimation Models under Corruptions
by: Ma, Sihan, et al.
Published: (2024)
by: Ma, Sihan, et al.
Published: (2024)
TeRA: Rethinking Text-guided Realistic 3D Avatar Generation
by: Wang, Yanwen, et al.
Published: (2025)
by: Wang, Yanwen, et al.
Published: (2025)
PolarMAE: Efficient Fetal Ultrasound Pre-training via Semantic Screening and Polar-Guided Masking
by: Lv, Meng, et al.
Published: (2026)
by: Lv, Meng, et al.
Published: (2026)
LAB-Det: Language as a Domain-Invariant Bridge for Training-Free One-Shot Domain Generalization in Object Detection
by: Zhang, Xu, et al.
Published: (2026)
by: Zhang, Xu, et al.
Published: (2026)
On Robust Cross-View Consistency in Self-Supervised Monocular Depth Estimation
by: Zhao, Haimei, et al.
Published: (2022)
by: Zhao, Haimei, et al.
Published: (2022)
Free-Form Motion Control: Controlling the 6D Poses of Camera and Objects in Video Generation
by: Shuai, Xincheng, et al.
Published: (2025)
by: Shuai, Xincheng, et al.
Published: (2025)
CogMorph: Cognitive Morphing Attacks for Text-to-Image Models
by: Jing, Zonglei, et al.
Published: (2025)
by: Jing, Zonglei, et al.
Published: (2025)
Rethinking Model Efficiency: Multi-Agent Inference with Large Models
by: Dong, Sixun, et al.
Published: (2026)
by: Dong, Sixun, et al.
Published: (2026)
Adapting Segment Anything Model for Power Transmission Corridor Hazard Segmentation
by: Chen, Hang, et al.
Published: (2025)
by: Chen, Hang, et al.
Published: (2025)
HandRefiner: Refining Malformed Hands in Generated Images by Diffusion-based Conditional Inpainting
by: Lu, Wenquan, et al.
Published: (2023)
by: Lu, Wenquan, et al.
Published: (2023)
Contact-aware Human Motion Generation from Textual Descriptions
by: Ma, Sihan, et al.
Published: (2024)
by: Ma, Sihan, et al.
Published: (2024)
Reverse Prompt: Cracking the Recipe Inside Text-to-Image Generation
by: Ren, Zhiyao, et al.
Published: (2025)
by: Ren, Zhiyao, et al.
Published: (2025)
Heuristic-inspired Reasoning Priors Facilitate Data-Efficient Referring Object Detection
by: Zhang, Xu, et al.
Published: (2026)
by: Zhang, Xu, et al.
Published: (2026)
Epistemic-aware Vision-Language Foundation Model for Fetal Ultrasound Interpretation
by: He, Xiao, et al.
Published: (2025)
by: He, Xiao, et al.
Published: (2025)
SimInversion: A Simple Framework for Inversion-Based Text-to-Image Editing
by: Qian, Qi, et al.
Published: (2024)
by: Qian, Qi, et al.
Published: (2024)
AvatarPose: Avatar-guided 3D Pose Estimation of Close Human Interaction from Sparse Multi-view Videos
by: Lu, Feichi, et al.
Published: (2024)
by: Lu, Feichi, et al.
Published: (2024)
Image Captions are Natural Prompts for Text-to-Image Models
by: Lei, Shiye, et al.
Published: (2023)
by: Lei, Shiye, et al.
Published: (2023)
PoseAnything: Universal Pose-guided Video Generation with Part-aware Temporal Coherence
by: Wang, Ruiyan, et al.
Published: (2025)
by: Wang, Ruiyan, et al.
Published: (2025)
Rethinking Structure Preservation in Text-Guided Image Editing with Visual Autoregressive Models
by: Xia, Tao, et al.
Published: (2026)
by: Xia, Tao, et al.
Published: (2026)
Event-based Simultaneous Localization and Mapping: A Comprehensive Survey
by: Huang, Kunping, et al.
Published: (2023)
by: Huang, Kunping, et al.
Published: (2023)
Object Gaussian for Monocular 6D Pose Estimation from Sparse Views
by: Luo, Luqing, et al.
Published: (2024)
by: Luo, Luqing, et al.
Published: (2024)
Echo-α: Large Agentic Multimodal Reasoning Model for Ultrasound Interpretation
by: Zhang, Jing, et al.
Published: (2026)
by: Zhang, Jing, et al.
Published: (2026)
GlyphPrinter: Region-Grouped Direct Preference Optimization for Glyph-Accurate Visual Text Rendering
by: Shuai, Xincheng, et al.
Published: (2026)
by: Shuai, Xincheng, et al.
Published: (2026)
Text-Guided Mixup Towards Long-Tailed Image Categorization
by: Franklin, Richard, et al.
Published: (2024)
by: Franklin, Richard, et al.
Published: (2024)
Stable-Pose: Leveraging Transformers for Pose-Guided Text-to-Image Generation
by: Wang, Jiajun, et al.
Published: (2024)
by: Wang, Jiajun, et al.
Published: (2024)
Text-guided Zero-Shot Object Localization
by: Wang, Jingjing, et al.
Published: (2024)
by: Wang, Jingjing, et al.
Published: (2024)
VORTA: Efficient Video Diffusion via Routing Sparse Attention
by: Sun, Wenhao, et al.
Published: (2025)
by: Sun, Wenhao, et al.
Published: (2025)
Similar Items
-
Learn to Think: Improving Multimodal Reasoning through Vision-Aware Self-Improvement Training
by: Zhong, Qihuang, et al.
Published: (2026) -
GoMatching++: Parameter- and Data-Efficient Arbitrary-Shaped Video Text Spotting and Benchmarking
by: He, Haibin, et al.
Published: (2025) -
GoMatching: A Simple Baseline for Video Text Spotting via Long and Short Term Matching
by: He, Haibin, et al.
Published: (2024) -
When ControlNet Meets Inexplicit Masks: A Case Study of ControlNet on its Contour-following Ability
by: Xuan, Wenjie, et al.
Published: (2024) -
DeepSolo++: Let Transformer Decoder with Explicit Points Solo for Multilingual Text Spotting
by: Ye, Maoyuan, et al.
Published: (2023)