Saved in:
| Main Authors: | Hu, Xiaoxing, Yang, Kaicheng, Wang, Jun, Xu, Haoran, Feng, Ziyong, Wang, Yupei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2504.16801 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ProCLIP: Progressive Vision-Language Alignment via LLM-based Embedder
by: Hu, Xiaoxing, et al.
Published: (2025)
by: Hu, Xiaoxing, et al.
Published: (2025)
Multi-label Cluster Discrimination for Visual Representation Learning
by: An, Xiang, et al.
Published: (2024)
by: An, Xiang, et al.
Published: (2024)
1st Place Solution to the 1st SkatingVerse Challenge
by: Sun, Tao, et al.
Published: (2024)
by: Sun, Tao, et al.
Published: (2024)
Gradient-Attention Guided Dual-Masking Synergetic Framework for Robust Text-based Person Retrieval
by: Zheng, Tianlu, et al.
Published: (2025)
by: Zheng, Tianlu, et al.
Published: (2025)
ORID: Organ-Regional Information Driven Framework for Radiology Report Generation
by: Gu, Tiancheng, et al.
Published: (2024)
by: Gu, Tiancheng, et al.
Published: (2024)
ViCToR: Improving Visual Comprehension via Token Reconstruction for Pretraining LMMs
by: Xie, Yin, et al.
Published: (2024)
by: Xie, Yin, et al.
Published: (2024)
RWKV-CLIP: A Robust Vision-Language Representation Learner
by: Gu, Tiancheng, et al.
Published: (2024)
by: Gu, Tiancheng, et al.
Published: (2024)
Efficient, Validation-Free Intrinsic Quality Estimation for Large-Scale Face Recognition Datasets
by: Chen, Zhichao, et al.
Published: (2026)
by: Chen, Zhichao, et al.
Published: (2026)
UniDoc-RL: Coarse-to-Fine Visual RAG with Hierarchical Actions and Dense Rewards
by: Wang, Jun, et al.
Published: (2026)
by: Wang, Jun, et al.
Published: (2026)
CLIP-CID: Efficient CLIP Distillation via Cluster-Instance Discrimination
by: Yang, Kaicheng, et al.
Published: (2024)
by: Yang, Kaicheng, et al.
Published: (2024)
Breaking the Modality Barrier: Universal Embedding Learning with Multimodal LLMs
by: Gu, Tiancheng, et al.
Published: (2025)
by: Gu, Tiancheng, et al.
Published: (2025)
DATR: Unsupervised Domain Adaptive Detection Transformer with Dataset-Level Adaptation and Prototypical Alignment
by: Han, Jianhong, et al.
Published: (2024)
by: Han, Jianhong, et al.
Published: (2024)
Efficiency Follows Global-Local Decoupling
by: Yang, Zhenyu, et al.
Published: (2026)
by: Yang, Zhenyu, et al.
Published: (2026)
RealSyn: An Effective and Scalable Multimodal Interleaved Document Transformation Paradigm
by: Gu, Tiancheng, et al.
Published: (2025)
by: Gu, Tiancheng, et al.
Published: (2025)
High-Fidelity Facial Albedo Estimation via Texture Quantization
by: Ran, Zimin, et al.
Published: (2024)
by: Ran, Zimin, et al.
Published: (2024)
Weakly-supervised Localization of Manipulated Image Regions Using Multi-resolution Learned Features
by: Wang, Ziyong, et al.
Published: (2025)
by: Wang, Ziyong, et al.
Published: (2025)
PaCo-FR: Patch-Pixel Aligned End-to-End Codebook Learning for Facial Representation Pre-training
by: Xie, Yin, et al.
Published: (2025)
by: Xie, Yin, et al.
Published: (2025)
Temporal Action Localization with Cross Layer Task Decoupling and Refinement
by: Li, Qiang, et al.
Published: (2024)
by: Li, Qiang, et al.
Published: (2024)
UniME-V2: MLLM-as-a-Judge for Universal Multimodal Embedding Learning
by: Gu, Tiancheng, et al.
Published: (2025)
by: Gu, Tiancheng, et al.
Published: (2025)
Region-based Cluster Discrimination for Visual Representation Learning
by: Xie, Yin, et al.
Published: (2025)
by: Xie, Yin, et al.
Published: (2025)
Earth-Adapter: Bridge the Geospatial Domain Gaps with Mixture of Frequency Adaptation
by: Hu, Xiaoxing, et al.
Published: (2025)
by: Hu, Xiaoxing, et al.
Published: (2025)
Local2Global query Alignment for Video Instance Segmentation
by: Koner, Rajat, et al.
Published: (2025)
by: Koner, Rajat, et al.
Published: (2025)
LocalDyGS: Multi-view Global Dynamic Scene Modeling via Adaptive Local Implicit Feature Decoupling
by: Wu, Jiahao, et al.
Published: (2025)
by: Wu, Jiahao, et al.
Published: (2025)
DanQing: An Up-to-Date Large-Scale Chinese Vision-Language Pre-training Dataset
by: Shen, Hengyu, et al.
Published: (2026)
by: Shen, Hengyu, et al.
Published: (2026)
IUP-Pose: Decoupled Iterative Uncertainty Propagation for Real-time Relative Pose Regression via Implicit Dense Alignment v1
by: Wang, Jun, et al.
Published: (2026)
by: Wang, Jun, et al.
Published: (2026)
AQUA-SLAM: Tightly-Coupled Underwater Acoustic-Visual-Inertial SLAM with Sensor Calibration
by: Xu, Shida, et al.
Published: (2025)
by: Xu, Shida, et al.
Published: (2025)
Improving Adversarial Robustness via Decoupled Visual Representation Masking
by: Liu, Decheng, et al.
Published: (2024)
by: Liu, Decheng, et al.
Published: (2024)
VFM-Guided Semi-Supervised Detection Transformer under Source-Free Constraints for Remote Sensing Object Detection
by: Han, Jianhong, et al.
Published: (2025)
by: Han, Jianhong, et al.
Published: (2025)
Style-Adaptive Detection Transformer for Single-Source Domain Generalized Object Detection
by: Han, Jianhong, et al.
Published: (2025)
by: Han, Jianhong, et al.
Published: (2025)
Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models
by: Xiao, Junyuan, et al.
Published: (2026)
by: Xiao, Junyuan, et al.
Published: (2026)
On the Global Photometric Alignment for Low-Level Vision
by: Li, Mingjia, et al.
Published: (2026)
by: Li, Mingjia, et al.
Published: (2026)
EvoTok: A Unified Image Tokenizer via Residual Latent Evolution for Visual Understanding and Generation
by: Li, Yan, et al.
Published: (2026)
by: Li, Yan, et al.
Published: (2026)
DeRA: Decoupled Representation Alignment for Video Tokenization
by: Guo, Pengbo, et al.
Published: (2025)
by: Guo, Pengbo, et al.
Published: (2025)
DVLO: Deep Visual-LiDAR Odometry with Local-to-Global Feature Fusion and Bi-Directional Structure Alignment
by: Liu, Jiuming, et al.
Published: (2024)
by: Liu, Jiuming, et al.
Published: (2024)
Multi-Level Embedding and Alignment Network with Consistency and Invariance Learning for Cross-View Geo-Localization
by: Chen, Zhongwei, et al.
Published: (2024)
by: Chen, Zhongwei, et al.
Published: (2024)
VAR-CLIP: Text-to-Image Generator with Visual Auto-Regressive Modeling
by: Zhang, Qian, et al.
Published: (2024)
by: Zhang, Qian, et al.
Published: (2024)
Locality Alignment Improves Vision-Language Models
by: Covert, Ian, et al.
Published: (2024)
by: Covert, Ian, et al.
Published: (2024)
Mocap-2-to-3: Multi-view Lifting for Monocular Motion Recovery with 2D Pretraining
by: Wang, Zhumei, et al.
Published: (2025)
by: Wang, Zhumei, et al.
Published: (2025)
DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies
by: Song, Wei, et al.
Published: (2025)
by: Song, Wei, et al.
Published: (2025)
Flowing Backwards: Improving Normalizing Flows via Reverse Representation Alignment
by: Chen, Yang, et al.
Published: (2025)
by: Chen, Yang, et al.
Published: (2025)
Similar Items
-
ProCLIP: Progressive Vision-Language Alignment via LLM-based Embedder
by: Hu, Xiaoxing, et al.
Published: (2025) -
Multi-label Cluster Discrimination for Visual Representation Learning
by: An, Xiang, et al.
Published: (2024) -
1st Place Solution to the 1st SkatingVerse Challenge
by: Sun, Tao, et al.
Published: (2024) -
Gradient-Attention Guided Dual-Masking Synergetic Framework for Robust Text-based Person Retrieval
by: Zheng, Tianlu, et al.
Published: (2025) -
ORID: Organ-Regional Information Driven Framework for Radiology Report Generation
by: Gu, Tiancheng, et al.
Published: (2024)