Saved in:
| Main Authors: | Cai, Peng, Li, Qiang, Yang, Kaicheng, Guo, Dong, Li, Jia, Zhou, Nan, An, Xiang, Yang, Ninghua, Deng, Jiankang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2507.19804 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multi-label Cluster Discrimination for Visual Representation Learning
by: An, Xiang, et al.
Published: (2024)
by: An, Xiang, et al.
Published: (2024)
RWKV-CLIP: A Robust Vision-Language Representation Learner
by: Gu, Tiancheng, et al.
Published: (2024)
by: Gu, Tiancheng, et al.
Published: (2024)
RealSyn: An Effective and Scalable Multimodal Interleaved Document Transformation Paradigm
by: Gu, Tiancheng, et al.
Published: (2025)
by: Gu, Tiancheng, et al.
Published: (2025)
BookNet: Book Image Rectification via Cross-Page Attention Network
by: Liu, Shaokai, et al.
Published: (2026)
by: Liu, Shaokai, et al.
Published: (2026)
CLIP-CID: Efficient CLIP Distillation via Cluster-Instance Discrimination
by: Yang, Kaicheng, et al.
Published: (2024)
by: Yang, Kaicheng, et al.
Published: (2024)
High-Fidelity Facial Albedo Estimation via Texture Quantization
by: Ran, Zimin, et al.
Published: (2024)
by: Ran, Zimin, et al.
Published: (2024)
PaCo-FR: Patch-Pixel Aligned End-to-End Codebook Learning for Facial Representation Pre-training
by: Xie, Yin, et al.
Published: (2025)
by: Xie, Yin, et al.
Published: (2025)
UniME-V2: MLLM-as-a-Judge for Universal Multimodal Embedding Learning
by: Gu, Tiancheng, et al.
Published: (2025)
by: Gu, Tiancheng, et al.
Published: (2025)
TextCenGen: Attention-Guided Text-Centric Background Adaptation for Text-to-Image Generation
by: Liang, Tianyi, et al.
Published: (2024)
by: Liang, Tianyi, et al.
Published: (2024)
G3DR: Generative 3D Reconstruction in ImageNet
by: Reddy, Pradyumna, et al.
Published: (2024)
by: Reddy, Pradyumna, et al.
Published: (2024)
VAR-CLIP: Text-to-Image Generator with Visual Auto-Regressive Modeling
by: Zhang, Qian, et al.
Published: (2024)
by: Zhang, Qian, et al.
Published: (2024)
Document Image Rectification Bases on Self-Adaptive Multitask Fusion
by: Li, Heng, et al.
Published: (2025)
by: Li, Heng, et al.
Published: (2025)
IDAdapter: Learning Mixed Features for Tuning-Free Personalization of Text-to-Image Models
by: Cui, Siying, et al.
Published: (2024)
by: Cui, Siying, et al.
Published: (2024)
Breaking the Modality Barrier: Universal Embedding Learning with Multimodal LLMs
by: Gu, Tiancheng, et al.
Published: (2025)
by: Gu, Tiancheng, et al.
Published: (2025)
Dual-frequency Selected Knowledge Distillation with Statistical-based Sample Rectification for PolSAR Image Classification
by: Xin, Xinyue, et al.
Published: (2025)
by: Xin, Xinyue, et al.
Published: (2025)
Cascaded Robust Rectification for Arbitrary Document Images
by: Wang, Chaoyun, et al.
Published: (2025)
by: Wang, Chaoyun, et al.
Published: (2025)
Background Fades, Foreground Leads: Curriculum-Guided Background Pruning for Efficient Foreground-Centric Collaborative Perception
by: Wu, Yuheng, et al.
Published: (2025)
by: Wu, Yuheng, et al.
Published: (2025)
ViCToR: Improving Visual Comprehension via Token Reconstruction for Pretraining LMMs
by: Xie, Yin, et al.
Published: (2024)
by: Xie, Yin, et al.
Published: (2024)
Beyond the Safety Tax: Mitigating Unsafe Text-to-Image Generation via External Safety Rectification
by: Meng, Xiangtao, et al.
Published: (2025)
by: Meng, Xiangtao, et al.
Published: (2025)
Region-based Cluster Discrimination for Visual Representation Learning
by: Xie, Yin, et al.
Published: (2025)
by: Xie, Yin, et al.
Published: (2025)
LMAD: Integrated End-to-End Vision-Language Model for Explainable Autonomous Driving
by: Song, Nan, et al.
Published: (2025)
by: Song, Nan, et al.
Published: (2025)
RigNet++: Semantic Assisted Repetitive Image Guided Network for Depth Completion
by: Yan, Zhiqiang, et al.
Published: (2023)
by: Yan, Zhiqiang, et al.
Published: (2023)
QueryCDR: Query-Based Controllable Distortion Rectification Network for Fisheye Images
by: Guo, Pengbo, et al.
Published: (2024)
by: Guo, Pengbo, et al.
Published: (2024)
ShelfRectNet: Single View Shelf Image Rectification with Homography Estimation
by: Tore, Onur Berk, et al.
Published: (2025)
by: Tore, Onur Berk, et al.
Published: (2025)
Future-Aware End-to-End Driving: Bidirectional Modeling of Trajectory Planning and Scene Evolution
by: Zhang, Bozhou, et al.
Published: (2025)
by: Zhang, Bozhou, et al.
Published: (2025)
ORID: Organ-Regional Information Driven Framework for Radiology Report Generation
by: Gu, Tiancheng, et al.
Published: (2024)
by: Gu, Tiancheng, et al.
Published: (2024)
Relation Rectification in Diffusion Model
by: Wu, Yinwei, et al.
Published: (2024)
by: Wu, Yinwei, et al.
Published: (2024)
See Tomorrow, Act Today: Foresight-Driven Autonomous Driving
by: Zhang, Bozhou, et al.
Published: (2026)
by: Zhang, Bozhou, et al.
Published: (2026)
NeuroPump: Simultaneous Geometric and Color Rectification for Underwater Images
by: Guo, Yue, et al.
Published: (2024)
by: Guo, Yue, et al.
Published: (2024)
RoFIR: Robust Fisheye Image Rectification Framework Impervious to Optical Center Deviation
by: Liao, Zhaokang, et al.
Published: (2024)
by: Liao, Zhaokang, et al.
Published: (2024)
Disentangling Foreground and Background for vision-Language Navigation via Online Augmentation
by: Xu, Yunbo, et al.
Published: (2025)
by: Xu, Yunbo, et al.
Published: (2025)
QArtSR: Quantization via Reverse-Module and Timestep-Retraining in One-Step Diffusion based Image Super-Resolution
by: Zhu, Libo, et al.
Published: (2025)
by: Zhu, Libo, et al.
Published: (2025)
Enhancing Few-Shot Out-of-Distribution Detection via the Refinement of Foreground and Background
by: Li, Tianyu, et al.
Published: (2026)
by: Li, Tianyu, et al.
Published: (2026)
Attribution-Guided Model Rectification of Unreliable Neural Network Behaviors
by: Yang, Peiyu, et al.
Published: (2026)
by: Yang, Peiyu, et al.
Published: (2026)
CSPR-Net: Self-supervised Curved Surface Projection Rectification Network for Geometric Distortion Correction in Non-planar Projections
by: Peng, Kejin, et al.
Published: (2026)
by: Peng, Kejin, et al.
Published: (2026)
LaPA: Latent Prompt Assist Model For Medical Visual Question Answering
by: Gu, Tiancheng, et al.
Published: (2024)
by: Gu, Tiancheng, et al.
Published: (2024)
A Deep Single Image Rectification Approach for Pan-Tilt-Zoom Cameras
by: Xiao, Teng, et al.
Published: (2025)
by: Xiao, Teng, et al.
Published: (2025)
LMS-Net: A Learned Mumford-Shah Network For Few-Shot Medical Image Segmentation
by: Zhang, Shengdong, et al.
Published: (2025)
by: Zhang, Shengdong, et al.
Published: (2025)
Generating Non-Stationary Textures using Self-Rectification
by: Zhou, Yang, et al.
Published: (2024)
by: Zhou, Yang, et al.
Published: (2024)
Spatio-temporal Prompting Network for Robust Video Feature Extraction
by: Sun, Guanxiong, et al.
Published: (2024)
by: Sun, Guanxiong, et al.
Published: (2024)
Similar Items
-
Multi-label Cluster Discrimination for Visual Representation Learning
by: An, Xiang, et al.
Published: (2024) -
RWKV-CLIP: A Robust Vision-Language Representation Learner
by: Gu, Tiancheng, et al.
Published: (2024) -
RealSyn: An Effective and Scalable Multimodal Interleaved Document Transformation Paradigm
by: Gu, Tiancheng, et al.
Published: (2025) -
BookNet: Book Image Rectification via Cross-Page Attention Network
by: Liu, Shaokai, et al.
Published: (2026) -
CLIP-CID: Efficient CLIP Distillation via Cluster-Instance Discrimination
by: Yang, Kaicheng, et al.
Published: (2024)