DocumentCLIP: Linking Figures and Main Body Text in Reflowed Documents
Fuente:
arXiv
Guardado en:
| Autores principales: | Liu, Fuxiao, Tan, Hao, Tensmeyer, Chris |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
ProReflow: Progressive Reflow with Decomposed Velocity
por: Ke, Lei, et al.
Publicado: (2025)
por: Ke, Lei, et al.
Publicado: (2025)
Text-Conditioned Background Generation for Editable Multi-Layer Documents
por: Kang, Taewon, et al.
Publicado: (2025)
por: Kang, Taewon, et al.
Publicado: (2025)
CLIP-based Synergistic Knowledge Transfer for Text-based Person Retrieval
por: Liu, Yating, et al.
Publicado: (2023)
por: Liu, Yating, et al.
Publicado: (2023)
TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
por: Liu, Yuliang, et al.
Publicado: (2024)
por: Liu, Yuliang, et al.
Publicado: (2024)
Parrot Captions Teach CLIP to Spot Text
por: Lin, Yiqi, et al.
Publicado: (2023)
por: Lin, Yiqi, et al.
Publicado: (2023)
LatteCLIP: Unsupervised CLIP Fine-Tuning via LMM-Synthetic Texts
por: Cao, Anh-Quan, et al.
Publicado: (2024)
por: Cao, Anh-Quan, et al.
Publicado: (2024)
Interpreting CLIP's Image Representation via Text-Based Decomposition
por: Gandelsman, Yossi, et al.
Publicado: (2023)
por: Gandelsman, Yossi, et al.
Publicado: (2023)
Enhancing Multimodal Understanding with CLIP-Based Image-to-Text Transformation
por: Che, Chang, et al.
Publicado: (2024)
por: Che, Chang, et al.
Publicado: (2024)
Enhancing Compositional Reasoning in CLIP via Reconstruction and Alignment of Text Descriptions
por: Kwon, Jihoon, et al.
Publicado: (2025)
por: Kwon, Jihoon, et al.
Publicado: (2025)
AdaptCLIP: Adapting CLIP for Universal Visual Anomaly Detection
por: Gao, Bin-Bin, et al.
Publicado: (2025)
por: Gao, Bin-Bin, et al.
Publicado: (2025)
BodyMetric: Evaluating the Realism of Human Bodies in Text-to-Image Generation
por: Andreou, Nefeli, et al.
Publicado: (2024)
por: Andreou, Nefeli, et al.
Publicado: (2024)
DocPedia: Unleashing the Power of Large Multimodal Model in the Frequency Domain for Versatile Document Understanding
por: Feng, Hao, et al.
Publicado: (2023)
por: Feng, Hao, et al.
Publicado: (2023)
Focus, Distinguish, and Prompt: Unleashing CLIP for Efficient and Flexible Scene Text Retrieval
por: Zeng, Gangyan, et al.
Publicado: (2024)
por: Zeng, Gangyan, et al.
Publicado: (2024)
Frame-Difference Guided Dynamic Region Perception for CLIP Adaptation in Text-Video Retrieval
por: Yu, Jiaao, et al.
Publicado: (2025)
por: Yu, Jiaao, et al.
Publicado: (2025)
Does CLIP's Generalization Performance Mainly Stem from High Train-Test Similarity?
por: Mayilvahanan, Prasanna, et al.
Publicado: (2023)
por: Mayilvahanan, Prasanna, et al.
Publicado: (2023)
CLIP-based Camera-Agnostic Feature Learning for Intra-camera Person Re-Identification
por: Tan, Xuan, et al.
Publicado: (2024)
por: Tan, Xuan, et al.
Publicado: (2024)
ComCLIP: Training-Free Compositional Image and Text Matching
por: Jiang, Kenan, et al.
Publicado: (2022)
por: Jiang, Kenan, et al.
Publicado: (2022)
DesignCLIP: Multimodal Learning with CLIP for Design Patent Understanding
por: Wang, Zhu, et al.
Publicado: (2025)
por: Wang, Zhu, et al.
Publicado: (2025)
CLIP-DQA: Blindly Evaluating Dehazed Images from Global and Local Perspectives Using CLIP
por: Zeng, Yirui, et al.
Publicado: (2025)
por: Zeng, Yirui, et al.
Publicado: (2025)
Document Haystacks: Vision-Language Reasoning Over Piles of 1000+ Documents
por: Chen, Jun, et al.
Publicado: (2024)
por: Chen, Jun, et al.
Publicado: (2024)
PriorCLIP: Visual Prior Guided Vision-Language Model for Remote Sensing Image-Text Retrieval
por: Pan, Jiancheng, et al.
Publicado: (2024)
por: Pan, Jiancheng, et al.
Publicado: (2024)
DGTRSD & DGTRS-CLIP: A Dual-Granularity Remote Sensing Image-Text Dataset and Vision Language Foundation Model for Alignment
por: Chen, Weizhi, et al.
Publicado: (2025)
por: Chen, Weizhi, et al.
Publicado: (2025)
M-DocSum: Do LVLMs Genuinely Comprehend Interleaved Image-Text in Document Summarization?
por: Yan, Haolong, et al.
Publicado: (2025)
por: Yan, Haolong, et al.
Publicado: (2025)
VisionCLIP: An Med-AIGC based Ethical Language-Image Foundation Model for Generalizable Retina Image Analysis
por: Wei, Hao, et al.
Publicado: (2024)
por: Wei, Hao, et al.
Publicado: (2024)
TNG-CLIP:Training-Time Negation Data Generation for Negation Awareness of CLIP
por: Cai, Yuliang, et al.
Publicado: (2025)
por: Cai, Yuliang, et al.
Publicado: (2025)
ViCLIP-OT: The First Foundation Vision-Language Model for Vietnamese Image-Text Retrieval with Optimal Transport
por: Tran, Quoc-Khang, et al.
Publicado: (2026)
por: Tran, Quoc-Khang, et al.
Publicado: (2026)
Hierarchical Representation Matching for CLIP-based Class-Incremental Learning
por: Wen, Zhen-Hao, et al.
Publicado: (2025)
por: Wen, Zhen-Hao, et al.
Publicado: (2025)
CLIP-MoE: Towards Building Mixture of Experts for CLIP with Diversified Multiplet Upcycling
por: Zhang, Jihai, et al.
Publicado: (2024)
por: Zhang, Jihai, et al.
Publicado: (2024)
AA-CLIP: Enhancing Zero-shot Anomaly Detection via Anomaly-Aware CLIP
por: Ma, Wenxin, et al.
Publicado: (2025)
por: Ma, Wenxin, et al.
Publicado: (2025)
CLIP-MUSED: CLIP-Guided Multi-Subject Visual Neural Information Semantic Decoding
por: Zhou, Qiongyi, et al.
Publicado: (2024)
por: Zhou, Qiongyi, et al.
Publicado: (2024)
Machine Unlearning for Document Classification
por: Kang, Lei, et al.
Publicado: (2024)
por: Kang, Lei, et al.
Publicado: (2024)
CLIP Architecture for Abdominal CT Image-Text Alignment and Zero-Shot Learning: Investigating Batch Composition and Data Scaling
por: Shivika, et al.
Publicado: (2026)
por: Shivika, et al.
Publicado: (2026)
Text2Avatar: Text to 3D Human Avatar Generation with Codebook-Driven Body Controllable Attribute
por: Gong, Chaoqun, et al.
Publicado: (2024)
por: Gong, Chaoqun, et al.
Publicado: (2024)
FoCLIP: A Feature-Space Misalignment Framework for CLIP-Based Image Manipulation and Detection
por: Chen, Yulin, et al.
Publicado: (2025)
por: Chen, Yulin, et al.
Publicado: (2025)
Fine-tuning CLIP Text Encoders with Two-step Paraphrasing
por: Kim, Hyunjae, et al.
Publicado: (2024)
por: Kim, Hyunjae, et al.
Publicado: (2024)
Omni-NegCLIP: Enhancing CLIP with Front-Layer Contrastive Fine-Tuning for Comprehensive Negation Understanding
por: Xu, Jingqi
Publicado: (2026)
por: Xu, Jingqi
Publicado: (2026)
No Captions, No Problem: Captionless 3D-CLIP Alignment with Hard Negatives via CLIP Knowledge and LLMs
por: Sbrolli, Cristian, et al.
Publicado: (2024)
por: Sbrolli, Cristian, et al.
Publicado: (2024)
Knowledge-Base based Semantic Image Transmission Using CLIP
por: Li, Chongyang, et al.
Publicado: (2025)
por: Li, Chongyang, et al.
Publicado: (2025)
InterCLIP-MEP: Interactive CLIP and Memory-Enhanced Predictor for Multi-modal Sarcasm Detection
por: Chen, Junjie, et al.
Publicado: (2024)
por: Chen, Junjie, et al.
Publicado: (2024)
Prefix-Adaptive Block Diffusion for Efficient Document Recognition
por: Chai, Mingxu, et al.
Publicado: (2026)
por: Chai, Mingxu, et al.
Publicado: (2026)
Ejemplares similares
-
ProReflow: Progressive Reflow with Decomposed Velocity
por: Ke, Lei, et al.
Publicado: (2025) -
Text-Conditioned Background Generation for Editable Multi-Layer Documents
por: Kang, Taewon, et al.
Publicado: (2025) -
CLIP-based Synergistic Knowledge Transfer for Text-based Person Retrieval
por: Liu, Yating, et al.
Publicado: (2023) -
TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
por: Liu, Yuliang, et al.
Publicado: (2024) -
Parrot Captions Teach CLIP to Spot Text
por: Lin, Yiqi, et al.
Publicado: (2023)