Patent Figure Classification using Large Vision-language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Awale, Sushil, Müller-Budack, Eric, Ewerth, Ralph |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multimodal Misinformation Detection using Large Vision-Language Models
by: Tahmasebi, Sahar, et al.
Published: (2024)
by: Tahmasebi, Sahar, et al.
Published: (2024)
Verifying Cross-modal Entity Consistency in News using Vision-language Models
by: Tahmasebi, Sahar, et al.
Published: (2025)
by: Tahmasebi, Sahar, et al.
Published: (2025)
Multi-Label Plant Species Classification with Self-Supervised Vision Transformers
by: Gustineli, Murilo, et al.
Published: (2024)
by: Gustineli, Murilo, et al.
Published: (2024)
Improving Visual Recommendation on E-commerce Platforms Using Vision-Language Models
by: Yada, Yuki, et al.
Published: (2025)
by: Yada, Yuki, et al.
Published: (2025)
Deep Learning for Technical Document Classification
by: Jiang, Shuo, et al.
Published: (2021)
by: Jiang, Shuo, et al.
Published: (2021)
Transfer Learning and Mixup for Fine-Grained Few-Shot Fungi Classification
by: Tam, Jason Kahei, et al.
Published: (2025)
by: Tam, Jason Kahei, et al.
Published: (2025)
Transfer Learning with Self-Supervised Vision Transformers for Snake Identification
by: Miyaguchi, Anthony, et al.
Published: (2024)
by: Miyaguchi, Anthony, et al.
Published: (2024)
Multi-Label Plant Species Prediction with Metadata-Enhanced Multi-Head Vision Transformers
by: Herasimchyk, Hanna, et al.
Published: (2025)
by: Herasimchyk, Hanna, et al.
Published: (2025)
Large Language Model Informed Patent Image Retrieval
by: Lo, Hao-Cheng, et al.
Published: (2024)
by: Lo, Hao-Cheng, et al.
Published: (2024)
Metric Compatible Training for Online Backfilling in Large-Scale Retrieval
by: Seo, Seonguk, et al.
Published: (2023)
by: Seo, Seonguk, et al.
Published: (2023)
NanoVDR: Distilling a 2B Vision-Language Retriever into a 70M Text-Only Encoder for Visual Document Retrieval
by: Liu, Zhuchenyang, et al.
Published: (2026)
by: Liu, Zhuchenyang, et al.
Published: (2026)
Sync from the Sea: Retrieving Alignable Videos from Large-Scale Datasets
by: Dave, Ishan Rajendrakumar, et al.
Published: (2024)
by: Dave, Ishan Rajendrakumar, et al.
Published: (2024)
Vision Search Assistant: Empower Vision-Language Models as Multimodal Search Engines
by: Zhang, Zhixin, et al.
Published: (2024)
by: Zhang, Zhixin, et al.
Published: (2024)
Image Outlier Detection Without Training using RANSAC
by: Tsai, Chen-Han, et al.
Published: (2023)
by: Tsai, Chen-Han, et al.
Published: (2023)
Towards Resource-Efficient Streaming of Large-Scale Medical Image Datasets for Deep Learning
by: Kulkarni, Pranav, et al.
Published: (2023)
by: Kulkarni, Pranav, et al.
Published: (2023)
GBSK: Skeleton Clustering via Granular-ball Computing and Multi-Sampling for Large-Scale Data
by: Chen, Yewang, et al.
Published: (2025)
by: Chen, Yewang, et al.
Published: (2025)
CTR-Driven Advertising Image Generation with Multimodal Large Language Models
by: Chen, Xingye, et al.
Published: (2025)
by: Chen, Xingye, et al.
Published: (2025)
A Fashion Item Recommendation Model in Hyperbolic Space
by: Shimizu, Ryotaro, et al.
Published: (2024)
by: Shimizu, Ryotaro, et al.
Published: (2024)
InvDiff: Invariant Guidance for Bias Mitigation in Diffusion Models
by: Hou, Min, et al.
Published: (2024)
by: Hou, Min, et al.
Published: (2024)
Image Hashing via Cross-View Code Alignment in the Age of Foundation Models
by: Moummad, Ilyass, et al.
Published: (2025)
by: Moummad, Ilyass, et al.
Published: (2025)
DREAM: Improving Video-Text Retrieval Through Relevance-Based Augmentation Using Large Foundation Models
by: Wang, Yimu, et al.
Published: (2024)
by: Wang, Yimu, et al.
Published: (2024)
EasyEdit: An Easy-to-use Knowledge Editing Framework for Large Language Models
by: Wang, Peng, et al.
Published: (2023)
by: Wang, Peng, et al.
Published: (2023)
AutoPP: Towards Automated Product Poster Generation and Optimization
by: Fan, Jiahao, et al.
Published: (2025)
by: Fan, Jiahao, et al.
Published: (2025)
Tile-Based ViT Inference with Visual-Cluster Priors for Zero-Shot Multi-Species Plant Identification
by: Gustineli, Murilo, et al.
Published: (2025)
by: Gustineli, Murilo, et al.
Published: (2025)
Uni-Layout: Integrating Human Feedback in Unified Layout Generation and Evaluation
by: Lu, Shuo, et al.
Published: (2025)
by: Lu, Shuo, et al.
Published: (2025)
Region-Point Joint Representation for Effective Trajectory Similarity Learning
by: Long, Hao, et al.
Published: (2025)
by: Long, Hao, et al.
Published: (2025)
$\texttt{InfoHier}$: Hierarchical Information Extraction via Encoding and Embedding
by: Zhang, Tianru, et al.
Published: (2025)
by: Zhang, Tianru, et al.
Published: (2025)
LD-DETR: Loop Decoder DEtection TRansformer for Video Moment Retrieval and Highlight Detection
by: Zhao, Pengcheng, et al.
Published: (2025)
by: Zhao, Pengcheng, et al.
Published: (2025)
LotusFilter: Fast Diverse Nearest Neighbor Search via a Learned Cutoff Table
by: Matsui, Yusuke
Published: (2025)
by: Matsui, Yusuke
Published: (2025)
Maximal Matching Matters: Preventing Representation Collapse for Robust Cross-Modal Retrieval
by: Alomari, Hani, et al.
Published: (2025)
by: Alomari, Hani, et al.
Published: (2025)
Studying Illustrations in Manuscripts: An Efficient Deep-Learning Approach
by: Evron, Yoav, et al.
Published: (2025)
by: Evron, Yoav, et al.
Published: (2025)
Embedding-based Retrieval in Multimodal Content Moderation
by: Liang, Hanzhong, et al.
Published: (2025)
by: Liang, Hanzhong, et al.
Published: (2025)
Domain-invariant feature learning in brain MR imaging for content-based image retrieval
by: Tobari, Shuya, et al.
Published: (2025)
by: Tobari, Shuya, et al.
Published: (2025)
Semantic-Cohesive Knowledge Distillation for Deep Cross-modal Hashing
by: Sun, Changchang, et al.
Published: (2025)
by: Sun, Changchang, et al.
Published: (2025)
When & How to Write for Personalized Demand-aware Query Rewriting in Video Search
by: cheng, Cheng, et al.
Published: (2025)
by: cheng, Cheng, et al.
Published: (2025)
VQPP: Video Query Performance Prediction Benchmark
by: Lutu, Adrian Catalin, et al.
Published: (2026)
by: Lutu, Adrian Catalin, et al.
Published: (2026)
SOLAR: SVD-Optimized Lifelong Attention for Recommendation
by: Zhang, Chenghao, et al.
Published: (2026)
by: Zhang, Chenghao, et al.
Published: (2026)
Hierarchy-of-Visual-Words: a Learning-based Approach for Trademark Image Retrieval
by: Lourenço, Vítor N., et al.
Published: (2019)
by: Lourenço, Vítor N., et al.
Published: (2019)
PixRec: Leveraging Visual Context for Next-Item Prediction in Sequential Recommendation
by: Chakrabarty, Sayak, et al.
Published: (2026)
by: Chakrabarty, Sayak, et al.
Published: (2026)
FOR: Finetuning for Object Level Open Vocabulary Image Retrieval
by: Levi, Hila, et al.
Published: (2024)
by: Levi, Hila, et al.
Published: (2024)
Similar Items
-
Multimodal Misinformation Detection using Large Vision-Language Models
by: Tahmasebi, Sahar, et al.
Published: (2024) -
Verifying Cross-modal Entity Consistency in News using Vision-language Models
by: Tahmasebi, Sahar, et al.
Published: (2025) -
Multi-Label Plant Species Classification with Self-Supervised Vision Transformers
by: Gustineli, Murilo, et al.
Published: (2024) -
Improving Visual Recommendation on E-commerce Platforms Using Vision-Language Models
by: Yada, Yuki, et al.
Published: (2025) -
Deep Learning for Technical Document Classification
by: Jiang, Shuo, et al.
Published: (2021)