PatentLMM: Large Multimodal Model for Generating Descriptions for Patent Figures
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shukla, Shreya, Sharma, Nakul, Gupta, Manish, Mishra, Anand |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DesignCLIP: Multimodal Learning with CLIP for Design Patent Understanding
von: Wang, Zhu, et al.
Veröffentlicht: (2025)
von: Wang, Zhu, et al.
Veröffentlicht: (2025)
Sketch-guided Image Inpainting with Partial Discrete Diffusion Process
von: Sharma, Nakul, et al.
Veröffentlicht: (2024)
von: Sharma, Nakul, et al.
Veröffentlicht: (2024)
LMM-VQA: Advancing Video Quality Assessment with Large Multimodal Models
von: Ge, Qihang, et al.
Veröffentlicht: (2024)
von: Ge, Qihang, et al.
Veröffentlicht: (2024)
AviationLMM: A Large Multimodal Foundation Model for Civil Aviation
von: Li, Wenbin, et al.
Veröffentlicht: (2026)
von: Li, Wenbin, et al.
Veröffentlicht: (2026)
SIRR-LMM: Single-image Reflection Removal via Large Multimodal Model
von: Guo, Yu, et al.
Veröffentlicht: (2026)
von: Guo, Yu, et al.
Veröffentlicht: (2026)
LMM-PCQA: Assisting Point Cloud Quality Assessment with LMM
von: Zhang, Zicheng, et al.
Veröffentlicht: (2024)
von: Zhang, Zicheng, et al.
Veröffentlicht: (2024)
MangaVQA and MangaLMM: A Benchmark and Specialized Model for Multimodal Manga Understanding
von: Baek, Jeonghun, et al.
Veröffentlicht: (2025)
von: Baek, Jeonghun, et al.
Veröffentlicht: (2025)
Towards Making Flowchart Images Machine Interpretable
von: Shukla, Shreya, et al.
Veröffentlicht: (2025)
von: Shukla, Shreya, et al.
Veröffentlicht: (2025)
Evaluating Large Vision-language Models for Surgical Tool Detection
von: Poudel, Nakul, et al.
Veröffentlicht: (2026)
von: Poudel, Nakul, et al.
Veröffentlicht: (2026)
Improvise, Adapt, Overcome -- Telescopic Adapters for Efficient Fine-tuning of Vision Language Models in Medical Imaging
von: Mishra, Ujjwal, et al.
Veröffentlicht: (2025)
von: Mishra, Ujjwal, et al.
Veröffentlicht: (2025)
Router-Suggest: Dynamic Routing for Multimodal Auto-Completion in Visually-Grounded Dialogs
von: Mishra, Sandeep, et al.
Veröffentlicht: (2026)
von: Mishra, Sandeep, et al.
Veröffentlicht: (2026)
PatientVLM Meets DocVLM: Pre-Consultation Dialogue Between Vision-Language Models for Efficient Diagnosis
von: Lokesh, K, et al.
Veröffentlicht: (2026)
von: Lokesh, K, et al.
Veröffentlicht: (2026)
Empower Vision Applications with LoRA LMM
von: Mi, Liang, et al.
Veröffentlicht: (2024)
von: Mi, Liang, et al.
Veröffentlicht: (2024)
Visual Text Matters: Improving Text-KVQA with Visual Text Entity Knowledge-aware Large Multimodal Assistant
von: Penamakuri, Abhirama Subramanyam, et al.
Veröffentlicht: (2024)
von: Penamakuri, Abhirama Subramanyam, et al.
Veröffentlicht: (2024)
Text-VQA Aug: Pipelined Harnessing of Large Multimodal Models for Automated Synthesis
von: Joshi, Soham, et al.
Veröffentlicht: (2025)
von: Joshi, Soham, et al.
Veröffentlicht: (2025)
Temporal Object-Aware Vision Transformer for Few-Shot Video Object Detection
von: Kumar, Yogesh, et al.
Veröffentlicht: (2025)
von: Kumar, Yogesh, et al.
Veröffentlicht: (2025)
Composite Sketch+Text Queries for Retrieving Objects with Elusive Names and Complex Interactions
von: Gatti, Prajwal, et al.
Veröffentlicht: (2025)
von: Gatti, Prajwal, et al.
Veröffentlicht: (2025)
LMM-IQA: Image Quality Assessment for Low-Dose CT Imaging
von: Celik, Kagan, et al.
Veröffentlicht: (2025)
von: Celik, Kagan, et al.
Veröffentlicht: (2025)
When Words Can't Capture It All: Towards Video-Based User Complaint Text Generation with Multimodal Video Complaint Dataset
von: Das, Sarmistha, et al.
Veröffentlicht: (2025)
von: Das, Sarmistha, et al.
Veröffentlicht: (2025)
Exploiting LMM-based knowledge for image classification tasks
von: Tzelepi, Maria, et al.
Veröffentlicht: (2024)
von: Tzelepi, Maria, et al.
Veröffentlicht: (2024)
Patent Figure Classification using Large Vision-language Models
von: Awale, Sushil, et al.
Veröffentlicht: (2025)
von: Awale, Sushil, et al.
Veröffentlicht: (2025)
LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
von: Xu, Ruyi, et al.
Veröffentlicht: (2024)
von: Xu, Ruyi, et al.
Veröffentlicht: (2024)
On the Out-Of-Distribution Generalization of Multimodal Large Language Models
von: Zhang, Xingxuan, et al.
Veröffentlicht: (2024)
von: Zhang, Xingxuan, et al.
Veröffentlicht: (2024)
LaMP-Cap: Personalized Figure Caption Generation With Multimodal Figure Profiles
von: Ng, Ho Yin 'Sam', et al.
Veröffentlicht: (2025)
von: Ng, Ho Yin 'Sam', et al.
Veröffentlicht: (2025)
Pixel-Grounded Retrieval for Knowledgeable Large Multimodal Models
von: Kim, Jeonghwan, et al.
Veröffentlicht: (2026)
von: Kim, Jeonghwan, et al.
Veröffentlicht: (2026)
AI-Generated Lecture Slides for Improving Slide Element Detection and Retrieval
von: Maniyar, Suyash, et al.
Veröffentlicht: (2025)
von: Maniyar, Suyash, et al.
Veröffentlicht: (2025)
GRAD-Former: Gated Robust Attention-based Differential Transformer for Change Detection
von: Ameta, Durgesh, et al.
Veröffentlicht: (2026)
von: Ameta, Durgesh, et al.
Veröffentlicht: (2026)
Do Large Multimodal Models Solve Caption Generation for Scientific Figures? Lessons Learned from SciCap Challenge 2023
von: Hsu, Ting-Yao E., et al.
Veröffentlicht: (2025)
von: Hsu, Ting-Yao E., et al.
Veröffentlicht: (2025)
MMR-AD: A Large-Scale Multimodal Dataset for Benchmarking General Anomaly Detection with Multimodal Large Language Models
von: Yao, Xincheng, et al.
Veröffentlicht: (2026)
von: Yao, Xincheng, et al.
Veröffentlicht: (2026)
OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning
von: Srivastava, Siddharth, et al.
Veröffentlicht: (2025)
von: Srivastava, Siddharth, et al.
Veröffentlicht: (2025)
LatteCLIP: Unsupervised CLIP Fine-Tuning via LMM-Synthetic Texts
von: Cao, Anh-Quan, et al.
Veröffentlicht: (2024)
von: Cao, Anh-Quan, et al.
Veröffentlicht: (2024)
Show Me the World in My Language: Establishing the First Baseline for Scene-Text to Scene-Text Translation
von: Vaidya, Shreyas, et al.
Veröffentlicht: (2023)
von: Vaidya, Shreyas, et al.
Veröffentlicht: (2023)
InsightX Agent: An LMM-based Agentic Framework with Integrated Tools for Reliable X-ray NDT Analysis
von: Liu, Jiale, et al.
Veröffentlicht: (2025)
von: Liu, Jiale, et al.
Veröffentlicht: (2025)
How Blind and Low-Vision Individuals Prefer Large Vision-Language Model-Generated Scene Descriptions
von: An, Na Min, et al.
Veröffentlicht: (2025)
von: An, Na Min, et al.
Veröffentlicht: (2025)
F-LMM: Grounding Frozen Large Multimodal Models
von: Wu, Size, et al.
Veröffentlicht: (2024)
von: Wu, Size, et al.
Veröffentlicht: (2024)
Genixer: Empowering Multimodal Large Language Models as a Powerful Data Generator
von: Zhao, Henry Hengyuan, et al.
Veröffentlicht: (2023)
von: Zhao, Henry Hengyuan, et al.
Veröffentlicht: (2023)
PAS : Prelim Attention Score for Detecting Object Hallucinations in Large Vision--Language Models
von: Hoang-Xuan, Nhat, et al.
Veröffentlicht: (2025)
von: Hoang-Xuan, Nhat, et al.
Veröffentlicht: (2025)
Cataract-LMM Large-Scale Multi-Source Multi-Task Benchmark for Deep Learning in Surgical Video Analysis
von: Ahmadi, Mohammad Javad, et al.
Veröffentlicht: (2025)
von: Ahmadi, Mohammad Javad, et al.
Veröffentlicht: (2025)
Augmenting Image Annotation: A Human-LMM Collaborative Framework for Efficient Object Selection and Label Generation
von: Zhang, He, et al.
Veröffentlicht: (2025)
von: Zhang, He, et al.
Veröffentlicht: (2025)
When Big Models Train Small Ones: Label-Free Model Parity Alignment for Efficient Visual Question Answering using Small VLMs
von: Penamakuri, Abhirama Subramanyam, et al.
Veröffentlicht: (2025)
von: Penamakuri, Abhirama Subramanyam, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
DesignCLIP: Multimodal Learning with CLIP for Design Patent Understanding
von: Wang, Zhu, et al.
Veröffentlicht: (2025) -
Sketch-guided Image Inpainting with Partial Discrete Diffusion Process
von: Sharma, Nakul, et al.
Veröffentlicht: (2024) -
LMM-VQA: Advancing Video Quality Assessment with Large Multimodal Models
von: Ge, Qihang, et al.
Veröffentlicht: (2024) -
AviationLMM: A Large Multimodal Foundation Model for Civil Aviation
von: Li, Wenbin, et al.
Veröffentlicht: (2026) -
SIRR-LMM: Single-image Reflection Removal via Large Multimodal Model
von: Guo, Yu, et al.
Veröffentlicht: (2026)