Harnessing PDF Data for Improving Japanese Large Multimodal Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Baek, Jeonghun, Aizawa, Akiko, Aizawa, Kiyoharu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
JMMMU-Pro: Image-based Japanese Multi-discipline Multimodal Understanding Benchmark via Vibe Benchmark Construction
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2025)
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2025)
Manga109-v2026: Revisiting Manga109 Annotations for Modern Manga Understanding
von: Baek, Jeonghun, et al.
Veröffentlicht: (2026)
von: Baek, Jeonghun, et al.
Veröffentlicht: (2026)
JMMMU: A Japanese Massive Multi-discipline Multimodal Understanding Benchmark for Culture-aware Evaluation
von: Onohara, Shota, et al.
Veröffentlicht: (2024)
von: Onohara, Shota, et al.
Veröffentlicht: (2024)
MangaVQA and MangaLMM: A Benchmark and Specialized Model for Multimodal Manga Understanding
von: Baek, Jeonghun, et al.
Veröffentlicht: (2025)
von: Baek, Jeonghun, et al.
Veröffentlicht: (2025)
FoodMLLM-JP: Leveraging Multimodal Large Language Models for Japanese Recipe Generation
von: Imajuku, Yuki, et al.
Veröffentlicht: (2024)
von: Imajuku, Yuki, et al.
Veröffentlicht: (2024)
Jr. AI Scientist and Its Risk Report: Autonomous Scientific Exploration from a Baseline Paper
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2025)
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2025)
Unsolvable Problem Detection: Robust Understanding Evaluation for Large Multimodal Models
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2024)
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2024)
MangaUB: A Manga Understanding Benchmark for Large Multimodal Models
von: Ikuta, Hikaru, et al.
Veröffentlicht: (2024)
von: Ikuta, Hikaru, et al.
Veröffentlicht: (2024)
Manga109Dialog: A Large-scale Dialogue Dataset for Comics Speaker Detection
von: Li, Yingxuan, et al.
Veröffentlicht: (2023)
von: Li, Yingxuan, et al.
Veröffentlicht: (2023)
PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling
von: Xie, Xudong, et al.
Veröffentlicht: (2024)
von: Xie, Xudong, et al.
Veröffentlicht: (2024)
Guided Image Synthesis via Initial Image Editing in Diffusion Model
von: Mao, Jiafeng, et al.
Veröffentlicht: (2023)
von: Mao, Jiafeng, et al.
Veröffentlicht: (2023)
Entity-NeRF: Detecting and Removing Moving Entities in Urban Scenes
von: Otonari, Takashi, et al.
Veröffentlicht: (2024)
von: Otonari, Takashi, et al.
Veröffentlicht: (2024)
The Lottery Ticket Hypothesis in Denoising: Towards Semantic-Driven Initialization
von: Mao, Jiafeng, et al.
Veröffentlicht: (2023)
von: Mao, Jiafeng, et al.
Veröffentlicht: (2023)
Mechanistic Diagnostics of Spatial Lexical Bias in Multimodal Large Language Model Spatial Reasoning
von: Ma, Chuang, et al.
Veröffentlicht: (2026)
von: Ma, Chuang, et al.
Veröffentlicht: (2026)
Zero-Shot Character Identification and Speaker Prediction in Comics via Iterative Multimodal Fusion
von: Li, Yingxuan, et al.
Veröffentlicht: (2024)
von: Li, Yingxuan, et al.
Veröffentlicht: (2024)
Training-Free Sketch-Guided Diffusion with Latent Optimization
von: Ding, Sandra Zhang, et al.
Veröffentlicht: (2024)
von: Ding, Sandra Zhang, et al.
Veröffentlicht: (2024)
A Highly Clean Recipe Dataset with Ingredient States Annotation for State Probing Task
von: Toyooka, Mashiro, et al.
Veröffentlicht: (2025)
von: Toyooka, Mashiro, et al.
Veröffentlicht: (2025)
First Logit Boosting: Visual Grounding Method to Mitigate Object Hallucination in Large Vision-Language Models
von: Ha, Jiwoo, et al.
Veröffentlicht: (2026)
von: Ha, Jiwoo, et al.
Veröffentlicht: (2026)
SUDER: Self-Improving Unified Large Multimodal Models for Understanding and Generation with Dual Self-Rewards
von: Hong, Jixiang, et al.
Veröffentlicht: (2025)
von: Hong, Jixiang, et al.
Veröffentlicht: (2025)
SPD-Faith Bench: Diagnosing and Improving Faithfulness in Chain-of-Thought for Multimodal Large Language Models
von: Lv, Weijiang, et al.
Veröffentlicht: (2026)
von: Lv, Weijiang, et al.
Veröffentlicht: (2026)
Model Composition for Multimodal Large Language Models
von: Chen, Chi, et al.
Veröffentlicht: (2024)
von: Chen, Chi, et al.
Veröffentlicht: (2024)
Graphic Design with Large Multimodal Model
von: Cheng, Yutao, et al.
Veröffentlicht: (2024)
von: Cheng, Yutao, et al.
Veröffentlicht: (2024)
DIAMOND: An LLM-Driven Agent for Context-Aware Baseball Highlight Summarization
von: Kang, Jeonghun, et al.
Veröffentlicht: (2025)
von: Kang, Jeonghun, et al.
Veröffentlicht: (2025)
PerFace: Metric Learning in Perceptual Facial Similarity for Enhanced Face Anonymization
von: Kumagai, Haruka, et al.
Veröffentlicht: (2025)
von: Kumagai, Haruka, et al.
Veröffentlicht: (2025)
GL-MCM: Global and Local Maximum Concept Matching for Zero-Shot Out-of-Distribution Detection
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2023)
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2023)
PP-DocBee2: Improved Baselines with Efficient Data for Multimodal Document Understanding
von: Huang, Kui, et al.
Veröffentlicht: (2025)
von: Huang, Kui, et al.
Veröffentlicht: (2025)
Benchmarking Multimodal Large Language Models for Face Recognition
von: Shahreza, Hatef Otroshi, et al.
Veröffentlicht: (2025)
von: Shahreza, Hatef Otroshi, et al.
Veröffentlicht: (2025)
A Survey on Agentic Multimodal Large Language Models
von: Yao, Huanjin, et al.
Veröffentlicht: (2025)
von: Yao, Huanjin, et al.
Veröffentlicht: (2025)
A Survey on Benchmarks of Multimodal Large Language Models
von: Li, Jian, et al.
Veröffentlicht: (2024)
von: Li, Jian, et al.
Veröffentlicht: (2024)
A Survey on Evaluation of Multimodal Large Language Models
von: Huang, Jiaxing, et al.
Veröffentlicht: (2024)
von: Huang, Jiaxing, et al.
Veröffentlicht: (2024)
mmE5: Improving Multimodal Multilingual Embeddings via High-quality Synthetic Data
von: Chen, Haonan, et al.
Veröffentlicht: (2025)
von: Chen, Haonan, et al.
Veröffentlicht: (2025)
Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation
von: Ye, Junyan, et al.
Veröffentlicht: (2025)
von: Ye, Junyan, et al.
Veröffentlicht: (2025)
Robust Multimodal Large Language Models Against Modality Conflict
von: Zhang, Zongmeng, et al.
Veröffentlicht: (2025)
von: Zhang, Zongmeng, et al.
Veröffentlicht: (2025)
MLLM-CL: Continual Learning for Multimodal Large Language Models
von: Zhao, Hongbo, et al.
Veröffentlicht: (2025)
von: Zhao, Hongbo, et al.
Veröffentlicht: (2025)
Forgotten Polygons: Multimodal Large Language Models are Shape-Blind
von: Rudman, William, et al.
Veröffentlicht: (2025)
von: Rudman, William, et al.
Veröffentlicht: (2025)
Evaluating Large Language Models on Multimodal Chemistry Olympiad Exams
von: Cui, Yiming, et al.
Veröffentlicht: (2025)
von: Cui, Yiming, et al.
Veröffentlicht: (2025)
Cross-modal Information Flow in Multimodal Large Language Models
von: Zhang, Zhi, et al.
Veröffentlicht: (2024)
von: Zhang, Zhi, et al.
Veröffentlicht: (2024)
Gemini in Reasoning: Unveiling Commonsense in Multimodal Large Language Models
von: Wang, Yuqing, et al.
Veröffentlicht: (2023)
von: Wang, Yuqing, et al.
Veröffentlicht: (2023)
BLINK: Multimodal Large Language Models Can See but Not Perceive
von: Fu, Xingyu, et al.
Veröffentlicht: (2024)
von: Fu, Xingyu, et al.
Veröffentlicht: (2024)
E$^2$AT: Multimodal Jailbreak Defense via Dynamic Joint Optimization for Multimodal Large Language Models
von: Lu, Liming, et al.
Veröffentlicht: (2025)
von: Lu, Liming, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
JMMMU-Pro: Image-based Japanese Multi-discipline Multimodal Understanding Benchmark via Vibe Benchmark Construction
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2025) -
Manga109-v2026: Revisiting Manga109 Annotations for Modern Manga Understanding
von: Baek, Jeonghun, et al.
Veröffentlicht: (2026) -
JMMMU: A Japanese Massive Multi-discipline Multimodal Understanding Benchmark for Culture-aware Evaluation
von: Onohara, Shota, et al.
Veröffentlicht: (2024) -
MangaVQA and MangaLMM: A Benchmark and Specialized Model for Multimodal Manga Understanding
von: Baek, Jeonghun, et al.
Veröffentlicht: (2025) -
FoodMLLM-JP: Leveraging Multimodal Large Language Models for Japanese Recipe Generation
von: Imajuku, Yuki, et al.
Veröffentlicht: (2024)