Doc-PP: Document Policy Preservation Benchmark for Large Vision-Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jang, Haeun, Chang, Hwan, Lee, Hwanhee |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Keep Security! Benchmarking Security Policy Preservation in Large Language Model Contexts Against Indirect Attacks in Question Answering
von: Chang, Hwan, et al.
Veröffentlicht: (2025)
von: Chang, Hwan, et al.
Veröffentlicht: (2025)
Probing-RAG: Self-Probing to Guide Language Models in Selective Document Retrieval
von: Baek, Ingeol, et al.
Veröffentlicht: (2024)
von: Baek, Ingeol, et al.
Veröffentlicht: (2024)
Which Retain Set Matters for LLM Unlearning? A Case Study on Entity Unlearning
von: Chang, Hwan, et al.
Veröffentlicht: (2025)
von: Chang, Hwan, et al.
Veröffentlicht: (2025)
ChatInject: Abusing Chat Templates for Prompt Injection in LLM Agents
von: Chang, Hwan, et al.
Veröffentlicht: (2025)
von: Chang, Hwan, et al.
Veröffentlicht: (2025)
Hallucinate at the Last in Long Response Generation: A Case Study on Long Document Summarization
von: Yang, Joonho, et al.
Veröffentlicht: (2025)
von: Yang, Joonho, et al.
Veröffentlicht: (2025)
Conversational Query Reformulation with the Guidance of Retrieved Documents
von: Park, Jeonghyun, et al.
Veröffentlicht: (2024)
von: Park, Jeonghyun, et al.
Veröffentlicht: (2024)
Low-Resource Cross-Lingual Summarization through Few-Shot Learning with Large Language Models
von: Park, Gyutae, et al.
Veröffentlicht: (2024)
von: Park, Gyutae, et al.
Veröffentlicht: (2024)
How Do Large Vision-Language Models See Text in Image? Unveiling the Distinctive Role of OCR Heads
von: Baek, Ingeol, et al.
Veröffentlicht: (2025)
von: Baek, Ingeol, et al.
Veröffentlicht: (2025)
ContraDoc: Understanding Self-Contradictions in Documents with Large Language Models
von: Li, Jierui, et al.
Veröffentlicht: (2023)
von: Li, Jierui, et al.
Veröffentlicht: (2023)
Investigating Language Preference of Multilingual RAG Systems
von: Park, Jeonghyun, et al.
Veröffentlicht: (2025)
von: Park, Jeonghyun, et al.
Veröffentlicht: (2025)
MARCH: Evaluating the Intersection of Ambiguity Interpretation and Multi-hop Inference
von: Park, Jeonghyun, et al.
Veröffentlicht: (2025)
von: Park, Jeonghyun, et al.
Veröffentlicht: (2025)
Enhancing Building Semantics Preservation in AI Model Training with Large Language Model Encodings
von: Jang, Suhyung, et al.
Veröffentlicht: (2026)
von: Jang, Suhyung, et al.
Veröffentlicht: (2026)
MedLayBench-V: A Large-Scale Benchmark for Expert-Lay Semantic Alignment in Medical Vision Language Models
von: Jang, Han, et al.
Veröffentlicht: (2026)
von: Jang, Han, et al.
Veröffentlicht: (2026)
PP-DocBee: Improving Multimodal Document Understanding Through a Bag of Tricks
von: Ni, Feng, et al.
Veröffentlicht: (2025)
von: Ni, Feng, et al.
Veröffentlicht: (2025)
PP-DocBee2: Improved Baselines with Efficient Data for Multimodal Document Understanding
von: Huang, Kui, et al.
Veröffentlicht: (2025)
von: Huang, Kui, et al.
Veröffentlicht: (2025)
Exploring Persona Sentiment Sensitivity in Personalized Dialogue Generation
von: Jun, Yonghyun, et al.
Veröffentlicht: (2025)
von: Jun, Yonghyun, et al.
Veröffentlicht: (2025)
Dynamic Order Template Prediction for Generative Aspect-Based Sentiment Analysis
von: Jun, Yonghyun, et al.
Veröffentlicht: (2024)
von: Jun, Yonghyun, et al.
Veröffentlicht: (2024)
FIZZ: Factual Inconsistency Detection by Zoom-in Summary and Zoom-out Document
von: Yang, Joonho, et al.
Veröffentlicht: (2024)
von: Yang, Joonho, et al.
Veröffentlicht: (2024)
CUB: Benchmarking Context Utilisation Techniques for Language Models
von: Hagström, Lovisa, et al.
Veröffentlicht: (2025)
von: Hagström, Lovisa, et al.
Veröffentlicht: (2025)
DocGraphLM: Documental Graph Language Model for Information Extraction
von: Wang, Dongsheng, et al.
Veröffentlicht: (2024)
von: Wang, Dongsheng, et al.
Veröffentlicht: (2024)
DocScope: Benchmarking Verifiable Reasoning for Trustworthy Long-Document Understanding
von: Feng, Xiang, et al.
Veröffentlicht: (2026)
von: Feng, Xiang, et al.
Veröffentlicht: (2026)
MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations
von: Ma, Yubo, et al.
Veröffentlicht: (2024)
von: Ma, Yubo, et al.
Veröffentlicht: (2024)
Decomposed On-Policy Distillation for Vision-Language Reasoning: Steering Gradients for Visual Grounding
von: Yoon, Hee Suk, et al.
Veröffentlicht: (2026)
von: Yoon, Hee Suk, et al.
Veröffentlicht: (2026)
Personality Editing for Language Models through Adjusting Self-Referential Queries
von: Hwang, Seojin, et al.
Veröffentlicht: (2025)
von: Hwang, Seojin, et al.
Veröffentlicht: (2025)
Pedagogical Alignment for Vision-Language-Action Models: A Comprehensive Framework for Data, Architecture, and Evaluation in Education
von: Lee, Unggi, et al.
Veröffentlicht: (2026)
von: Lee, Unggi, et al.
Veröffentlicht: (2026)
OpenLearnLM Benchmark: A Unified Framework for Evaluating Knowledge, Skill, and Attitude in Educational Large Language Models
von: Lee, Unggi, et al.
Veröffentlicht: (2026)
von: Lee, Unggi, et al.
Veröffentlicht: (2026)
BenchmarkCards: Standardized Documentation for Large Language Model Benchmarks
von: Sokol, Anna, et al.
Veröffentlicht: (2024)
von: Sokol, Anna, et al.
Veröffentlicht: (2024)
DocReLM: Mastering Document Retrieval with Language Model
von: Wei, Gengchen, et al.
Veröffentlicht: (2024)
von: Wei, Gengchen, et al.
Veröffentlicht: (2024)
Selective Demonstration Retrieval for Improved Implicit Hate Speech Detection
von: Kim, Yumin, et al.
Veröffentlicht: (2025)
von: Kim, Yumin, et al.
Veröffentlicht: (2025)
Are Vision-Language Models Safe in the Wild? A Meme-Based Benchmark Study
von: Lee, DongGeon, et al.
Veröffentlicht: (2025)
von: Lee, DongGeon, et al.
Veröffentlicht: (2025)
WinoQueer: A Community-in-the-Loop Benchmark for Anti-LGBTQ+ Bias in Large Language Models
von: Felkner, Virginia K., et al.
Veröffentlicht: (2023)
von: Felkner, Virginia K., et al.
Veröffentlicht: (2023)
DocEDA: Automated Extraction and Design of Analog Circuits from Documents with Large Language Model
von: Chen, Hong Cai, et al.
Veröffentlicht: (2024)
von: Chen, Hong Cai, et al.
Veröffentlicht: (2024)
Enhancing Multilingual RAG Systems with Debiased Language Preference-Guided Query Fusion
von: Park, Jeonghyun, et al.
Veröffentlicht: (2026)
von: Park, Jeonghyun, et al.
Veröffentlicht: (2026)
AmbigDocs: Reasoning across Documents on Different Entities under the Same Name
von: Lee, Yoonsang, et al.
Veröffentlicht: (2024)
von: Lee, Yoonsang, et al.
Veröffentlicht: (2024)
DocMEdit: Towards Document-Level Model Editing
von: Zeng, Li, et al.
Veröffentlicht: (2025)
von: Zeng, Li, et al.
Veröffentlicht: (2025)
DocMIA: Document-Level Membership Inference Attacks against DocVQA Models
von: Nguyen, Khanh, et al.
Veröffentlicht: (2025)
von: Nguyen, Khanh, et al.
Veröffentlicht: (2025)
KoDialogBench: Evaluating Conversational Understanding of Language Models with Korean Dialogue Benchmark
von: Jang, Seongbo, et al.
Veröffentlicht: (2024)
von: Jang, Seongbo, et al.
Veröffentlicht: (2024)
Benchmarking Direct Preference Optimization for Medical Large Vision-Language Models
von: Kim, Dain, et al.
Veröffentlicht: (2026)
von: Kim, Dain, et al.
Veröffentlicht: (2026)
Leave No Document Behind: Benchmarking Long-Context LLMs with Extended Multi-Doc QA
von: Wang, Minzheng, et al.
Veröffentlicht: (2024)
von: Wang, Minzheng, et al.
Veröffentlicht: (2024)
DocAtlas: Multilingual Document Understanding Across 80+ Languages
von: Heakl, Ahmed, et al.
Veröffentlicht: (2026)
von: Heakl, Ahmed, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Keep Security! Benchmarking Security Policy Preservation in Large Language Model Contexts Against Indirect Attacks in Question Answering
von: Chang, Hwan, et al.
Veröffentlicht: (2025) -
Probing-RAG: Self-Probing to Guide Language Models in Selective Document Retrieval
von: Baek, Ingeol, et al.
Veröffentlicht: (2024) -
Which Retain Set Matters for LLM Unlearning? A Case Study on Entity Unlearning
von: Chang, Hwan, et al.
Veröffentlicht: (2025) -
ChatInject: Abusing Chat Templates for Prompt Injection in LLM Agents
von: Chang, Hwan, et al.
Veröffentlicht: (2025) -
Hallucinate at the Last in Long Response Generation: A Case Study on Long Document Summarization
von: Yang, Joonho, et al.
Veröffentlicht: (2025)