Extracting Training Data from Document-Based VQA Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Pinto, Francesco, Rauschmayr, Nathalie, Tramèr, Florian, Torr, Philip, Tombari, Federico |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Seeing Is No Longer Believing: Frontier Image Generation Models, Synthetic Visual Evidence, and Real-World Risk
por: Wu, Shuai, et al.
Publicado: (2026)
por: Wu, Shuai, et al.
Publicado: (2026)
AI vs. Human Moderators: A Comparative Evaluation of Multimodal LLMs in Content Moderation for Brand Safety
por: Levi, Adi, et al.
Publicado: (2025)
por: Levi, Adi, et al.
Publicado: (2025)
A Roadmap for Multilingual, Multimodal Domain Independent Deception Detection
por: Boumber, Dainis, et al.
Publicado: (2024)
por: Boumber, Dainis, et al.
Publicado: (2024)
FlexDoc: Parameterized Sampling for Diverse Multilingual Synthetic Documents for Training Document Understanding Models
por: Dua, Karan, et al.
Publicado: (2025)
por: Dua, Karan, et al.
Publicado: (2025)
Memory-Efficient Differentially Private Training with Gradient Random Projection
por: Mulrooney, Alex, et al.
Publicado: (2025)
por: Mulrooney, Alex, et al.
Publicado: (2025)
Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning
por: Tong, Jingqi, et al.
Publicado: (2025)
por: Tong, Jingqi, et al.
Publicado: (2025)
Exploring and Mitigating Gender Bias in Encoder-Based Transformer Models
por: Hossain, Ariyan, et al.
Publicado: (2025)
por: Hossain, Ariyan, et al.
Publicado: (2025)
Leum-VL Technical Report
por: He, Yuxuan, et al.
Publicado: (2026)
por: He, Yuxuan, et al.
Publicado: (2026)
Reframing linguistic bootstrapping as joint inference using visually-grounded grammar induction models
por: Portelance, Eva, et al.
Publicado: (2024)
por: Portelance, Eva, et al.
Publicado: (2024)
PhysicsArena: The First Multimodal Physics Reasoning Benchmark Exploring Variable, Process, and Solution Dimensions
por: Dai, Song, et al.
Publicado: (2025)
por: Dai, Song, et al.
Publicado: (2025)
Universal Adversarial Attack on Aligned Multimodal LLMs
por: Rahmatullaev, Temurbek, et al.
Publicado: (2025)
por: Rahmatullaev, Temurbek, et al.
Publicado: (2025)
PhysNote: Self-Knowledge Notes for Evolvable Physical Reasoning in Vision-Language Model
por: Zhang, Sinin, et al.
Publicado: (2026)
por: Zhang, Sinin, et al.
Publicado: (2026)
ICG: Improving Cover Image Generation via MLLM-based Prompting and Personalized Preference Alignment
por: Bian, Zhipeng, et al.
Publicado: (2026)
por: Bian, Zhipeng, et al.
Publicado: (2026)
Qwerty AI: Explainable Automated Age Rating and Content Safety Assessment for Russian-Language Screenplays
por: Zmanovskii, Nikita
Publicado: (2025)
por: Zmanovskii, Nikita
Publicado: (2025)
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
por: Raoufi, Behnam, et al.
Publicado: (2025)
por: Raoufi, Behnam, et al.
Publicado: (2025)
OpenMap: Instruction Grounding via Open-Vocabulary Visual-Language Mapping
por: Li, Danyang, et al.
Publicado: (2025)
por: Li, Danyang, et al.
Publicado: (2025)
Taking Flight with Dialogue: Enabling Natural Language Control for PX4-based Drone Agent
por: Lim, Shoon Kit, et al.
Publicado: (2025)
por: Lim, Shoon Kit, et al.
Publicado: (2025)
StratXplore: Strategic Novelty-seeking and Instruction-aligned Exploration for Vision and Language Navigation
por: Gopinathan, Muraleekrishna, et al.
Publicado: (2024)
por: Gopinathan, Muraleekrishna, et al.
Publicado: (2024)
Beyond Localization: A Comprehensive Diagnosis of Perspective-Conditioned Spatial Reasoning in MLLMs from Omnidirectional Images
por: Chen, Yuangong, et al.
Publicado: (2026)
por: Chen, Yuangong, et al.
Publicado: (2026)
MetaCloak-JPEG: JPEG-Robust Adversarial Perturbation for Preventing Unauthorized DreamBooth-Based Deepfake Generation
por: Fardin, Tanjim Rahaman, et al.
Publicado: (2026)
por: Fardin, Tanjim Rahaman, et al.
Publicado: (2026)
Enhancing Sports Strategy with Video Analytics and Data Mining: Assessing the effectiveness of Multimodal LLMs in tennis video analysis
por: Teo, Charlton
Publicado: (2025)
por: Teo, Charlton
Publicado: (2025)
Replicating TEMPEST at Scale: Multi-Turn Adversarial Attacks Against Trillion-Parameter Frontier Models
por: Young, Richard
Publicado: (2025)
por: Young, Richard
Publicado: (2025)
Bye Bye Perspective API: Lessons for Measurement Infrastructure in NLP, CSS and LLM Evaluation
por: Hartmann, David, et al.
Publicado: (2026)
por: Hartmann, David, et al.
Publicado: (2026)
Whose wife is it anyway? Assessing bias against same-gender relationships in machine translation
por: Stewart, Ian, et al.
Publicado: (2024)
por: Stewart, Ian, et al.
Publicado: (2024)
The Company You Keep: How LLMs Respond to Dark Triad Traits
por: Lu, Zeyi, et al.
Publicado: (2026)
por: Lu, Zeyi, et al.
Publicado: (2026)
Tatemae: Detecting Alignment Faking via Tool Selection in LLMs
por: Leonesi, Matteo, et al.
Publicado: (2026)
por: Leonesi, Matteo, et al.
Publicado: (2026)
SALLIE: Safeguarding Against Latent Language & Image Exploits
por: Azov, Guy, et al.
Publicado: (2026)
por: Azov, Guy, et al.
Publicado: (2026)
Learning the meanings of function words from grounded language using a visual question answering model
por: Portelance, Eva, et al.
Publicado: (2023)
por: Portelance, Eva, et al.
Publicado: (2023)
Physics-R1: An Audited Olympiad Corpus and Recipe for Visual Physics Reasoning
por: Yang, Shan
Publicado: (2026)
por: Yang, Shan
Publicado: (2026)
Spatially-Aware Speaker for Vision-and-Language Navigation Instruction Generation
por: Gopinathan, Muraleekrishna, et al.
Publicado: (2024)
por: Gopinathan, Muraleekrishna, et al.
Publicado: (2024)
Evaluating Voice Command Pipelines for Drone Control: From STT and LLM to Direct Classification and Siamese Networks
por: Simões, Lucca Emmanuel Pineli, et al.
Publicado: (2024)
por: Simões, Lucca Emmanuel Pineli, et al.
Publicado: (2024)
LLM-FACETS: A Privacy-Preserving Framework for Evaluating LLM Transparency and Accountability
por: Lucas, Tom, et al.
Publicado: (2026)
por: Lucas, Tom, et al.
Publicado: (2026)
Analyzing Quality, Bias, and Performance in Text-to-Image Generative Models
por: Masrourisaadat, Nila, et al.
Publicado: (2024)
por: Masrourisaadat, Nila, et al.
Publicado: (2024)
Vision Token Masking Alone Cannot Prevent PHI Leakage in Medical Document OCR: A Systematic Evaluation
por: Young, Richard J.
Publicado: (2025)
por: Young, Richard J.
Publicado: (2025)
Correspondence of high-dimensional emotion structures elicited by video clips between humans and Multimodal LLMs
por: Asanuma, Haruka, et al.
Publicado: (2025)
por: Asanuma, Haruka, et al.
Publicado: (2025)
CCVA-FL: Cross-Client Variations Adaptive Federated Learning for Medical Imaging
por: Gupta, Sunny, et al.
Publicado: (2024)
por: Gupta, Sunny, et al.
Publicado: (2024)
Taming the Tail: Leveraging Asymmetric Loss and Pade Approximation to Overcome Medical Image Long-Tailed Class Imbalance
por: Kashyap, Pankhi, et al.
Publicado: (2024)
por: Kashyap, Pankhi, et al.
Publicado: (2024)
MixAT: Combining Continuous and Discrete Adversarial Training for LLMs
por: Dékány, Csaba, et al.
Publicado: (2025)
por: Dékány, Csaba, et al.
Publicado: (2025)
DriveMRP: Enhancing Vision-Language Models with Synthetic Motion Data for Motion Risk Prediction
por: Hou, Zhiyi, et al.
Publicado: (2025)
por: Hou, Zhiyi, et al.
Publicado: (2025)
StoryMovie: A Dataset for Semantic Alignment of Visual Stories with Movie Scripts and Subtitles
por: Oliveira, Daniel, et al.
Publicado: (2026)
por: Oliveira, Daniel, et al.
Publicado: (2026)
Ejemplares similares
-
Seeing Is No Longer Believing: Frontier Image Generation Models, Synthetic Visual Evidence, and Real-World Risk
por: Wu, Shuai, et al.
Publicado: (2026) -
AI vs. Human Moderators: A Comparative Evaluation of Multimodal LLMs in Content Moderation for Brand Safety
por: Levi, Adi, et al.
Publicado: (2025) -
A Roadmap for Multilingual, Multimodal Domain Independent Deception Detection
por: Boumber, Dainis, et al.
Publicado: (2024) -
FlexDoc: Parameterized Sampling for Diverse Multilingual Synthetic Documents for Training Document Understanding Models
por: Dua, Karan, et al.
Publicado: (2025) -
Memory-Efficient Differentially Private Training with Gradient Random Projection
por: Mulrooney, Alex, et al.
Publicado: (2025)