From Masks to Pixels and Meaning: A New Taxonomy, Benchmark, and Metrics for VLM Image Tampering
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shang, Xinyi, Tang, Yi, Cui, Jiacheng, Elhagry, Ahmed, Khatib, Salwa K. Al, Bsharat, Sondos Mahmoud, Liu, Jiacheng, Zhao, Xiaohan, Xue, Jing-Hao, Li, Hao, Khan, Salman, Shen, Zhiqiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Prompting Test-Time Scaling Is A Strong LLM Reasoning Data Augmentation
von: Bsharat, Sondos Mahmoud, et al.
Veröffentlicht: (2025)
von: Bsharat, Sondos Mahmoud, et al.
Veröffentlicht: (2025)
Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena
von: Myrzakhan, Aidar, et al.
Veröffentlicht: (2024)
von: Myrzakhan, Aidar, et al.
Veröffentlicht: (2024)
Principled Instructions Are All You Need for Questioning LLaMA-1/2, GPT-3.5/4
von: Bsharat, Sondos Mahmoud, et al.
Veröffentlicht: (2023)
von: Bsharat, Sondos Mahmoud, et al.
Veröffentlicht: (2023)
Mobile-MMLU: A Mobile Intelligence Language Understanding Benchmark
von: Bsharat, Sondos Mahmoud, et al.
Veröffentlicht: (2025)
von: Bsharat, Sondos Mahmoud, et al.
Veröffentlicht: (2025)
Next-Gen CAPTCHAs: Leveraging the Cognitive Gap for Scalable and Diverse GUI-Agent Defense
von: Liu, Jiacheng, et al.
Veröffentlicht: (2026)
von: Liu, Jiacheng, et al.
Veröffentlicht: (2026)
Dive into Claude Code: The Design Space of Today's and Future AI Agent Systems
von: Liu, Jiacheng, et al.
Veröffentlicht: (2026)
von: Liu, Jiacheng, et al.
Veröffentlicht: (2026)
Exploring 3D Dataset Pruning
von: Zhao, Xiaohan, et al.
Veröffentlicht: (2026)
von: Zhao, Xiaohan, et al.
Veröffentlicht: (2026)
DRAG: Distilling RAG for SLMs from LLMs to Transfer Knowledge and Mitigate Hallucination via Evidence and Graph-based Distillation
von: Chen, Jennifer, et al.
Veröffentlicht: (2025)
von: Chen, Jennifer, et al.
Veröffentlicht: (2025)
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents
von: Luo, Yaxin, et al.
Veröffentlicht: (2025)
von: Luo, Yaxin, et al.
Veröffentlicht: (2025)
LLMSurgeon: Diagnosing Data Mixture of Large Language Models
von: Luo, Yaxin, et al.
Veröffentlicht: (2026)
von: Luo, Yaxin, et al.
Veröffentlicht: (2026)
FADRM: Fast and Accurate Data Residual Matching for Dataset Distillation
von: Cui, Jiacheng, et al.
Veröffentlicht: (2025)
von: Cui, Jiacheng, et al.
Veröffentlicht: (2025)
Hard Labels In! Rethinking the Role of Hard Labels in Mitigating Local Semantic Drift
von: Cui, Jiacheng, et al.
Veröffentlicht: (2025)
von: Cui, Jiacheng, et al.
Veröffentlicht: (2025)
OD3: Optimization-free Dataset Distillation for Object Detection
von: Khatib, Salwa K. Al, et al.
Veröffentlicht: (2025)
von: Khatib, Salwa K. Al, et al.
Veröffentlicht: (2025)
Knowledge Consultation for Semi-Supervised Semantic Segmentation
von: Than, Thuan, et al.
Veröffentlicht: (2025)
von: Than, Thuan, et al.
Veröffentlicht: (2025)
Pushing the Frontier of Black-Box LVLM Attacks via Fine-Grained Detail Targeting
von: Zhao, Xiaohan, et al.
Veröffentlicht: (2026)
von: Zhao, Xiaohan, et al.
Veröffentlicht: (2026)
A Frustratingly Simple Yet Highly Effective Attack Baseline: Over 90% Success Rate Against the Strong Black-box Models of GPT-4.5/4o/o1
von: Li, Zhaoyi, et al.
Veröffentlicht: (2025)
von: Li, Zhaoyi, et al.
Veröffentlicht: (2025)
BiGain: Unified Token Compression for Joint Generation and Classification
von: Liu, Jiacheng, et al.
Veröffentlicht: (2026)
von: Liu, Jiacheng, et al.
Veröffentlicht: (2026)
Novel OCT mosaicking pipeline with Feature- and Pixel-based registration
von: Wang, Jiacheng, et al.
Veröffentlicht: (2023)
von: Wang, Jiacheng, et al.
Veröffentlicht: (2023)
GeoPixel: Pixel Grounding Large Multimodal Model in Remote Sensing
von: Shabbir, Akashah, et al.
Veröffentlicht: (2025)
von: Shabbir, Akashah, et al.
Veröffentlicht: (2025)
The Ergodic Linear-Quadratic Optimal Control Problems for Stochastic Mean-Field Systems with Periodic Coefficients
von: Wu, Jiacheng, et al.
Veröffentlicht: (2025)
von: Wu, Jiacheng, et al.
Veröffentlicht: (2025)
GEOBench-VLM: Benchmarking Vision-Language Models for Geospatial Tasks
von: Danish, Muhammad Sohail, et al.
Veröffentlicht: (2024)
von: Danish, Muhammad Sohail, et al.
Veröffentlicht: (2024)
Bench-MFG: A Benchmark Suite for Learning in Stationary Mean Field Games
von: Magnino, Lorenzo, et al.
Veröffentlicht: (2026)
von: Magnino, Lorenzo, et al.
Veröffentlicht: (2026)
On the Casimir number and formal codegree of Haagerup-Izumi fusion rings
von: Zheng, Ying, et al.
Veröffentlicht: (2025)
von: Zheng, Ying, et al.
Veröffentlicht: (2025)
CEEMDAN-Based Multiscale CNN for Wind Turbine Gearbox Fault Detection
von: Alagha, Nejad, et al.
Veröffentlicht: (2026)
von: Alagha, Nejad, et al.
Veröffentlicht: (2026)
Understanding Implementation Barriers for Lean Magnet Accreditation in the United Arab Emirates: A Qualitative Approach
von: Inas Al Khatib, et al.
Veröffentlicht: (2025)
von: Inas Al Khatib, et al.
Veröffentlicht: (2025)
Federated Self-supervised Domain Generalization for Label-efficient Polyp Segmentation
von: Tan, Xinyi, et al.
Veröffentlicht: (2025)
von: Tan, Xinyi, et al.
Veröffentlicht: (2025)
"Newspaper Eat" Means "Not Tasty": A Taxonomy and Benchmark for Coded Language in Real-World Chinese Online Reviews
von: Wan, Ruyuan, et al.
Veröffentlicht: (2026)
von: Wan, Ruyuan, et al.
Veröffentlicht: (2026)
Change Detection of Markov Kernels with Unknown Pre and Post Change Kernel
von: Chen, Hao, et al.
Veröffentlicht: (2022)
von: Chen, Hao, et al.
Veröffentlicht: (2022)
Tripod in uniform spanning tree and three-sided radial SLE$_2$
von: Ding, Jiacheng, et al.
Veröffentlicht: (2025)
von: Ding, Jiacheng, et al.
Veröffentlicht: (2025)
Reinforcement Learning for Finite Space Mean-Field Type Games
von: Shao, Kai, et al.
Veröffentlicht: (2024)
von: Shao, Kai, et al.
Veröffentlicht: (2024)
BSNet: Box-Supervised Simulation-assisted Mean Teacher for 3D Instance Segmentation
von: Lu, Jiahao, et al.
Veröffentlicht: (2024)
von: Lu, Jiahao, et al.
Veröffentlicht: (2024)
Towards Pixel-Level VLM Perception via Simple Points Prediction
von: Song, Tianhui, et al.
Veröffentlicht: (2026)
von: Song, Tianhui, et al.
Veröffentlicht: (2026)
Can LLMs See Without Pixels? Benchmarking Spatial Intelligence from Textual Descriptions
von: Guo, Zhongbin, et al.
Veröffentlicht: (2026)
von: Guo, Zhongbin, et al.
Veröffentlicht: (2026)
Pixel-wise Smoothing for Certified Robustness against Camera Motion Perturbations
von: Hu, Hanjiang, et al.
Veröffentlicht: (2023)
von: Hu, Hanjiang, et al.
Veröffentlicht: (2023)
Learning Robust Diffusion Models from Imprecise Supervision
von: Wu, Dong-Dong, et al.
Veröffentlicht: (2025)
von: Wu, Dong-Dong, et al.
Veröffentlicht: (2025)
Robust Image Self-Recovery against Tampering using Watermark Generation with Pixel Shuffling
von: Kim, Minyoung, et al.
Veröffentlicht: (2025)
von: Kim, Minyoung, et al.
Veröffentlicht: (2025)
Improving Accuracy-robustness Trade-off via Pixel Reweighted Adversarial Training
von: Zhang, Jiacheng, et al.
Veröffentlicht: (2024)
von: Zhang, Jiacheng, et al.
Veröffentlicht: (2024)
The Judaization of Jerusalem / by Rouhi Al-Khatib
von: Khatib, Rouhi Al
Veröffentlicht: (1973)
von: Khatib, Rouhi Al
Veröffentlicht: (1973)
"One Hole for Benefits" in The Book of Lord Shang: Practical Vindication of an Early Classic of Organizational Management, Clarification of Disputed Passages, and Reconstruction of Its Value
von: Yang, Jiacheng
Veröffentlicht: (2026)
von: Yang, Jiacheng
Veröffentlicht: (2026)
Study on the Service Characteristics of Planetary Roller Screw Mechanisms with Complex Conditions
von: Miao, Jiacheng
Veröffentlicht: (2025)
von: Miao, Jiacheng
Veröffentlicht: (2025)
Ähnliche Einträge
-
Prompting Test-Time Scaling Is A Strong LLM Reasoning Data Augmentation
von: Bsharat, Sondos Mahmoud, et al.
Veröffentlicht: (2025) -
Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena
von: Myrzakhan, Aidar, et al.
Veröffentlicht: (2024) -
Principled Instructions Are All You Need for Questioning LLaMA-1/2, GPT-3.5/4
von: Bsharat, Sondos Mahmoud, et al.
Veröffentlicht: (2023) -
Mobile-MMLU: A Mobile Intelligence Language Understanding Benchmark
von: Bsharat, Sondos Mahmoud, et al.
Veröffentlicht: (2025) -
Next-Gen CAPTCHAs: Leveraging the Cognitive Gap for Scalable and Diverse GUI-Agent Defense
von: Liu, Jiacheng, et al.
Veröffentlicht: (2026)