The Coherence Trap: When MLLM-Crafted Narratives Exploit Manipulated Visual Contexts
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhang, Yuchen, Wang, Yaxiong, Wu, Yujiao, Wu, Lianwei, Zhu, Li, Zheng, Zhedong |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Cultivating Forensic Reasoning for Generalizable Multimodal Manipulation Detection
por: Zhang, Yuchen, et al.
Publicado: (2026)
por: Zhang, Yuchen, et al.
Publicado: (2026)
Generating Attribution Reports for Manipulated Facial Images: A Dataset and Baseline
por: Lian, Jingchun, et al.
Publicado: (2024)
por: Lian, Jingchun, et al.
Publicado: (2024)
EntityCLIP: Entity-Centric Image-Text Matching via Multimodal Attentive Contrastive Learning
por: Wang, Yaxiong, et al.
Publicado: (2024)
por: Wang, Yaxiong, et al.
Publicado: (2024)
REVEAL: Reference-Grounded Reasoning for Multimodal Manipulation Detection
por: Zhou, Jun, et al.
Publicado: (2026)
por: Zhou, Jun, et al.
Publicado: (2026)
Beyond Walking: A Large-Scale Image-Text Benchmark for Text-based Person Anomaly Search
por: Yang, Shuyu, et al.
Publicado: (2024)
por: Yang, Shuyu, et al.
Publicado: (2024)
Every Painting Awakened: A Training-free Framework for Painting-to-Animation Generation
por: Liu, Lingyu, et al.
Publicado: (2025)
por: Liu, Lingyu, et al.
Publicado: (2025)
Can Video Diffusion Models Predict Past Frames? Bidirectional Cycle Consistency for Reversible Interpolation
por: Liu, Lingyu, et al.
Publicado: (2026)
por: Liu, Lingyu, et al.
Publicado: (2026)
Minimizing the Pretraining Gap: Domain-aligned Text-Based Person Retrieval
por: Yang, Shuyu, et al.
Publicado: (2025)
por: Yang, Shuyu, et al.
Publicado: (2025)
Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark
por: Yang, Shuyu, et al.
Publicado: (2025)
por: Yang, Shuyu, et al.
Publicado: (2025)
Look, Compare and Draw: Differential Query Transformer for Automatic Oil Painting
por: Liu, Lingyu, et al.
Publicado: (2026)
por: Liu, Lingyu, et al.
Publicado: (2026)
Pretrain-then-Adapt: Uncertainty-Aware Test-Time Adaptation for Text-based Person Search
por: Zhang, Jiahao, et al.
Publicado: (2026)
por: Zhang, Jiahao, et al.
Publicado: (2026)
AINet+: Advancing Superpixel Segmentation via Cascaded Association Implantation
por: Wang, Yaxiong, et al.
Publicado: (2021)
por: Wang, Yaxiong, et al.
Publicado: (2021)
Towards Unified Multimodal Misinformation Detection in Social Media: A Benchmark Dataset and Baseline
por: Li, Haiyang, et al.
Publicado: (2025)
por: Li, Haiyang, et al.
Publicado: (2025)
Dual Relation Alignment for Composed Image Retrieval
por: Jiang, Xintong, et al.
Publicado: (2023)
por: Jiang, Xintong, et al.
Publicado: (2023)
Scale Up Composed Image Retrieval Learning via Modification Text Generation
por: Zhou, Yinan, et al.
Publicado: (2025)
por: Zhou, Yinan, et al.
Publicado: (2025)
GUI Action Narrator: Where and When Did That Action Take Place?
por: Wu, Qinchen, et al.
Publicado: (2024)
por: Wu, Qinchen, et al.
Publicado: (2024)
TIGeR: Text-Instructed Generation and Refinement for Template-Free Hand-Object Interaction
por: Huang, Yiyao, et al.
Publicado: (2025)
por: Huang, Yiyao, et al.
Publicado: (2025)
Beyond Known Clusters: Probe New Prototypes for Efficient Generalized Class Discovery
por: Wang, Ye, et al.
Publicado: (2024)
por: Wang, Ye, et al.
Publicado: (2024)
InstructX: Towards Unified Visual Editing with MLLM Guidance
por: Mou, Chong, et al.
Publicado: (2025)
por: Mou, Chong, et al.
Publicado: (2025)
OmniVL-Guard Pro: A Tool-Augmented Agent for Omnibus Vision-Language Forensics
por: Shen, Jinjie, et al.
Publicado: (2026)
por: Shen, Jinjie, et al.
Publicado: (2026)
Beyond Boundary Frames: Context-Centric Video Interpolation with Audio-Visual Semantics
por: Deng, Yuchen, et al.
Publicado: (2025)
por: Deng, Yuchen, et al.
Publicado: (2025)
WeatherPrompt: Multi-modality Representation Learning for All-Weather Drone Visual Geo-Localization
por: Wen, Jiahao, et al.
Publicado: (2025)
por: Wen, Jiahao, et al.
Publicado: (2025)
MEJO: MLLM-Engaged Surgical Triplet Recognition via Inter- and Intra-Task Joint Optimization
por: Zhang, Yiyi, et al.
Publicado: (2025)
por: Zhang, Yiyi, et al.
Publicado: (2025)
Collaborative Group: Composed Image Retrieval via Consensus Learning from Noisy Annotations
por: Zhang, Xu, et al.
Publicado: (2023)
por: Zhang, Xu, et al.
Publicado: (2023)
VL-Uncertainty: Detecting Hallucination in Large Vision-Language Model via Uncertainty Estimation
por: Zhang, Ruiyang, et al.
Publicado: (2024)
por: Zhang, Ruiyang, et al.
Publicado: (2024)
Are MLMs Trapped in the Visual Room?
por: Zhang, Yazhou, et al.
Publicado: (2025)
por: Zhang, Yazhou, et al.
Publicado: (2025)
CaLa: Complementary Association Learning for Augmenting Composed Image Retrieval
por: Jiang, Xintong, et al.
Publicado: (2024)
por: Jiang, Xintong, et al.
Publicado: (2024)
From Instruction to Event: Sound-Triggered Mobile Manipulation
por: Ju, Hao, et al.
Publicado: (2026)
por: Ju, Hao, et al.
Publicado: (2026)
TDEdit: A Unified Diffusion Framework for Text-Drag Guided Image Manipulation
por: Wang, Qihang, et al.
Publicado: (2025)
por: Wang, Qihang, et al.
Publicado: (2025)
SketchThinker-R1: Towards Efficient Sketch-Style Reasoning in Large Multimodal Models
por: Zhang, Ruiyang, et al.
Publicado: (2026)
por: Zhang, Ruiyang, et al.
Publicado: (2026)
AnomalyLMM: Bridging Generative Knowledge and Discriminative Retrieval for Text-Based Person Anomaly Search
por: Ju, Hao, et al.
Publicado: (2025)
por: Ju, Hao, et al.
Publicado: (2025)
AIC MLLM: Autonomous Interactive Correction MLLM for Robust Robotic Manipulation
por: Xiong, Chuyan, et al.
Publicado: (2024)
por: Xiong, Chuyan, et al.
Publicado: (2024)
IPCV: Information-Preserving Compression for MLLM Visual Encoders
por: Chen, Yuan, et al.
Publicado: (2025)
por: Chen, Yuan, et al.
Publicado: (2025)
PIP-MM: Pre-Integrating Prompt Information into Visual Encoding via Existing MLLM Structures
por: Wu, Tianxiang, et al.
Publicado: (2024)
por: Wu, Tianxiang, et al.
Publicado: (2024)
BevSplat: Resolving Height Ambiguity via Feature-Based Gaussian Primitives for Weakly-Supervised Cross-View Localization
por: Wang, Qiwei, et al.
Publicado: (2025)
por: Wang, Qiwei, et al.
Publicado: (2025)
MLLM as Video Narrator: Mitigating Modality Imbalance in Video Moment Retrieval
por: Cai, Weitong, et al.
Publicado: (2024)
por: Cai, Weitong, et al.
Publicado: (2024)
WorldCraft: From Camera Navigation to Object Manipulation in Interactive Video World Models
por: Gu, Bohai, et al.
Publicado: (2026)
por: Gu, Bohai, et al.
Publicado: (2026)
PUMA: Empowering Unified MLLM with Multi-granular Visual Generation
por: Fang, Rongyao, et al.
Publicado: (2024)
por: Fang, Rongyao, et al.
Publicado: (2024)
Creation-MMBench: Assessing Context-Aware Creative Intelligence in MLLM
por: Fang, Xinyu, et al.
Publicado: (2025)
por: Fang, Xinyu, et al.
Publicado: (2025)
Approaching Outside: Scaling Unsupervised 3D Object Detection from 2D Scene
por: Zhang, Ruiyang, et al.
Publicado: (2024)
por: Zhang, Ruiyang, et al.
Publicado: (2024)
Ejemplares similares
-
Cultivating Forensic Reasoning for Generalizable Multimodal Manipulation Detection
por: Zhang, Yuchen, et al.
Publicado: (2026) -
Generating Attribution Reports for Manipulated Facial Images: A Dataset and Baseline
por: Lian, Jingchun, et al.
Publicado: (2024) -
EntityCLIP: Entity-Centric Image-Text Matching via Multimodal Attentive Contrastive Learning
por: Wang, Yaxiong, et al.
Publicado: (2024) -
REVEAL: Reference-Grounded Reasoning for Multimodal Manipulation Detection
por: Zhou, Jun, et al.
Publicado: (2026) -
Beyond Walking: A Large-Scale Image-Text Benchmark for Text-based Person Anomaly Search
por: Yang, Shuyu, et al.
Publicado: (2024)