OmniVL-Guard Pro: A Tool-Augmented Agent for Omnibus Vision-Language Forensics
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shen, Jinjie, Huang, Zheng, Zhang, Yuchen, Wu, Yujiao, Wang, Yaxiong, Cheng, Lechao, Tang, Shengeng, Hui, Tianrui, Pu, Nan, Zhong, Zhun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
OmniVL-Guard: Towards Unified Vision-Language Forgery Detection and Grounding via Balanced RL
von: Shen, Jinjie, et al.
Veröffentlicht: (2026)
von: Shen, Jinjie, et al.
Veröffentlicht: (2026)
Beyond Artificial Misalignment: Detecting and Grounding Semantic-Coordinated Multimodal Manipulations
von: Shen, Jinjie, et al.
Veröffentlicht: (2025)
von: Shen, Jinjie, et al.
Veröffentlicht: (2025)
Towards Fine-Grained Emotion Understanding via Skeleton-Based Micro-Gesture Recognition
von: Xu, Hao, et al.
Veröffentlicht: (2025)
von: Xu, Hao, et al.
Veröffentlicht: (2025)
Knowledge Swapping via Learning and Unlearning
von: Xing, Mingyu, et al.
Veröffentlicht: (2025)
von: Xing, Mingyu, et al.
Veröffentlicht: (2025)
Towards Unified Multimodal Misinformation Detection in Social Media: A Benchmark Dataset and Baseline
von: Li, Haiyang, et al.
Veröffentlicht: (2025)
von: Li, Haiyang, et al.
Veröffentlicht: (2025)
EntityCLIP: Entity-Centric Image-Text Matching via Multimodal Attentive Contrastive Learning
von: Wang, Yaxiong, et al.
Veröffentlicht: (2024)
von: Wang, Yaxiong, et al.
Veröffentlicht: (2024)
FakeSV-VLM: Taming VLM for Detecting Fake Short-Video News via Progressive Mixture-Of-Experts Adapter
von: Wang, Junxi, et al.
Veröffentlicht: (2025)
von: Wang, Junxi, et al.
Veröffentlicht: (2025)
TDEdit: A Unified Diffusion Framework for Text-Drag Guided Image Manipulation
von: Wang, Qihang, et al.
Veröffentlicht: (2025)
von: Wang, Qihang, et al.
Veröffentlicht: (2025)
Towards Micro-Action Recognition with Limited Annotations: An Asynchronous Pseudo Labeling and Training Approach
von: Zhang, Yan, et al.
Veröffentlicht: (2025)
von: Zhang, Yan, et al.
Veröffentlicht: (2025)
CanonSLR: Canonical-View Guided Multi-View Continuous Sign Language Recognition
von: Wang, Xu, et al.
Veröffentlicht: (2026)
von: Wang, Xu, et al.
Veröffentlicht: (2026)
Motion is the Choreographer: Learning Latent Pose Dynamics for Seamless Sign Language Generation
von: He, Jiayi, et al.
Veröffentlicht: (2025)
von: He, Jiayi, et al.
Veröffentlicht: (2025)
Text-Driven Diffusion Model for Sign Language Production
von: He, Jiayi, et al.
Veröffentlicht: (2025)
von: He, Jiayi, et al.
Veröffentlicht: (2025)
Open-World 3D Scene Graph Generation for Retrieval-Augmented Reasoning
von: Yu, Fei, et al.
Veröffentlicht: (2025)
von: Yu, Fei, et al.
Veröffentlicht: (2025)
Shaping a Stabilized Video by Mitigating Unintended Changes for Concept-Augmented Video Editing
von: Guo, Mingce, et al.
Veröffentlicht: (2024)
von: Guo, Mingce, et al.
Veröffentlicht: (2024)
ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
von: Zhang, Zhenxing, et al.
Veröffentlicht: (2024)
von: Zhang, Zhenxing, et al.
Veröffentlicht: (2024)
SSAM: Self-Supervised Association Modeling for Test-Time Adaption
von: Wang, Yaxiong, et al.
Veröffentlicht: (2025)
von: Wang, Yaxiong, et al.
Veröffentlicht: (2025)
FedHPL: Efficient Heterogeneous Federated Learning with Prompt Tuning and Logit Distillation
von: Ma, Yuting, et al.
Veröffentlicht: (2024)
von: Ma, Yuting, et al.
Veröffentlicht: (2024)
Cultivating Forensic Reasoning for Generalizable Multimodal Manipulation Detection
von: Zhang, Yuchen, et al.
Veröffentlicht: (2026)
von: Zhang, Yuchen, et al.
Veröffentlicht: (2026)
Modality Alignment Meets Federated Broadcasting
von: Ma, Yuting, et al.
Veröffentlicht: (2024)
von: Ma, Yuting, et al.
Veröffentlicht: (2024)
Efficient Vision Language Model Fine-tuning for Text-based Person Anomaly Search
von: He, Jiayi, et al.
Veröffentlicht: (2025)
von: He, Jiayi, et al.
Veröffentlicht: (2025)
Discrete to Continuous: Generating Smooth Transition Poses from Sign Language Observation
von: Tang, Shengeng, et al.
Veröffentlicht: (2024)
von: Tang, Shengeng, et al.
Veröffentlicht: (2024)
Dataset Distillers Are Good Label Denoisers In the Wild
von: Cheng, Lechao, et al.
Veröffentlicht: (2024)
von: Cheng, Lechao, et al.
Veröffentlicht: (2024)
SignAligner: Harmonizing Complementary Pose Modalities for Coherent Sign Language Generation
von: Wang, Xu, et al.
Veröffentlicht: (2025)
von: Wang, Xu, et al.
Veröffentlicht: (2025)
StgcDiff: Spatial-Temporal Graph Condition Diffusion for Sign Language Transition Generation
von: He, Jiashu, et al.
Veröffentlicht: (2025)
von: He, Jiashu, et al.
Veröffentlicht: (2025)
Agent0-VL: Exploring Self-Evolving Agent for Tool-Integrated Vision-Language Reasoning
von: Liu, Jiaqi, et al.
Veröffentlicht: (2025)
von: Liu, Jiaqi, et al.
Veröffentlicht: (2025)
OmniGuard: Hybrid Manipulation Localization via Augmented Versatile Deep Image Watermarking
von: Zhang, Xuanyu, et al.
Veröffentlicht: (2024)
von: Zhang, Xuanyu, et al.
Veröffentlicht: (2024)
OmniGuard: Unified Omni-Modal Guardrails with Deliberate Reasoning
von: Zhu, Boyu, et al.
Veröffentlicht: (2025)
von: Zhu, Boyu, et al.
Veröffentlicht: (2025)
Navigating Semantic Drift in Task-Agnostic Class-Incremental Learning
von: Wu, Fangwen, et al.
Veröffentlicht: (2025)
von: Wu, Fangwen, et al.
Veröffentlicht: (2025)
SplitGaussian: Reconstructing Dynamic Scenes via Visual Geometry Decomposition
von: Li, Jiahui, et al.
Veröffentlicht: (2025)
von: Li, Jiahui, et al.
Veröffentlicht: (2025)
Text2Lip: Progressive Lip-Synced Talking Face Generation from Text via Viseme-Guided Rendering
von: Wang, Xu, et al.
Veröffentlicht: (2025)
von: Wang, Xu, et al.
Veröffentlicht: (2025)
CaLa: Complementary Association Learning for Augmenting Composed Image Retrieval
von: Jiang, Xintong, et al.
Veröffentlicht: (2024)
von: Jiang, Xintong, et al.
Veröffentlicht: (2024)
GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning
von: Liu, Yue, et al.
Veröffentlicht: (2025)
von: Liu, Yue, et al.
Veröffentlicht: (2025)
Prototypical Hash Encoding for On-the-Fly Fine-Grained Category Discovery
von: Zheng, Haiyang, et al.
Veröffentlicht: (2024)
von: Zheng, Haiyang, et al.
Veröffentlicht: (2024)
Generalized Fine-Grained Category Discovery with Multi-Granularity Conceptual Experts
von: Zheng, Haiyang, et al.
Veröffentlicht: (2025)
von: Zheng, Haiyang, et al.
Veröffentlicht: (2025)
Multi-Scale Global-Instance Prompt Tuning for Continual Test-time Adaptation in Medical Image Segmentation
von: Li, Lingrui, et al.
Veröffentlicht: (2026)
von: Li, Lingrui, et al.
Veröffentlicht: (2026)
Textual Knowledge Matters: Cross-Modality Co-Teaching for Generalized Visual Class Discovery
von: Zheng, Haiyang, et al.
Veröffentlicht: (2024)
von: Zheng, Haiyang, et al.
Veröffentlicht: (2024)
OmniParser for Pure Vision Based GUI Agent
von: Lu, Yadong, et al.
Veröffentlicht: (2024)
von: Lu, Yadong, et al.
Veröffentlicht: (2024)
The Coherence Trap: When MLLM-Crafted Narratives Exploit Manipulated Visual Contexts
von: Zhang, Yuchen, et al.
Veröffentlicht: (2025)
von: Zhang, Yuchen, et al.
Veröffentlicht: (2025)
VL-Reader: Vision and Language Reconstructor is an Effective Scene Text Recognizer
von: Zhong, Humen, et al.
Veröffentlicht: (2024)
von: Zhong, Humen, et al.
Veröffentlicht: (2024)
Spoofing-aware Prompt Learning for Unified Physical-Digital Facial Attack Detection
von: Guo, Jiabao, et al.
Veröffentlicht: (2025)
von: Guo, Jiabao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
OmniVL-Guard: Towards Unified Vision-Language Forgery Detection and Grounding via Balanced RL
von: Shen, Jinjie, et al.
Veröffentlicht: (2026) -
Beyond Artificial Misalignment: Detecting and Grounding Semantic-Coordinated Multimodal Manipulations
von: Shen, Jinjie, et al.
Veröffentlicht: (2025) -
Towards Fine-Grained Emotion Understanding via Skeleton-Based Micro-Gesture Recognition
von: Xu, Hao, et al.
Veröffentlicht: (2025) -
Knowledge Swapping via Learning and Unlearning
von: Xing, Mingyu, et al.
Veröffentlicht: (2025) -
Towards Unified Multimodal Misinformation Detection in Social Media: A Benchmark Dataset and Baseline
von: Li, Haiyang, et al.
Veröffentlicht: (2025)