Unleashing the Potential of Consistency Learning for Detecting and Grounding Multi-Modal Media Manipulation
Fuente:
arXiv
Guardado en:
| Autores principales: | Li, Yiheng, Yang, Yang, Tan, Zichang, Liu, Huan, Chen, Weihua, Zhou, Xu, Lei, Zhen |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Towards Generalizable AI-Generated Image Detection via Image-Adaptive Prompt Learning
por: Li, Yiheng, et al.
Publicado: (2025)
por: Li, Yiheng, et al.
Publicado: (2025)
CTForensics: A Comprehensive Dataset and Method for AI-Generated CT Image Detection
por: Li, Yiheng, et al.
Publicado: (2026)
por: Li, Yiheng, et al.
Publicado: (2026)
Reduce the Artifacts Bias for More Generalizable AI-Generated Image Detection
por: Li, Yiheng, et al.
Publicado: (2026)
por: Li, Yiheng, et al.
Publicado: (2026)
PVLR: Prompt-driven Visual-Linguistic Representation Learning for Multi-Label Image Recognition
por: Tan, Hao, et al.
Publicado: (2024)
por: Tan, Hao, et al.
Publicado: (2024)
Exploiting Modality-Specific Features For Multi-Modal Manipulation Detection And Grounding
por: Wang, Jiazhen, et al.
Publicado: (2023)
por: Wang, Jiazhen, et al.
Publicado: (2023)
CoreNet: Conflict Resolution Network for Point-Pixel Misalignment and Sub-Task Suppression of 3D LiDAR-Camera Object Detection
por: Li, Yiheng, et al.
Publicado: (2025)
por: Li, Yiheng, et al.
Publicado: (2025)
Recover and Match: Open-Vocabulary Multi-Label Recognition through Knowledge-Constrained Optimal Transport
por: Tan, Hao, et al.
Publicado: (2025)
por: Tan, Hao, et al.
Publicado: (2025)
RCTrans: Radar-Camera Transformer via Radar Densifier and Sequential Decoder for 3D Object Detection
por: Li, Yiheng, et al.
Publicado: (2024)
por: Li, Yiheng, et al.
Publicado: (2024)
Unleashing the Potential of Multimodal LLMs for Zero-Shot Spatio-Temporal Video Grounding
por: Yang, Zaiquan, et al.
Publicado: (2025)
por: Yang, Zaiquan, et al.
Publicado: (2025)
DECO: Unleashing the Potential of ConvNets for Query-based Detection and Segmentation
por: Chen, Xinghao, et al.
Publicado: (2023)
por: Chen, Xinghao, et al.
Publicado: (2023)
GLOVER++: Unleashing the Potential of Affordance Learning from Human Behaviors for Robotic Manipulation
por: Ma, Teli, et al.
Publicado: (2025)
por: Ma, Teli, et al.
Publicado: (2025)
SSPA: Split-and-Synthesize Prompting with Gated Alignments for Multi-Label Image Recognition
por: Tan, Hao, et al.
Publicado: (2024)
por: Tan, Hao, et al.
Publicado: (2024)
Unleashing VLA Potentials in Autonomous Driving via Explicit Learning from Failures
por: Luo, Yuechen, et al.
Publicado: (2026)
por: Luo, Yuechen, et al.
Publicado: (2026)
MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing
por: Zheng, Junjie, et al.
Publicado: (2025)
por: Zheng, Junjie, et al.
Publicado: (2025)
VideoVeritas: AI-Generated Video Detection via Perception Pretext Reinforcement Learning
por: Tan, Hao, et al.
Publicado: (2026)
por: Tan, Hao, et al.
Publicado: (2026)
NCL++: Nested Collaborative Learning for Long-Tailed Visual Recognition
por: Tan, Zichang, et al.
Publicado: (2023)
por: Tan, Zichang, et al.
Publicado: (2023)
Unleashing the Potential of the Diffusion Model in Few-shot Semantic Segmentation
por: Zhu, Muzhi, et al.
Publicado: (2024)
por: Zhu, Muzhi, et al.
Publicado: (2024)
DAG: Unleash the Potential of Diffusion Model for Open-Vocabulary 3D Affordance Grounding
por: Wang, Hanqing, et al.
Publicado: (2025)
por: Wang, Hanqing, et al.
Publicado: (2025)
ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
por: Zhang, Zhenxing, et al.
Publicado: (2024)
por: Zhang, Zhenxing, et al.
Publicado: (2024)
Quaternion Sparse Decomposition for Multi-focus Color Image Fusion
por: Yang, Weihua, et al.
Publicado: (2025)
por: Yang, Weihua, et al.
Publicado: (2025)
LatentUM: Unleashing the Potential of Interleaved Cross-Modal Reasoning via a Latent-Space Unified Model
por: Jin, Jiachun, et al.
Publicado: (2026)
por: Jin, Jiachun, et al.
Publicado: (2026)
Exploring Conditional Multi-Modal Prompts for Zero-shot HOI Detection
por: Lei, Ting, et al.
Publicado: (2024)
por: Lei, Ting, et al.
Publicado: (2024)
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding
por: Yin, Xingyilang, et al.
Publicado: (2025)
por: Yin, Xingyilang, et al.
Publicado: (2025)
RMMSS: Towards Advanced Robust Multi-Modal Semantic Segmentation with Hybrid Prototype Distillation and Feature Selection
por: Tan, Jiaqi, et al.
Publicado: (2025)
por: Tan, Jiaqi, et al.
Publicado: (2025)
Unleashing In-context Learning of Autoregressive Models for Few-shot Image Manipulation
por: Lai, Bolin, et al.
Publicado: (2024)
por: Lai, Bolin, et al.
Publicado: (2024)
Teaching with Uncertainty: Unleashing the Potential of Knowledge Distillation in Object Detection
por: Yi, Junfei, et al.
Publicado: (2024)
por: Yi, Junfei, et al.
Publicado: (2024)
How Auxiliary Reasoning Unleashes GUI Grounding in VLMs
por: Li, Weiming, et al.
Publicado: (2025)
por: Li, Weiming, et al.
Publicado: (2025)
CoVAR: Co-generation of Video and Action for Robotic Manipulation via Multi-Modal Diffusion
por: Yang, Liudi, et al.
Publicado: (2025)
por: Yang, Liudi, et al.
Publicado: (2025)
REVEAL: Reference-Grounded Reasoning for Multimodal Manipulation Detection
por: Zhou, Jun, et al.
Publicado: (2026)
por: Zhou, Jun, et al.
Publicado: (2026)
Instruction Guided Multi Object Image Editing with Quantity and Layout Consistency
por: Tan, Jiaqi, et al.
Publicado: (2025)
por: Tan, Jiaqi, et al.
Publicado: (2025)
Modality-Agnostic Prompt Learning for Multi-Modal Camouflaged Object Detection
por: Wang, Hao, et al.
Publicado: (2026)
por: Wang, Hao, et al.
Publicado: (2026)
Multi-clue Consistency Learning to Bridge Gaps Between General and Oriented Object in Semi-supervised Detection
por: Wang, Chenxu, et al.
Publicado: (2024)
por: Wang, Chenxu, et al.
Publicado: (2024)
VoCoT: Unleashing Visually Grounded Multi-Step Reasoning in Large Multi-Modal Models
por: Li, Zejun, et al.
Publicado: (2024)
por: Li, Zejun, et al.
Publicado: (2024)
Unleashing the Capabilities of Large Vision-Language Models for Intelligent Perception of Roadside Infrastructure
por: Fu, Luxuan, et al.
Publicado: (2026)
por: Fu, Luxuan, et al.
Publicado: (2026)
RealisHuman: A Two-Stage Approach for Refining Malformed Human Parts in Generated Images
por: Wang, Benzhi, et al.
Publicado: (2024)
por: Wang, Benzhi, et al.
Publicado: (2024)
DGFusion: Dual-guided Fusion for Robust Multi-Modal 3D Object Detection
por: Jia, Feiyang, et al.
Publicado: (2025)
por: Jia, Feiyang, et al.
Publicado: (2025)
Source-Free Cross-Modal Knowledge Transfer by Unleashing the Potential of Task-Irrelevant Data
por: Zhu, Jinjing, et al.
Publicado: (2024)
por: Zhu, Jinjing, et al.
Publicado: (2024)
SEEC: Segmentation-Assisted Multi-Entropy Models for Learned Lossless Image Compression
por: Zheng, Chunhang, et al.
Publicado: (2025)
por: Zheng, Chunhang, et al.
Publicado: (2025)
GraphBEV: Towards Robust BEV Feature Alignment for Multi-Modal 3D Object Detection
por: Song, Ziying, et al.
Publicado: (2024)
por: Song, Ziying, et al.
Publicado: (2024)
On Learning Multi-Modal Forgery Representation for Diffusion Generated Video Detection
por: Song, Xiufeng, et al.
Publicado: (2024)
por: Song, Xiufeng, et al.
Publicado: (2024)
Ejemplares similares
-
Towards Generalizable AI-Generated Image Detection via Image-Adaptive Prompt Learning
por: Li, Yiheng, et al.
Publicado: (2025) -
CTForensics: A Comprehensive Dataset and Method for AI-Generated CT Image Detection
por: Li, Yiheng, et al.
Publicado: (2026) -
Reduce the Artifacts Bias for More Generalizable AI-Generated Image Detection
por: Li, Yiheng, et al.
Publicado: (2026) -
PVLR: Prompt-driven Visual-Linguistic Representation Learning for Multi-Label Image Recognition
por: Tan, Hao, et al.
Publicado: (2024) -
Exploiting Modality-Specific Features For Multi-Modal Manipulation Detection And Grounding
por: Wang, Jiazhen, et al.
Publicado: (2023)