PROVE: A Perceptual RemOVal cohErence Benchmark for Visual Media
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Fuhao, You, Shaofeng, Hu, Jiagao, Liu, Yu, Chen, Yuxuan, Wang, Zepeng, Wang, Fei, Zhou, Daiguo, Luan, Jian |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AutoAWG: Adverse Weather Generation with Adaptive Multi-Controls for Automotive Videos
by: Hu, Jiagao, et al.
Published: (2026)
by: Hu, Jiagao, et al.
Published: (2026)
From Ideal to Real: Stable Video Object Removal under Imperfect Conditions
by: Hu, Jiagao, et al.
Published: (2026)
by: Hu, Jiagao, et al.
Published: (2026)
RemEdit: Efficient Diffusion Editing with Riemannian Geometry
by: Adhikarla, Eashan, et al.
Published: (2026)
by: Adhikarla, Eashan, et al.
Published: (2026)
Controllable Pedestrian Video Editing for Multi-View Driving Scenarios via Motion Sequence
by: Fu, Danzhen, et al.
Published: (2025)
by: Fu, Danzhen, et al.
Published: (2025)
Perceptual-oriented Learned Image Compression with Dynamic Kernel
by: Fu, Nianxiang, et al.
Published: (2024)
by: Fu, Nianxiang, et al.
Published: (2024)
Perceptual Quality Assessment of Octree-RAHT Encoded 3D Point Clouds
by: Duan, Dongshuai, et al.
Published: (2024)
by: Duan, Dongshuai, et al.
Published: (2024)
Harmony-Aware Music-driven Motion Synthesis with Perceptual Constraint on UGC Datasets
by: Wu, Xinyi, et al.
Published: (2025)
by: Wu, Xinyi, et al.
Published: (2025)
PP-Motion: Physical-Perceptual Fidelity Evaluation for Human Motion Generation
by: Zhao, Sihan, et al.
Published: (2025)
by: Zhao, Sihan, et al.
Published: (2025)
Perceptual Visual Quality Assessment: Principles, Methods, and Future Directions
by: Zhou, Wei, et al.
Published: (2025)
by: Zhou, Wei, et al.
Published: (2025)
Perceptual Depth Quality Assessment of Stereoscopic Omnidirectional Images
by: Zhou, Wei, et al.
Published: (2024)
by: Zhou, Wei, et al.
Published: (2024)
Perceptual Crack Detection for Rendered 3D Textured Meshes
by: Sarvestani, Armin Shafiee, et al.
Published: (2024)
by: Sarvestani, Armin Shafiee, et al.
Published: (2024)
SpeechEE: A Novel Benchmark for Speech Event Extraction
by: Wang, Bin, et al.
Published: (2024)
by: Wang, Bin, et al.
Published: (2024)
MMED: A Multimodal Micro-Expression Dataset based on Audio-Visual Fusion
by: Wang, Junbo, et al.
Published: (2025)
by: Wang, Junbo, et al.
Published: (2025)
SFQA: A Comprehensive Perceptual Quality Assessment Dataset for Singing Face Generation
by: Gao, Zhilin, et al.
Published: (2026)
by: Gao, Zhilin, et al.
Published: (2026)
AVID: A Benchmark for Omni-Modal Audio-Visual Inconsistency Understanding via Agent-Driven Construction
by: Chen, Zixuan, et al.
Published: (2026)
by: Chen, Zixuan, et al.
Published: (2026)
AV-Edit: Multimodal Generative Sound Effect Editing via Audio-Visual Semantic Joint Control
by: Guo, Xinyue, et al.
Published: (2025)
by: Guo, Xinyue, et al.
Published: (2025)
TeMTG: Text-Enhanced Multi-Hop Temporal Graph Modeling for Audio-Visual Video Parsing
by: Chen, Yaru, et al.
Published: (2025)
by: Chen, Yaru, et al.
Published: (2025)
TPIFM: A Task-Aware Model for Evaluating Perceptual Interaction Fluency in Remote AR Collaboration
by: Song, Jiarun, et al.
Published: (2026)
by: Song, Jiarun, et al.
Published: (2026)
VSpeechLM: A Visual Speech Language Model for Visual Text-to-Speech Task
by: Wang, Yuyue, et al.
Published: (2025)
by: Wang, Yuyue, et al.
Published: (2025)
FLARE: Full-Modality Long-Video Audiovisual Retrieval Benchmark with User-Simulated Queries
by: You, Qijie, et al.
Published: (2026)
by: You, Qijie, et al.
Published: (2026)
Do LLMs Understand Visual Anomalies? Uncovering LLM's Capabilities in Zero-shot Anomaly Detection
by: Zhu, Jiaqi, et al.
Published: (2024)
by: Zhu, Jiaqi, et al.
Published: (2024)
PopSim: Social Network Simulation for Social Media Popularity Prediction
by: Liu, Yijun, et al.
Published: (2025)
by: Liu, Yijun, et al.
Published: (2025)
Revisiting Vision-Language Features Adaptation and Inconsistency for Social Media Popularity Prediction
by: Hsu, Chih-Chung, et al.
Published: (2024)
by: Hsu, Chih-Chung, et al.
Published: (2024)
Perceptual Quality Optimization of Image Super-Resolution
by: Zhou, Wei, et al.
Published: (2026)
by: Zhou, Wei, et al.
Published: (2026)
Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark
by: Zhang, Han, et al.
Published: (2025)
by: Zhang, Han, et al.
Published: (2025)
Semantic Communication-Enabled Cloud-Edge-End-collaborative Metaverse Services Architecure
by: Li, Yuxuan, et al.
Published: (2025)
by: Li, Yuxuan, et al.
Published: (2025)
Content-Adaptive Rate-Quality Curve Prediction Model in Media Processing System
by: Yin, Shibo, et al.
Published: (2024)
by: Yin, Shibo, et al.
Published: (2024)
fMRI Exploration of Visual Quality Assessment
by: Zhang, Yiming, et al.
Published: (2024)
by: Zhang, Yiming, et al.
Published: (2024)
Contextual Wireless Video Semantic Communication in MIMO-OFDM Systems
by: Xie, Bingyan, et al.
Published: (2026)
by: Xie, Bingyan, et al.
Published: (2026)
RealBench: A Chinese Multi-image Understanding Benchmark Close to Real-world Scenarios
by: Zhao, Fei, et al.
Published: (2025)
by: Zhao, Fei, et al.
Published: (2025)
Recognizing Everything from All Modalities at Once: Grounded Multimodal Universal Information Extraction
by: Zhang, Meishan, et al.
Published: (2024)
by: Zhang, Meishan, et al.
Published: (2024)
Pistachio: Towards Synthetic, Balanced, and Long-Form Video Anomaly Benchmarks
by: Li, Jie, et al.
Published: (2025)
by: Li, Jie, et al.
Published: (2025)
EidetiCom: A Cross-modal Brain-Computer Semantic Communication Paradigm for Decoding Visual Perception
by: Zheng, Linfeng, et al.
Published: (2024)
by: Zheng, Linfeng, et al.
Published: (2024)
MViR: Multi-View Visual-Semantic Representation for Fake News Detection
by: Liang, Haochen, et al.
Published: (2026)
by: Liang, Haochen, et al.
Published: (2026)
SMTPD: A New Benchmark for Temporal Prediction of Social Media Popularity
by: Xu, Yijie, et al.
Published: (2025)
by: Xu, Yijie, et al.
Published: (2025)
Cross-Modality and Within-Modality Regularization for Audio-Visual DeepFake Detection
by: Zou, Heqing, et al.
Published: (2024)
by: Zou, Heqing, et al.
Published: (2024)
Orthogonal Disentanglement with Projected Feature Alignment for Multimodal Emotion Recognition in Conversation
by: Che, Xinyi, et al.
Published: (2025)
by: Che, Xinyi, et al.
Published: (2025)
Beyond Text: Multimodal Jailbreaking of Vision-Language and Audio Models through Perceptually Simple Transformations
by: Kumar, Divyanshu, et al.
Published: (2025)
by: Kumar, Divyanshu, et al.
Published: (2025)
SFE-Net: Harnessing Biological Principles of Differential Gene Expression for Improved Feature Selection in Deep Learning Networks
by: Li, Yuqi, et al.
Published: (2024)
by: Li, Yuqi, et al.
Published: (2024)
SocialDF: Benchmark Dataset and Detection Model for Mitigating Harmful Deepfake Content on Social Media Platforms
by: Batra, Arnesh, et al.
Published: (2025)
by: Batra, Arnesh, et al.
Published: (2025)
Similar Items
-
AutoAWG: Adverse Weather Generation with Adaptive Multi-Controls for Automotive Videos
by: Hu, Jiagao, et al.
Published: (2026) -
From Ideal to Real: Stable Video Object Removal under Imperfect Conditions
by: Hu, Jiagao, et al.
Published: (2026) -
RemEdit: Efficient Diffusion Editing with Riemannian Geometry
by: Adhikarla, Eashan, et al.
Published: (2026) -
Controllable Pedestrian Video Editing for Multi-View Driving Scenarios via Motion Sequence
by: Fu, Danzhen, et al.
Published: (2025) -
Perceptual-oriented Learned Image Compression with Dynamic Kernel
by: Fu, Nianxiang, et al.
Published: (2024)