RelayFormer: A Unified Local-Global Attention Framework for Scalable Image and Video Manipulation Localization
Fuente:
arXiv
Salvato in:
| Autori principali: | Huang, Wen, Yang, Jiarui, Dai, Tao, Li, Jiawei, Zhan, Shaoxiong, Wang, Bin, Xia, Shu-Tao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Looking Back and Forth: Cross-Image Attention Calibration and Attentive Preference Learning for Multi-Image Hallucination Mitigation
di: Yang, Xiaochen, et al.
Pubblicazione: (2026)
di: Yang, Xiaochen, et al.
Pubblicazione: (2026)
GMMFormer v2: An Uncertainty-aware Framework for Partially Relevant Video Retrieval
di: Wang, Yuting, et al.
Pubblicazione: (2024)
di: Wang, Yuting, et al.
Pubblicazione: (2024)
Many-for-Many: Unify the Training of Multiple Video and Image Generation and Manipulation Tasks
di: Li, Ruibin, et al.
Pubblicazione: (2025)
di: Li, Ruibin, et al.
Pubblicazione: (2025)
Personalized Face Super-Resolution with Identity Decoupling and Fitting
di: Yang, Jiarui, et al.
Pubblicazione: (2025)
di: Yang, Jiarui, et al.
Pubblicazione: (2025)
DuoFormer: Leveraging Hierarchical Visual Representations by Local and Global Attention
di: Tang, Xiaoya, et al.
Pubblicazione: (2024)
di: Tang, Xiaoya, et al.
Pubblicazione: (2024)
Context-Aware Weakly Supervised Image Manipulation Localization with SAM Refinement
di: Wang, Xinghao, et al.
Pubblicazione: (2025)
di: Wang, Xinghao, et al.
Pubblicazione: (2025)
DuoFormer: Leveraging Hierarchical Representations by Local and Global Attention Vision Transformer
di: Tang, Xiaoya, et al.
Pubblicazione: (2025)
di: Tang, Xiaoya, et al.
Pubblicazione: (2025)
Global2Local: A Joint-Hierarchical Attention for Video Captioning
di: Dai, Chengpeng, et al.
Pubblicazione: (2022)
di: Dai, Chengpeng, et al.
Pubblicazione: (2022)
ManipShield: A Unified Framework for Image Manipulation Detection, Localization and Explanation
di: Xu, Zitong, et al.
Pubblicazione: (2025)
di: Xu, Zitong, et al.
Pubblicazione: (2025)
VIA: Unified Spatiotemporal Video Adaptation Framework for Global and Local Video Editing
di: Gu, Jing, et al.
Pubblicazione: (2024)
di: Gu, Jing, et al.
Pubblicazione: (2024)
CLIP-Guided Generative Networks for Transferable Targeted Adversarial Attacks
di: Fang, Hao, et al.
Pubblicazione: (2024)
di: Fang, Hao, et al.
Pubblicazione: (2024)
LCM: Locally Constrained Compact Point Cloud Model for Masked Point Modeling
di: Zha, Yaohua, et al.
Pubblicazione: (2024)
di: Zha, Yaohua, et al.
Pubblicazione: (2024)
UniVST: A Unified Framework for Training-free Localized Video Style Transfer
di: Song, Quanjian, et al.
Pubblicazione: (2024)
di: Song, Quanjian, et al.
Pubblicazione: (2024)
Training-Free In-Context Forensic Chain for Image Manipulation Detection and Localization
di: Chen, Rui, et al.
Pubblicazione: (2025)
di: Chen, Rui, et al.
Pubblicazione: (2025)
Omni-IML: Towards Unified Image Manipulation Localization
di: Qu, Chenfan, et al.
Pubblicazione: (2024)
di: Qu, Chenfan, et al.
Pubblicazione: (2024)
Unsupervised Deformable Image Registration with Local-Global Attention and Image Decomposition
di: Huang, Zhengyong, et al.
Pubblicazione: (2026)
di: Huang, Zhengyong, et al.
Pubblicazione: (2026)
LoFormer: Local Frequency Transformer for Image Deblurring
di: Mao, Xintian, et al.
Pubblicazione: (2024)
di: Mao, Xintian, et al.
Pubblicazione: (2024)
UGD-IML: A Unified Generative Diffusion-based Framework for Constrained and Unconstrained Image Manipulation Localization
di: Mi, Yachun, et al.
Pubblicazione: (2025)
di: Mi, Yachun, et al.
Pubblicazione: (2025)
Improving Deepfake Detection with Reinforcement Learning-Based Adaptive Data Augmentation
di: Zhou, Yuxuan, et al.
Pubblicazione: (2025)
di: Zhou, Yuxuan, et al.
Pubblicazione: (2025)
SpatialFormer: Semantic and Target Aware Attentions for Few-Shot Learning
di: Lai, Jinxiang, et al.
Pubblicazione: (2023)
di: Lai, Jinxiang, et al.
Pubblicazione: (2023)
Protecting Your Video Content: Disrupting Automated Video-based LLM Annotations
di: Liu, Haitong, et al.
Pubblicazione: (2025)
di: Liu, Haitong, et al.
Pubblicazione: (2025)
Parameter Efficient Adaptation for Image Restoration with Heterogeneous Mixture-of-Experts
di: Guo, Hang, et al.
Pubblicazione: (2023)
di: Guo, Hang, et al.
Pubblicazione: (2023)
PHPQ: Pyramid Hybrid Pooling Quantization for Efficient Fine-Grained Image Retrieval
di: Zeng, Ziyun, et al.
Pubblicazione: (2021)
di: Zeng, Ziyun, et al.
Pubblicazione: (2021)
Suite-IN++: A FlexiWear BodyNet Integrating Global and Local Motion Features from Apple Suite for Robust Inertial Navigation
di: Sun, Lan, et al.
Pubblicazione: (2025)
di: Sun, Lan, et al.
Pubblicazione: (2025)
Proto-Former: Unified Facial Landmark Detection by Prototype Transformer
di: Hu, Shengkai, et al.
Pubblicazione: (2025)
di: Hu, Shengkai, et al.
Pubblicazione: (2025)
Self-supervised Representation Learning with Local Aggregation for Image-based Profiling
di: Dai, Siran, et al.
Pubblicazione: (2025)
di: Dai, Siran, et al.
Pubblicazione: (2025)
GLGait: A Global-Local Temporal Receptive Field Network for Gait Recognition in the Wild
di: Peng, Guozhen, et al.
Pubblicazione: (2024)
di: Peng, Guozhen, et al.
Pubblicazione: (2024)
Unifying Global-Local Representations in Salient Object Detection with Transformer
di: Ren, Sucheng, et al.
Pubblicazione: (2021)
di: Ren, Sucheng, et al.
Pubblicazione: (2021)
Efficiency Follows Global-Local Decoupling
di: Yang, Zhenyu, et al.
Pubblicazione: (2026)
di: Yang, Zhenyu, et al.
Pubblicazione: (2026)
MUSE: Manipulating Unified Framework for Synthesizing Emotions in Images via Test-Time Optimization
di: Xia, Yingjie, et al.
Pubblicazione: (2025)
di: Xia, Yingjie, et al.
Pubblicazione: (2025)
UVL2: A Unified Framework for Video Tampering Localization
di: Pei, Pengfei
Pubblicazione: (2023)
di: Pei, Pengfei
Pubblicazione: (2023)
Attention to Detail: Global-Local Attention for High-Resolution AI-Generated Image Detection
di: Han, Lawrence
Pubblicazione: (2026)
di: Han, Lawrence
Pubblicazione: (2026)
Fractal-IR: A Unified Framework for Efficient and Scalable Image Restoration
di: Li, Yawei, et al.
Pubblicazione: (2025)
di: Li, Yawei, et al.
Pubblicazione: (2025)
FullTransNet: Full Transformer with Local-Global Attention for Video Summarization
di: Lan, Libin, et al.
Pubblicazione: (2025)
di: Lan, Libin, et al.
Pubblicazione: (2025)
IML-ViT: Benchmarking Image Manipulation Localization by Vision Transformer
di: Ma, Xiaochen, et al.
Pubblicazione: (2023)
di: Ma, Xiaochen, et al.
Pubblicazione: (2023)
HierSum: A Global and Local Attention Mechanism for Video Summarization
di: Beedu, Apoorva, et al.
Pubblicazione: (2025)
di: Beedu, Apoorva, et al.
Pubblicazione: (2025)
Unified Local and Global Attention Interaction Modeling for Vision Transformers
di: Nguyen, Tan, et al.
Pubblicazione: (2024)
di: Nguyen, Tan, et al.
Pubblicazione: (2024)
Multi-spectral Class Center Network for Face Manipulation Detection and Localization
di: Miao, Changtao, et al.
Pubblicazione: (2023)
di: Miao, Changtao, et al.
Pubblicazione: (2023)
GLCF: A Global-Local Multimodal Coherence Analysis Framework for Talking Face Generation Detection
di: Chen, Xiaocan, et al.
Pubblicazione: (2024)
di: Chen, Xiaocan, et al.
Pubblicazione: (2024)
RelayGS: Reconstructing Dynamic Scenes with Large-Scale and Complex Motions via Relay Gaussians
di: Gao, Qiankun, et al.
Pubblicazione: (2024)
di: Gao, Qiankun, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Looking Back and Forth: Cross-Image Attention Calibration and Attentive Preference Learning for Multi-Image Hallucination Mitigation
di: Yang, Xiaochen, et al.
Pubblicazione: (2026) -
GMMFormer v2: An Uncertainty-aware Framework for Partially Relevant Video Retrieval
di: Wang, Yuting, et al.
Pubblicazione: (2024) -
Many-for-Many: Unify the Training of Multiple Video and Image Generation and Manipulation Tasks
di: Li, Ruibin, et al.
Pubblicazione: (2025) -
Personalized Face Super-Resolution with Identity Decoupling and Fitting
di: Yang, Jiarui, et al.
Pubblicazione: (2025) -
DuoFormer: Leveraging Hierarchical Visual Representations by Local and Global Attention
di: Tang, Xiaoya, et al.
Pubblicazione: (2024)