Guardado en:
| Autores principales: | Huang, Wen, Yang, Jiarui, Dai, Tao, Li, Jiawei, Zhan, Shaoxiong, Wang, Bin, Xia, Shu-Tao |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2508.09459 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Looking Back and Forth: Cross-Image Attention Calibration and Attentive Preference Learning for Multi-Image Hallucination Mitigation
por: Yang, Xiaochen, et al.
Publicado: (2026)
por: Yang, Xiaochen, et al.
Publicado: (2026)
GMMFormer v2: An Uncertainty-aware Framework for Partially Relevant Video Retrieval
por: Wang, Yuting, et al.
Publicado: (2024)
por: Wang, Yuting, et al.
Publicado: (2024)
Personalized Face Super-Resolution with Identity Decoupling and Fitting
por: Yang, Jiarui, et al.
Publicado: (2025)
por: Yang, Jiarui, et al.
Publicado: (2025)
Many-for-Many: Unify the Training of Multiple Video and Image Generation and Manipulation Tasks
por: Li, Ruibin, et al.
Publicado: (2025)
por: Li, Ruibin, et al.
Publicado: (2025)
Context-Aware Weakly Supervised Image Manipulation Localization with SAM Refinement
por: Wang, Xinghao, et al.
Publicado: (2025)
por: Wang, Xinghao, et al.
Publicado: (2025)
DuoFormer: Leveraging Hierarchical Visual Representations by Local and Global Attention
por: Tang, Xiaoya, et al.
Publicado: (2024)
por: Tang, Xiaoya, et al.
Publicado: (2024)
CLIP-Guided Generative Networks for Transferable Targeted Adversarial Attacks
por: Fang, Hao, et al.
Publicado: (2024)
por: Fang, Hao, et al.
Publicado: (2024)
DuoFormer: Leveraging Hierarchical Representations by Local and Global Attention Vision Transformer
por: Tang, Xiaoya, et al.
Publicado: (2025)
por: Tang, Xiaoya, et al.
Publicado: (2025)
LCM: Locally Constrained Compact Point Cloud Model for Masked Point Modeling
por: Zha, Yaohua, et al.
Publicado: (2024)
por: Zha, Yaohua, et al.
Publicado: (2024)
Global2Local: A Joint-Hierarchical Attention for Video Captioning
por: Dai, Chengpeng, et al.
Publicado: (2022)
por: Dai, Chengpeng, et al.
Publicado: (2022)
VIA: Unified Spatiotemporal Video Adaptation Framework for Global and Local Video Editing
por: Gu, Jing, et al.
Publicado: (2024)
por: Gu, Jing, et al.
Publicado: (2024)
ManipShield: A Unified Framework for Image Manipulation Detection, Localization and Explanation
por: Xu, Zitong, et al.
Publicado: (2025)
por: Xu, Zitong, et al.
Publicado: (2025)
Training-Free In-Context Forensic Chain for Image Manipulation Detection and Localization
por: Chen, Rui, et al.
Publicado: (2025)
por: Chen, Rui, et al.
Publicado: (2025)
Improving Deepfake Detection with Reinforcement Learning-Based Adaptive Data Augmentation
por: Zhou, Yuxuan, et al.
Publicado: (2025)
por: Zhou, Yuxuan, et al.
Publicado: (2025)
UniVST: A Unified Framework for Training-free Localized Video Style Transfer
por: Song, Quanjian, et al.
Publicado: (2024)
por: Song, Quanjian, et al.
Publicado: (2024)
Omni-IML: Towards Unified Image Manipulation Localization
por: Qu, Chenfan, et al.
Publicado: (2024)
por: Qu, Chenfan, et al.
Publicado: (2024)
PHPQ: Pyramid Hybrid Pooling Quantization for Efficient Fine-Grained Image Retrieval
por: Zeng, Ziyun, et al.
Publicado: (2021)
por: Zeng, Ziyun, et al.
Publicado: (2021)
Parameter Efficient Adaptation for Image Restoration with Heterogeneous Mixture-of-Experts
por: Guo, Hang, et al.
Publicado: (2023)
por: Guo, Hang, et al.
Publicado: (2023)
Protecting Your Video Content: Disrupting Automated Video-based LLM Annotations
por: Liu, Haitong, et al.
Publicado: (2025)
por: Liu, Haitong, et al.
Publicado: (2025)
SpatialFormer: Semantic and Target Aware Attentions for Few-Shot Learning
por: Lai, Jinxiang, et al.
Publicado: (2023)
por: Lai, Jinxiang, et al.
Publicado: (2023)
Suite-IN++: A FlexiWear BodyNet Integrating Global and Local Motion Features from Apple Suite for Robust Inertial Navigation
por: Sun, Lan, et al.
Publicado: (2025)
por: Sun, Lan, et al.
Publicado: (2025)
Unsupervised Deformable Image Registration with Local-Global Attention and Image Decomposition
por: Huang, Zhengyong, et al.
Publicado: (2026)
por: Huang, Zhengyong, et al.
Publicado: (2026)
LoFormer: Local Frequency Transformer for Image Deblurring
por: Mao, Xintian, et al.
Publicado: (2024)
por: Mao, Xintian, et al.
Publicado: (2024)
GLGait: A Global-Local Temporal Receptive Field Network for Gait Recognition in the Wild
por: Peng, Guozhen, et al.
Publicado: (2024)
por: Peng, Guozhen, et al.
Publicado: (2024)
Proto-Former: Unified Facial Landmark Detection by Prototype Transformer
por: Hu, Shengkai, et al.
Publicado: (2025)
por: Hu, Shengkai, et al.
Publicado: (2025)
UGD-IML: A Unified Generative Diffusion-based Framework for Constrained and Unconstrained Image Manipulation Localization
por: Mi, Yachun, et al.
Publicado: (2025)
por: Mi, Yachun, et al.
Publicado: (2025)
Self-supervised Representation Learning with Local Aggregation for Image-based Profiling
por: Dai, Siran, et al.
Publicado: (2025)
por: Dai, Siran, et al.
Publicado: (2025)
Efficiency Follows Global-Local Decoupling
por: Yang, Zhenyu, et al.
Publicado: (2026)
por: Yang, Zhenyu, et al.
Publicado: (2026)
Unifying Global-Local Representations in Salient Object Detection with Transformer
por: Ren, Sucheng, et al.
Publicado: (2021)
por: Ren, Sucheng, et al.
Publicado: (2021)
3D-LMVIC: Learning-based Multi-View Image Coding with 3D Gaussian Geometric Priors
por: Huang, Yujun, et al.
Publicado: (2024)
por: Huang, Yujun, et al.
Publicado: (2024)
MUSE: Manipulating Unified Framework for Synthesizing Emotions in Images via Test-Time Optimization
por: Xia, Yingjie, et al.
Publicado: (2025)
por: Xia, Yingjie, et al.
Publicado: (2025)
BoostAdapter: Improving Vision-Language Test-Time Adaptation via Regional Bootstrapping
por: Zhang, Taolin, et al.
Publicado: (2024)
por: Zhang, Taolin, et al.
Publicado: (2024)
RelayGS: Reconstructing Dynamic Scenes with Large-Scale and Complex Motions via Relay Gaussians
por: Gao, Qiankun, et al.
Publicado: (2024)
por: Gao, Qiankun, et al.
Publicado: (2024)
MambaIR: A Simple Baseline for Image Restoration with State-Space Model
por: Guo, Hang, et al.
Publicado: (2024)
por: Guo, Hang, et al.
Publicado: (2024)
Fractal-IR: A Unified Framework for Efficient and Scalable Image Restoration
por: Li, Yawei, et al.
Publicado: (2025)
por: Li, Yawei, et al.
Publicado: (2025)
Pre-training Point Cloud Compact Model with Partial-aware Reconstruction
por: Zha, Yaohua, et al.
Publicado: (2024)
por: Zha, Yaohua, et al.
Publicado: (2024)
IML-ViT: Benchmarking Image Manipulation Localization by Vision Transformer
por: Ma, Xiaochen, et al.
Publicado: (2023)
por: Ma, Xiaochen, et al.
Publicado: (2023)
UVL2: A Unified Framework for Video Tampering Localization
por: Pei, Pengfei
Publicado: (2023)
por: Pei, Pengfei
Publicado: (2023)
Unified Local and Global Attention Interaction Modeling for Vision Transformers
por: Nguyen, Tan, et al.
Publicado: (2024)
por: Nguyen, Tan, et al.
Publicado: (2024)
Attention to Detail: Global-Local Attention for High-Resolution AI-Generated Image Detection
por: Han, Lawrence
Publicado: (2026)
por: Han, Lawrence
Publicado: (2026)
Ejemplares similares
-
Looking Back and Forth: Cross-Image Attention Calibration and Attentive Preference Learning for Multi-Image Hallucination Mitigation
por: Yang, Xiaochen, et al.
Publicado: (2026) -
GMMFormer v2: An Uncertainty-aware Framework for Partially Relevant Video Retrieval
por: Wang, Yuting, et al.
Publicado: (2024) -
Personalized Face Super-Resolution with Identity Decoupling and Fitting
por: Yang, Jiarui, et al.
Publicado: (2025) -
Many-for-Many: Unify the Training of Multiple Video and Image Generation and Manipulation Tasks
por: Li, Ruibin, et al.
Publicado: (2025) -
Context-Aware Weakly Supervised Image Manipulation Localization with SAM Refinement
por: Wang, Xinghao, et al.
Publicado: (2025)