Toward Real Text Manipulation Detection: New Dataset and New Solution
Fuente:
arXiv
Saved in:
| Main Authors: | Luo, Dongliang, Liu, Yuliang, Yang, Rui, Liu, Xianjin, Zeng, Jishen, Zhou, Yu, Bai, Xiang |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The First Swahili Language Scene Text Detection and Recognition Dataset
by: Douamba, Fadila Wendigoundi, et al.
Published: (2024)
by: Douamba, Fadila Wendigoundi, et al.
Published: (2024)
SemiETS: Integrating Spatial and Content Consistencies for Semi-Supervised End-to-end Text Spotting
by: Luo, Dongliang, et al.
Published: (2025)
by: Luo, Dongliang, et al.
Published: (2025)
Dataset and Benchmark for Urdu Natural Scenes Text Detection, Recognition and Visual Question Answering
by: Maryam, Hiba, et al.
Published: (2024)
by: Maryam, Hiba, et al.
Published: (2024)
Towards UAV Detection in the Real World: A New Multispectral Dataset UAVNet-MS and a New Method
by: Luo, Yihang, et al.
Published: (2026)
by: Luo, Yihang, et al.
Published: (2026)
DocThinker: Explainable Multimodal Large Language Models with Rule-based Reinforcement Learning for Document Understanding
by: Yu, Wenwen, et al.
Published: (2025)
by: Yu, Wenwen, et al.
Published: (2025)
SwinTextSpotter v2: Towards Better Synergy for Scene Text Spotting
by: Huang, Mingxin, et al.
Published: (2024)
by: Huang, Mingxin, et al.
Published: (2024)
Bridging the Gap Between End-to-End and Two-Step Text Spotting
by: Huang, Mingxin, et al.
Published: (2024)
by: Huang, Mingxin, et al.
Published: (2024)
MSTAR: Box-free Multi-query Scene Text Retrieval with Attention Recycling
by: Yin, Liang, et al.
Published: (2025)
by: Yin, Liang, et al.
Published: (2025)
Towards Automatic Power Battery Detection: New Challenge, Benchmark Dataset and Baseline
by: Zhao, Xiaoqi, et al.
Published: (2023)
by: Zhao, Xiaoqi, et al.
Published: (2023)
SIGMA: Semantic-Difference Instruction-Grounding Mask Annotator for Text-Driven Image Manipulation Localization
by: Zhuang, Peiyu, et al.
Published: (2026)
by: Zhuang, Peiyu, et al.
Published: (2026)
A New Dataset and Framework for Real-World Blurred Images Super-Resolution
by: Qin, Rui, et al.
Published: (2024)
by: Qin, Rui, et al.
Published: (2024)
First Creating Backgrounds Then Rendering Texts: A New Paradigm for Visual Text Blending
by: Li, Zhenhang, et al.
Published: (2024)
by: Li, Zhenhang, et al.
Published: (2024)
Super-resolving Real-world Image Illumination Enhancement: A New Dataset and A Conditional Diffusion Model
by: Liu, Yang, et al.
Published: (2024)
by: Liu, Yang, et al.
Published: (2024)
Progressive Evolution from Single-Point to Polygon for Scene Text
by: Deng, Linger, et al.
Published: (2023)
by: Deng, Linger, et al.
Published: (2023)
The Devil is in Fine-tuning and Long-tailed Problems:A New Benchmark for Scene Text Detection
by: Cao, Tianjiao, et al.
Published: (2025)
by: Cao, Tianjiao, et al.
Published: (2025)
TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
by: Liu, Yuliang, et al.
Published: (2024)
by: Liu, Yuliang, et al.
Published: (2024)
TokBench: Evaluating Your Visual Tokenizer before Visual Generation
by: Wu, Junfeng, et al.
Published: (2025)
by: Wu, Junfeng, et al.
Published: (2025)
DeltaEdit: Exploring Text-free Training for Text-Driven Image Manipulation
by: Lyu, Yueming, et al.
Published: (2023)
by: Lyu, Yueming, et al.
Published: (2023)
OmniParser: A Unified Framework for Text Spotting, Key Information Extraction and Table Recognition
by: Wan, Jianqiang, et al.
Published: (2024)
by: Wan, Jianqiang, et al.
Published: (2024)
TextSleuth: Towards Explainable Tampered Text Detection
by: Qu, Chenfan, et al.
Published: (2024)
by: Qu, Chenfan, et al.
Published: (2024)
Towards Generalized and Training-Free Text-Guided Semantic Manipulation
by: Hong, Yu, et al.
Published: (2025)
by: Hong, Yu, et al.
Published: (2025)
The Evolution of Dataset Distillation: Toward Scalable and Generalizable Solutions
by: Liu, Ping, et al.
Published: (2025)
by: Liu, Ping, et al.
Published: (2025)
VidText: Towards Comprehensive Evaluation for Video Text Understanding
by: Yang, Zhoufaran, et al.
Published: (2025)
by: Yang, Zhoufaran, et al.
Published: (2025)
TextPecker: Rewarding Structural Anomaly Quantification for Enhancing Visual Text Rendering
by: Zhu, Hanshen, et al.
Published: (2026)
by: Zhu, Hanshen, et al.
Published: (2026)
Towards Active Real-to-Twin Inspection: A New Paradigm for Zero-Shot Anomaly Detection
by: Liu, Jiaxuan, et al.
Published: (2026)
by: Liu, Jiaxuan, et al.
Published: (2026)
WAS: Dataset and Methods for Artistic Text Segmentation
by: Xie, Xudong, et al.
Published: (2024)
by: Xie, Xudong, et al.
Published: (2024)
GeoFocus: Blending Efficient Global-to-Local Perception for Multimodal Geometry Problem-Solving
by: Deng, Linger, et al.
Published: (2026)
by: Deng, Linger, et al.
Published: (2026)
TextBlockV2: Towards Precise-Detection-Free Scene Text Spotting with Pre-trained Language Model
by: Lyu, Jiahao, et al.
Published: (2024)
by: Lyu, Jiahao, et al.
Published: (2024)
Towards Real-World Deepfake Detection: A Diverse In-the-wild Dataset of Forgery Faces
by: Shi, Junyu, et al.
Published: (2025)
by: Shi, Junyu, et al.
Published: (2025)
OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models
by: Yu, Wenwen, et al.
Published: (2025)
by: Yu, Wenwen, et al.
Published: (2025)
Mini-Monkey: Alleviating the Semantic Sawtooth Effect for Lightweight MLLMs via Complementary Image Pyramid
by: Huang, Mingxin, et al.
Published: (2024)
by: Huang, Mingxin, et al.
Published: (2024)
Towards Student Actions in Classroom Scenes: New Dataset and Baseline
by: Tan, Zhuolin, et al.
Published: (2024)
by: Tan, Zhuolin, et al.
Published: (2024)
MSPT: A Lightweight Face Image Quality Assessment Method with Multi-stage Progressive Training
by: Xiao, Xiongwei, et al.
Published: (2025)
by: Xiao, Xiongwei, et al.
Published: (2025)
Constructing a Real-World Benchmark for Early Wildfire Detection with the New PYRONEAR-2025 Dataset
by: Lostanlen, Mateo, et al.
Published: (2024)
by: Lostanlen, Mateo, et al.
Published: (2024)
PathVG: A New Benchmark and Dataset for Pathology Visual Grounding
by: Zhong, Chunlin, et al.
Published: (2025)
by: Zhong, Chunlin, et al.
Published: (2025)
Towards Effective Multi-Moving-Camera Tracking: A New Dataset and Lightweight Link Model
by: Zhang, Yanting, et al.
Published: (2023)
by: Zhang, Yanting, et al.
Published: (2023)
ViTEraser: Harnessing the Power of Vision Transformers for Scene Text Removal with SegMIM Pretraining
by: Peng, Dezhi, et al.
Published: (2023)
by: Peng, Dezhi, et al.
Published: (2023)
EgoMe: A New Dataset and Challenge for Following Me via Egocentric View in Real World
by: Qiu, Heqian, et al.
Published: (2025)
by: Qiu, Heqian, et al.
Published: (2025)
A New Benchmark and Model for Challenging Image Manipulation Detection
by: Zhang, Zhenfei, et al.
Published: (2023)
by: Zhang, Zhenfei, et al.
Published: (2023)
TrAME: Trajectory-Anchored Multi-View Editing for Text-Guided 3D Gaussian Splatting Manipulation
by: Luo, Chaofan, et al.
Published: (2024)
by: Luo, Chaofan, et al.
Published: (2024)
Similar Items
-
The First Swahili Language Scene Text Detection and Recognition Dataset
by: Douamba, Fadila Wendigoundi, et al.
Published: (2024) -
SemiETS: Integrating Spatial and Content Consistencies for Semi-Supervised End-to-end Text Spotting
by: Luo, Dongliang, et al.
Published: (2025) -
Dataset and Benchmark for Urdu Natural Scenes Text Detection, Recognition and Visual Question Answering
by: Maryam, Hiba, et al.
Published: (2024) -
Towards UAV Detection in the Real World: A New Multispectral Dataset UAVNet-MS and a New Method
by: Luo, Yihang, et al.
Published: (2026) -
DocThinker: Explainable Multimodal Large Language Models with Rule-based Reinforcement Learning for Document Understanding
by: Yu, Wenwen, et al.
Published: (2025)