Towards Unified Multi-granularity Text Detection with Interactive Attention
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wan, Xingyu, Zhang, Chengquan, Lyu, Pengyuan, Fan, Sen, Ni, Zihan, Yao, Kun, Ding, Errui, Wang, Jingdong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
StrucTexTv3: An Efficient Vision-Language Model for Text-rich Image Perception, Comprehension, and Beyond
von: Lyu, Pengyuan, et al.
Veröffentlicht: (2024)
von: Lyu, Pengyuan, et al.
Veröffentlicht: (2024)
Towards Real-World Document Parsing via Realistic Scene Synthesis and Document-Aware Training
von: Li, Gengluo, et al.
Veröffentlicht: (2026)
von: Li, Gengluo, et al.
Veröffentlicht: (2026)
MMTIT-Bench: A Multilingual and Multi-Scenario Benchmark with Cognition-Perception-Reasoning Guided Text-Image Machine Translation
von: Li, Gengluo, et al.
Veröffentlicht: (2026)
von: Li, Gengluo, et al.
Veröffentlicht: (2026)
Recognition-Synergistic Scene Text Editing
von: Fang, Zhengyao, et al.
Veröffentlicht: (2025)
von: Fang, Zhengyao, et al.
Veröffentlicht: (2025)
OVLW-DETR: Open-Vocabulary Light-Weighted Detection Transformer
von: Wang, Yu, et al.
Veröffentlicht: (2024)
von: Wang, Yu, et al.
Veröffentlicht: (2024)
MGMapNet: Multi-Granularity Representation Learning for End-to-End Vectorized HD Map Construction
von: Yang, Jing, et al.
Veröffentlicht: (2024)
von: Yang, Jing, et al.
Veröffentlicht: (2024)
PUMA: Empowering Unified MLLM with Multi-granular Visual Generation
von: Fang, Rongyao, et al.
Veröffentlicht: (2024)
von: Fang, Rongyao, et al.
Veröffentlicht: (2024)
Prune Redundancy, Preserve Essence: Vision Token Compression in VLMs via Synergistic Importance-Diversity
von: Fang, Zhengyao, et al.
Veröffentlicht: (2026)
von: Fang, Zhengyao, et al.
Veröffentlicht: (2026)
Interpretable Face Anti-Spoofing: Enhancing Generalization with Multimodal Large Language Models
von: Zhang, Guosheng, et al.
Veröffentlicht: (2025)
von: Zhang, Guosheng, et al.
Veröffentlicht: (2025)
WeCromCL: Weakly Supervised Cross-Modality Contrastive Learning for Transcription-only Supervised Text Spotting
von: Wu, Jingjing, et al.
Veröffentlicht: (2024)
von: Wu, Jingjing, et al.
Veröffentlicht: (2024)
Uni$^2$Det: Unified and Universal Framework for Prompt-Guided Multi-dataset 3D Detection
von: Wang, Yubin, et al.
Veröffentlicht: (2024)
von: Wang, Yubin, et al.
Veröffentlicht: (2024)
OPEN: Object-wise Position Embedding for Multi-view 3D Object Detection
von: Hou, Jinghua, et al.
Veröffentlicht: (2024)
von: Hou, Jinghua, et al.
Veröffentlicht: (2024)
Re-HOLD: Video Hand Object Interaction Reenactment via adaptive Layout-instructed Diffusion Model
von: Fan, Yingying, et al.
Veröffentlicht: (2025)
von: Fan, Yingying, et al.
Veröffentlicht: (2025)
SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation
von: Ge, Yuying, et al.
Veröffentlicht: (2024)
von: Ge, Yuying, et al.
Veröffentlicht: (2024)
Skim then Focus: Integrating Contextual and Fine-grained Views for Repetitive Action Counting
von: Zhao, Zhengqi, et al.
Veröffentlicht: (2024)
von: Zhao, Zhengqi, et al.
Veröffentlicht: (2024)
FullAnno: A Data Engine for Enhancing Image Comprehension of MLLMs
von: Hao, Jing, et al.
Veröffentlicht: (2024)
von: Hao, Jing, et al.
Veröffentlicht: (2024)
Explanatory Instructions: Towards Unified Vision Tasks Understanding and Zero-shot Generalization
von: Shen, Yang, et al.
Veröffentlicht: (2024)
von: Shen, Yang, et al.
Veröffentlicht: (2024)
Add-SD: Rational Generation without Manual Reference
von: Yang, Lingfeng, et al.
Veröffentlicht: (2024)
von: Yang, Lingfeng, et al.
Veröffentlicht: (2024)
CODER: Coupled Diversity-Sensitive Momentum Contrastive Learning for Image-Text Retrieval
von: Wang, Haoran, et al.
Veröffentlicht: (2022)
von: Wang, Haoran, et al.
Veröffentlicht: (2022)
Gradient-based Sampling for Class Imbalanced Semi-supervised Object Detection
von: Li, Jiaming, et al.
Veröffentlicht: (2024)
von: Li, Jiaming, et al.
Veröffentlicht: (2024)
Decoupled Pseudo-labeling for Semi-Supervised Monocular 3D Object Detection
von: Zhang, Jiacheng, et al.
Veröffentlicht: (2024)
von: Zhang, Jiacheng, et al.
Veröffentlicht: (2024)
MS-DETR: Efficient DETR Training with Mixed Supervision
von: Zhao, Chuyang, et al.
Veröffentlicht: (2024)
von: Zhao, Chuyang, et al.
Veröffentlicht: (2024)
GGRt: Towards Pose-free Generalizable 3D Gaussian Splatting in Real-time
von: Li, Hao, et al.
Veröffentlicht: (2024)
von: Li, Hao, et al.
Veröffentlicht: (2024)
LaMI-DETR: Open-Vocabulary Detection with Language Model Instruction
von: Du, Penghui, et al.
Veröffentlicht: (2024)
von: Du, Penghui, et al.
Veröffentlicht: (2024)
LW-DETR: A Transformer Replacement to YOLO for Real-Time Detection
von: Chen, Qiang, et al.
Veröffentlicht: (2024)
von: Chen, Qiang, et al.
Veröffentlicht: (2024)
MonoFormer: One Transformer for Both Diffusion and Autoregression
von: Zhao, Chuyang, et al.
Veröffentlicht: (2024)
von: Zhao, Chuyang, et al.
Veröffentlicht: (2024)
TopoSD: Topology-Enhanced Lane Segment Perception with SDMap Prior
von: Yang, Sen, et al.
Veröffentlicht: (2024)
von: Yang, Sen, et al.
Veröffentlicht: (2024)
A Mutual Learning Method for Salient Object Detection with intertwined Multi-Supervision--Revised
von: Wu, Runmin, et al.
Veröffentlicht: (2025)
von: Wu, Runmin, et al.
Veröffentlicht: (2025)
OpenGaussian: Towards Point-Level 3D Gaussian-based Open Vocabulary Understanding
von: Wu, Yanmin, et al.
Veröffentlicht: (2024)
von: Wu, Yanmin, et al.
Veröffentlicht: (2024)
GVA: Reconstructing Vivid 3D Gaussian Avatars from Monocular Videos
von: Liu, Xinqi, et al.
Veröffentlicht: (2024)
von: Liu, Xinqi, et al.
Veröffentlicht: (2024)
PaintFlow: A Unified Framework for Interactive Oil Paintings Editing and Generation
von: Hu, Zhangli, et al.
Veröffentlicht: (2025)
von: Hu, Zhangli, et al.
Veröffentlicht: (2025)
TCMA: Text-Conditioned Multi-granularity Alignment for Drone Cross-Modal Text-Video Retrieval
von: Zhao, Zixu, et al.
Veröffentlicht: (2025)
von: Zhao, Zixu, et al.
Veröffentlicht: (2025)
Uni-ISP: Toward Unifying the Learning of ISPs from Multiple Mobile Cameras
von: Li, Lingen, et al.
Veröffentlicht: (2024)
von: Li, Lingen, et al.
Veröffentlicht: (2024)
UniLION: Towards Unified Autonomous Driving Model with Linear Group RNNs
von: Liu, Zhe, et al.
Veröffentlicht: (2025)
von: Liu, Zhe, et al.
Veröffentlicht: (2025)
A Novel Unified Approach to Deepfake Detection
von: Sen, Lord, et al.
Veröffentlicht: (2026)
von: Sen, Lord, et al.
Veröffentlicht: (2026)
M2DAO-Talker: Harmonizing Multi-granular Motion Decoupling and Alternating Optimization for Talking-head Generation
von: Jiang, Kui, et al.
Veröffentlicht: (2025)
von: Jiang, Kui, et al.
Veröffentlicht: (2025)
VRP-SAM: SAM with Visual Reference Prompt
von: Sun, Yanpeng, et al.
Veröffentlicht: (2024)
von: Sun, Yanpeng, et al.
Veröffentlicht: (2024)
XLD: A Cross-Lane Dataset for Benchmarking Novel Driving View Synthesis
von: Li, Hao, et al.
Veröffentlicht: (2024)
von: Li, Hao, et al.
Veröffentlicht: (2024)
TexRO: Generating Delicate Textures of 3D Models by Recursive Optimization
von: Wu, Jinbo, et al.
Veröffentlicht: (2024)
von: Wu, Jinbo, et al.
Veröffentlicht: (2024)
GIR: 3D Gaussian Inverse Rendering for Relightable Scene Factorization
von: Shi, Yahao, et al.
Veröffentlicht: (2023)
von: Shi, Yahao, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
StrucTexTv3: An Efficient Vision-Language Model for Text-rich Image Perception, Comprehension, and Beyond
von: Lyu, Pengyuan, et al.
Veröffentlicht: (2024) -
Towards Real-World Document Parsing via Realistic Scene Synthesis and Document-Aware Training
von: Li, Gengluo, et al.
Veröffentlicht: (2026) -
MMTIT-Bench: A Multilingual and Multi-Scenario Benchmark with Cognition-Perception-Reasoning Guided Text-Image Machine Translation
von: Li, Gengluo, et al.
Veröffentlicht: (2026) -
Recognition-Synergistic Scene Text Editing
von: Fang, Zhengyao, et al.
Veröffentlicht: (2025) -
OVLW-DETR: Open-Vocabulary Light-Weighted Detection Transformer
von: Wang, Yu, et al.
Veröffentlicht: (2024)