TiCLS : Tightly Coupled Language Text Spotter
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jang, Leeje, Lin, Yijun, Chiang, Yao-Yi, Weinman, Jerod |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LIGHT: Multi-Modal Text Linking on Historical Maps
von: Lin, Yijun, et al.
Veröffentlicht: (2025)
von: Lin, Yijun, et al.
Veröffentlicht: (2025)
Hyper-Local Deformable Transformers for Text Spotting on Historical Maps
von: Lin, Yijun, et al.
Veröffentlicht: (2025)
von: Lin, Yijun, et al.
Veröffentlicht: (2025)
EdgeSpotter: Multi-Scale Dense Text Spotting for Industrial Panel Monitoring
von: Fu, Changhong, et al.
Veröffentlicht: (2025)
von: Fu, Changhong, et al.
Veröffentlicht: (2025)
SwinTextSpotter v2: Towards Better Synergy for Scene Text Spotting
von: Huang, Mingxin, et al.
Veröffentlicht: (2024)
von: Huang, Mingxin, et al.
Veröffentlicht: (2024)
Arbitrary Reading Order Scene Text Spotter with Local Semantics Guidance
von: Lyu, Jiahao, et al.
Veröffentlicht: (2024)
von: Lyu, Jiahao, et al.
Veröffentlicht: (2024)
FastTextSpotter: A High-Efficiency Transformer for Multilingual Scene Text Spotting
von: Das, Alloy, et al.
Veröffentlicht: (2024)
von: Das, Alloy, et al.
Veröffentlicht: (2024)
FRIEDA: Benchmarking Multi-Step Cartographic Reasoning in Vision-Language Models
von: Pyo, Jiyoon, et al.
Veröffentlicht: (2025)
von: Pyo, Jiyoon, et al.
Veröffentlicht: (2025)
DIGMAPPER: A Modular System for Automated Geologic Map Digitization
von: Duan, Weiwei, et al.
Veröffentlicht: (2025)
von: Duan, Weiwei, et al.
Veröffentlicht: (2025)
TextFormer: A Query-based End-to-End Text Spotter with Mixed Supervision
von: Zhai, Yukun, et al.
Veröffentlicht: (2023)
von: Zhai, Yukun, et al.
Veröffentlicht: (2023)
HotSpotter - Patterned Species Instance Recognition
von: Crall, Jonathan P., et al.
Veröffentlicht: (2025)
von: Crall, Jonathan P., et al.
Veröffentlicht: (2025)
Spotter+GPT: Turning Sign Spottings into Sentences with LLMs
von: Sincan, Ozge Mercanoglu, et al.
Veröffentlicht: (2024)
von: Sincan, Ozge Mercanoglu, et al.
Veröffentlicht: (2024)
"ScatSpotter" -- A Dog Poop Detection Dataset
von: Crall, Jon
Veröffentlicht: (2024)
von: Crall, Jon
Veröffentlicht: (2024)
DNTextSpotter: Arbitrary-Shaped Scene Text Spotting via Improved Denoising Training
von: Xie, Yu, et al.
Veröffentlicht: (2024)
von: Xie, Yu, et al.
Veröffentlicht: (2024)
[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs
von: Wang, Ao, et al.
Veröffentlicht: (2024)
von: Wang, Ao, et al.
Veröffentlicht: (2024)
VimTS: A Unified Video and Image Text Spotter for Enhancing the Cross-domain Generalization
von: Liu, Yuliang, et al.
Veröffentlicht: (2024)
von: Liu, Yuliang, et al.
Veröffentlicht: (2024)
CASP: Few-Shot Class-Incremental Learning with CLS Token Attention Steering Prompts
von: Huang, Shuai, et al.
Veröffentlicht: (2026)
von: Huang, Shuai, et al.
Veröffentlicht: (2026)
[CLS] is Not Enough: Multi-Label Recognition via Patch-Level Inference and Adaptive Aggregation
von: Wang, Akang, et al.
Veröffentlicht: (2026)
von: Wang, Akang, et al.
Veröffentlicht: (2026)
Unsqueeze [CLS] Bottleneck to Learn Rich Representations
von: Su, Qing, et al.
Veröffentlicht: (2024)
von: Su, Qing, et al.
Veröffentlicht: (2024)
Unlocking [CLS] Features for Continual Post-Training
von: Yildirim, Murat Onur, et al.
Veröffentlicht: (2025)
von: Yildirim, Murat Onur, et al.
Veröffentlicht: (2025)
Multi-Level CLS Token Fusion for Contrastive Learning in Endoscopy Image Classification
von: Nguyen, Y Hop, et al.
Veröffentlicht: (2025)
von: Nguyen, Y Hop, et al.
Veröffentlicht: (2025)
OMNI-Dent: Towards an Accessible and Explainable AI Framework for Automated Dental Diagnosis
von: Jang, Leeje, et al.
Veröffentlicht: (2026)
von: Jang, Leeje, et al.
Veröffentlicht: (2026)
Revisiting [CLS] and Patch Token Interaction in Vision Transformers
von: Marouani, Alexis, et al.
Veröffentlicht: (2026)
von: Marouani, Alexis, et al.
Veröffentlicht: (2026)
KM-ViPE: Online Tightly Coupled Vision-Language-Geometry Fusion for Open-Vocabulary Semantic SLAM
von: Nasser, Zaid, et al.
Veröffentlicht: (2025)
von: Nasser, Zaid, et al.
Veröffentlicht: (2025)
FlexID: Training-Free Flexible Identity Injection via Intent-Aware Modulation for Text-to-Image Generation
von: Li, Guandong, et al.
Veröffentlicht: (2026)
von: Li, Guandong, et al.
Veröffentlicht: (2026)
TCLC-GS: Tightly Coupled LiDAR-Camera Gaussian Splatting for Autonomous Driving
von: Zhao, Cheng, et al.
Veröffentlicht: (2024)
von: Zhao, Cheng, et al.
Veröffentlicht: (2024)
Register and [CLS] tokens yield a decoupling of local and global features in large ViTs
von: Lappe, Alexander, et al.
Veröffentlicht: (2025)
von: Lappe, Alexander, et al.
Veröffentlicht: (2025)
Tightly-Coupled, Speed-aided Monocular Visual-Inertial Localization in Topological Map
von: Yang, Chanuk, et al.
Veröffentlicht: (2024)
von: Yang, Chanuk, et al.
Veröffentlicht: (2024)
WalkCLIP: Multimodal Learning for Urban Walkability Prediction
von: Xiang, Shilong, et al.
Veröffentlicht: (2025)
von: Xiang, Shilong, et al.
Veröffentlicht: (2025)
Cross-Modal Action Recognition in Egocentric Video Using Mamba: Integrating RGB and Hand Skeleton Streams via CLS Token Fusion Strategies
von: Gorostegui, Juan Ignacio Bustos, et al.
Veröffentlicht: (2026)
von: Gorostegui, Juan Ignacio Bustos, et al.
Veröffentlicht: (2026)
X-Oscar: A Progressive Framework for High-quality Text-guided 3D Animatable Avatar Generation
von: Ma, Yiwei, et al.
Veröffentlicht: (2024)
von: Ma, Yiwei, et al.
Veröffentlicht: (2024)
Fine-Scale Soil Mapping in Alaska with Multimodal Machine Learning
von: Lin, Yijun, et al.
Veröffentlicht: (2025)
von: Lin, Yijun, et al.
Veröffentlicht: (2025)
LayerFusion: Harmonized Multi-Layer Text-to-Image Generation with Generative Priors
von: Dalva, Yusuf, et al.
Veröffentlicht: (2024)
von: Dalva, Yusuf, et al.
Veröffentlicht: (2024)
TiP4GEN: Text to Immersive Panorama 4D Scene Generation
von: Xing, Ke, et al.
Veröffentlicht: (2025)
von: Xing, Ke, et al.
Veröffentlicht: (2025)
Text-Anchored Score Composition: Tackling Condition Misalignment in Text-to-Image Diffusion Models
von: Wang, Luozhou, et al.
Veröffentlicht: (2023)
von: Wang, Luozhou, et al.
Veröffentlicht: (2023)
TransPixeler: Advancing Text-to-Video Generation with Transparency
von: Wang, Luozhou, et al.
Veröffentlicht: (2025)
von: Wang, Luozhou, et al.
Veröffentlicht: (2025)
AQUA-SLAM: Tightly-Coupled Underwater Acoustic-Visual-Inertial SLAM with Sensor Calibration
von: Xu, Shida, et al.
Veröffentlicht: (2025)
von: Xu, Shida, et al.
Veröffentlicht: (2025)
Beyond Optimal Transport: Model-Aligned Coupling for Flow Matching
von: Lin, Yexiong, et al.
Veröffentlicht: (2025)
von: Lin, Yexiong, et al.
Veröffentlicht: (2025)
RADIO-ViPE: Online Tightly Coupled Multi-Modal Fusion for Open-Vocabulary Semantic SLAM in Dynamic Environments
von: Nasser, Zaid, et al.
Veröffentlicht: (2026)
von: Nasser, Zaid, et al.
Veröffentlicht: (2026)
Removing Distributional Discrepancies in Captions Improves Image-Text Alignment
von: Li, Yuheng, et al.
Veröffentlicht: (2024)
von: Li, Yuheng, et al.
Veröffentlicht: (2024)
GuardT2I: Defending Text-to-Image Models from Adversarial Prompts
von: Yang, Yijun, et al.
Veröffentlicht: (2024)
von: Yang, Yijun, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
LIGHT: Multi-Modal Text Linking on Historical Maps
von: Lin, Yijun, et al.
Veröffentlicht: (2025) -
Hyper-Local Deformable Transformers for Text Spotting on Historical Maps
von: Lin, Yijun, et al.
Veröffentlicht: (2025) -
EdgeSpotter: Multi-Scale Dense Text Spotting for Industrial Panel Monitoring
von: Fu, Changhong, et al.
Veröffentlicht: (2025) -
SwinTextSpotter v2: Towards Better Synergy for Scene Text Spotting
von: Huang, Mingxin, et al.
Veröffentlicht: (2024) -
Arbitrary Reading Order Scene Text Spotter with Local Semantics Guidance
von: Lyu, Jiahao, et al.
Veröffentlicht: (2024)