Show Me the World in My Language: Establishing the First Baseline for Scene-Text to Scene-Text Translation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Vaidya, Shreyas, Sharma, Arvind Kumar, Gatti, Prajwal, Mishra, Anand |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Composite Sketch+Text Queries for Retrieving Objects with Elusive Names and Complex Interactions
von: Gatti, Prajwal, et al.
Veröffentlicht: (2025)
von: Gatti, Prajwal, et al.
Veröffentlicht: (2025)
Scene-Text Grounding for Text-Based Video Question Answering
von: Zhou, Sheng, et al.
Veröffentlicht: (2024)
von: Zhou, Sheng, et al.
Veröffentlicht: (2024)
MorphText: Deep Morphology Regularized Arbitrary-shape Scene Text Detection
von: Xu, Chengpei, et al.
Veröffentlicht: (2024)
von: Xu, Chengpei, et al.
Veröffentlicht: (2024)
EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering
von: Zhou, Sheng, et al.
Veröffentlicht: (2025)
von: Zhou, Sheng, et al.
Veröffentlicht: (2025)
SceneDreamer360: Text-Driven 3D-Consistent Scene Generation with Panoramic Gaussian Splatting
von: Li, Wenrui, et al.
Veröffentlicht: (2024)
von: Li, Wenrui, et al.
Veröffentlicht: (2024)
Bharat Scene Text: A Novel Comprehensive Dataset and Benchmark for Indian Language Scene Text Understanding
von: De, Anik, et al.
Veröffentlicht: (2025)
von: De, Anik, et al.
Veröffentlicht: (2025)
STEFANN: Scene Text Editor using Font Adaptive Neural Network
von: Roy, Prasun, et al.
Veröffentlicht: (2019)
von: Roy, Prasun, et al.
Veröffentlicht: (2019)
FASTER: A Font-Agnostic Scene Text Editing and Rendering Framework
von: Das, Alloy, et al.
Veröffentlicht: (2023)
von: Das, Alloy, et al.
Veröffentlicht: (2023)
Teach Me Sign: Stepwise Prompting LLM for Sign Language Production
von: An, Zhaoyi, et al.
Veröffentlicht: (2025)
von: An, Zhaoyi, et al.
Veröffentlicht: (2025)
MuLTI: Efficient Video-and-Language Understanding with Text-Guided MultiWay-Sampler and Multiple Choice Modeling
von: Xu, Jiaqi, et al.
Veröffentlicht: (2023)
von: Xu, Jiaqi, et al.
Veröffentlicht: (2023)
Beyond Coarse-Grained Matching in Video-Text Retrieval
von: Chen, Aozhu, et al.
Veröffentlicht: (2024)
von: Chen, Aozhu, et al.
Veröffentlicht: (2024)
DeepMoLM: Leveraging Visual and Geometric Structural Information for Molecule-Text Modeling
von: Lan, Jing, et al.
Veröffentlicht: (2026)
von: Lan, Jing, et al.
Veröffentlicht: (2026)
Improving Gloss-free Sign Language Translation by Reducing Representation Density
von: Ye, Jinhui, et al.
Veröffentlicht: (2024)
von: Ye, Jinhui, et al.
Veröffentlicht: (2024)
DreamArtist++: Controllable One-Shot Text-to-Image Generation via Positive-Negative Adapter
von: Dong, Ziyi, et al.
Veröffentlicht: (2022)
von: Dong, Ziyi, et al.
Veröffentlicht: (2022)
ControlText: Unlocking Controllable Fonts in Multilingual Text Rendering without Font Annotations
von: Jiang, Bowen, et al.
Veröffentlicht: (2025)
von: Jiang, Bowen, et al.
Veröffentlicht: (2025)
Toward Automatic Relevance Judgment using Vision--Language Models for Image--Text Retrieval Evaluation
von: Yang, Jheng-Hong, et al.
Veröffentlicht: (2024)
von: Yang, Jheng-Hong, et al.
Veröffentlicht: (2024)
Scene Graph Generation with Role-Playing Large Language Models
von: Chen, Guikun, et al.
Veröffentlicht: (2024)
von: Chen, Guikun, et al.
Veröffentlicht: (2024)
Discriminative Probing and Tuning for Text-to-Image Generation
von: Qu, Leigang, et al.
Veröffentlicht: (2024)
von: Qu, Leigang, et al.
Veröffentlicht: (2024)
Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion
von: Lv, Zheqi, et al.
Veröffentlicht: (2025)
von: Lv, Zheqi, et al.
Veröffentlicht: (2025)
V-FAT: Benchmarking Visual Fidelity Against Text-bias
von: Wang, Ziteng, et al.
Veröffentlicht: (2026)
von: Wang, Ziteng, et al.
Veröffentlicht: (2026)
Evaluating Text-to-Visual Generation with Image-to-Text Generation
von: Lin, Zhiqiu, et al.
Veröffentlicht: (2024)
von: Lin, Zhiqiu, et al.
Veröffentlicht: (2024)
Spatio-Temporal Data Enhanced Vision-Language Model for Traffic Scene Understanding
von: Ma, Jingtian, et al.
Veröffentlicht: (2025)
von: Ma, Jingtian, et al.
Veröffentlicht: (2025)
Layout-Aware Text Editing for Efficient Transformation of Academic PDFs to Markdown
von: Duan, Changxu
Veröffentlicht: (2025)
von: Duan, Changxu
Veröffentlicht: (2025)
Unleashing Text-to-Image Diffusion Prior for Zero-Shot Image Captioning
von: Luo, Jianjie, et al.
Veröffentlicht: (2024)
von: Luo, Jianjie, et al.
Veröffentlicht: (2024)
HiScene: Creating Hierarchical 3D Scenes with Isometric View Generation
von: Dong, Wenqi, et al.
Veröffentlicht: (2025)
von: Dong, Wenqi, et al.
Veröffentlicht: (2025)
CLOSP: A Unified Semantic Space for SAR, MSI, and Text in Remote Sensing
von: Cambrin, Daniele Rege, et al.
Veröffentlicht: (2025)
von: Cambrin, Daniele Rege, et al.
Veröffentlicht: (2025)
Self-Adaptive Sampling for Efficient Video Question-Answering on Image--Text Models
von: Han, Wei, et al.
Veröffentlicht: (2023)
von: Han, Wei, et al.
Veröffentlicht: (2023)
Both Text and Images Leaked! A Systematic Analysis of Data Contamination in Multimodal LLM
von: Song, Dingjie, et al.
Veröffentlicht: (2024)
von: Song, Dingjie, et al.
Veröffentlicht: (2024)
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
von: Han, Jiaming, et al.
Veröffentlicht: (2025)
von: Han, Jiaming, et al.
Veröffentlicht: (2025)
SonoWorld: From One Image to a 3D Audio-Visual Scene
von: Jin, Derong, et al.
Veröffentlicht: (2026)
von: Jin, Derong, et al.
Veröffentlicht: (2026)
Single Image Dehazing Using Scene Depth Ordering
von: Ling, Pengyang, et al.
Veröffentlicht: (2024)
von: Ling, Pengyang, et al.
Veröffentlicht: (2024)
Dual-Modal Attention-Enhanced Text-Video Retrieval with Triplet Partial Margin Contrastive Learning
von: Jiang, Chen, et al.
Veröffentlicht: (2023)
von: Jiang, Chen, et al.
Veröffentlicht: (2023)
Words or Vision: Do Vision-Language Models Have Blind Faith in Text?
von: Deng, Ailin, et al.
Veröffentlicht: (2025)
von: Deng, Ailin, et al.
Veröffentlicht: (2025)
Modality-Aware Shot Relating and Comparing for Video Scene Detection
von: Tan, Jiawei, et al.
Veröffentlicht: (2024)
von: Tan, Jiawei, et al.
Veröffentlicht: (2024)
IsoSignVid2Aud: Sign Language Video to Audio Conversion without Text Intermediaries
von: Kavediya, Harsh, et al.
Veröffentlicht: (2025)
von: Kavediya, Harsh, et al.
Veröffentlicht: (2025)
UniScene: Multi-Camera Unified Pre-training via 3D Scene Reconstruction for Autonomous Driving
von: Min, Chen, et al.
Veröffentlicht: (2023)
von: Min, Chen, et al.
Veröffentlicht: (2023)
TextBraTS: Text-Guided Volumetric Brain Tumor Segmentation with Innovative Dataset Development and Fusion Module Exploration
von: Shi, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Shi, Xiaoyu, et al.
Veröffentlicht: (2025)
CPSL: Representing Volumetric Video via Content-Promoted Scene Layers
von: Hu, Kaiyuan, et al.
Veröffentlicht: (2025)
von: Hu, Kaiyuan, et al.
Veröffentlicht: (2025)
Scene Aware Person Image Generation through Global Contextual Conditioning
von: Roy, Prasun, et al.
Veröffentlicht: (2022)
von: Roy, Prasun, et al.
Veröffentlicht: (2022)
SCENEFORGE: Enhancing 3D-text alignment with Structured Scene Compositions
von: Sbrolli, Cristian, et al.
Veröffentlicht: (2025)
von: Sbrolli, Cristian, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Composite Sketch+Text Queries for Retrieving Objects with Elusive Names and Complex Interactions
von: Gatti, Prajwal, et al.
Veröffentlicht: (2025) -
Scene-Text Grounding for Text-Based Video Question Answering
von: Zhou, Sheng, et al.
Veröffentlicht: (2024) -
MorphText: Deep Morphology Regularized Arbitrary-shape Scene Text Detection
von: Xu, Chengpei, et al.
Veröffentlicht: (2024) -
EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering
von: Zhou, Sheng, et al.
Veröffentlicht: (2025) -
SceneDreamer360: Text-Driven 3D-Consistent Scene Generation with Panoramic Gaussian Splatting
von: Li, Wenrui, et al.
Veröffentlicht: (2024)