Winner Team Mia at TextVQA Challenge 2021: Vision-and-Language Representation Learning with Pre-trained Sequence-to-Sequence Model
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Qiao, Yixuan, Chen, Hao, Wang, Jun, Zhao, Shanshan, Chen, Yihao, Ye, Xianbin, Li, Ziliang, Qi, Xianbiao, Gao, Peng, Xie, Guotong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2021
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VTAgent: Agentic Keyframe Anchoring for Evidence-Aware Video TextVQA
von: He, Haibin, et al.
Veröffentlicht: (2026)
von: He, Haibin, et al.
Veröffentlicht: (2026)
PASH at TREC 2021 Deep Learning Track: Generative Enhanced Model for Multi-stage Ranking
von: Qiao, Yixuan, et al.
Veröffentlicht: (2022)
von: Qiao, Yixuan, et al.
Veröffentlicht: (2022)
Gather and Trace: Rethinking Video TextVQA from an Instance-oriented Perspective
von: Zhang, Yan, et al.
Veröffentlicht: (2025)
von: Zhang, Yan, et al.
Veröffentlicht: (2025)
Track the Answer: Extending TextVQA from Image to Video with Spatio-Temporal Clues
von: Zhang, Yan, et al.
Veröffentlicht: (2024)
von: Zhang, Yan, et al.
Veröffentlicht: (2024)
SFA: Scan, Focus, and Amplify toward Guidance-aware Answering for Video TextVQA
von: He, Haibin, et al.
Veröffentlicht: (2025)
von: He, Haibin, et al.
Veröffentlicht: (2025)
Sequence-to-Sequence Spanish Pre-trained Language Models
von: Araujo, Vladimir, et al.
Veröffentlicht: (2023)
von: Araujo, Vladimir, et al.
Veröffentlicht: (2023)
AdaMR: Adaptable Molecular Representation for Unified Pre-training Strategy
von: Ding, Yan, et al.
Veröffentlicht: (2023)
von: Ding, Yan, et al.
Veröffentlicht: (2023)
SFM-Protein: Integrative Co-evolutionary Pre-training for Advanced Protein Sequence Representation
von: He, Liang, et al.
Veröffentlicht: (2024)
von: He, Liang, et al.
Veröffentlicht: (2024)
Exploiting Pre-trained Encoder-Decoder Transformers for Sequence-to-Sequence Constituent Parsing
von: Fernández-González, Daniel, et al.
Veröffentlicht: (2026)
von: Fernández-González, Daniel, et al.
Veröffentlicht: (2026)
Spatial-Temporal Cross-View Contrastive Pre-training for Check-in Sequence Representation Learning
von: Gong, Letian, et al.
Veröffentlicht: (2024)
von: Gong, Letian, et al.
Veröffentlicht: (2024)
S$^2$ALM: Sequence-Structure Pre-trained Large Language Model for Comprehensive Antibody Representation Learning
von: Yin, Mingze, et al.
Veröffentlicht: (2024)
von: Yin, Mingze, et al.
Veröffentlicht: (2024)
PrahokBART: A Pre-trained Sequence-to-Sequence Model for Khmer Natural Language Generation
von: Kaing, Hour, et al.
Veröffentlicht: (2025)
von: Kaing, Hour, et al.
Veröffentlicht: (2025)
Non-autoregressive Sequence-to-Sequence Vision-Language Models
von: Shi, Kunyu, et al.
Veröffentlicht: (2024)
von: Shi, Kunyu, et al.
Veröffentlicht: (2024)
Modularized Zero-shot VQA with Pre-trained Models
von: Cao, Rui, et al.
Veröffentlicht: (2023)
von: Cao, Rui, et al.
Veröffentlicht: (2023)
Chinese Sequence Labeling with Semi-Supervised Boundary-Aware Language Model Pre-training
von: Zhang, Longhui, et al.
Veröffentlicht: (2024)
von: Zhang, Longhui, et al.
Veröffentlicht: (2024)
Superpixel Semantics Representation and Pre-training for Vision-Language Task
von: Zhang, Siyu, et al.
Veröffentlicht: (2023)
von: Zhang, Siyu, et al.
Veröffentlicht: (2023)
Efficient Vision-and-Language Pre-training with Text-Relevant Image Patch Selection
von: Ye, Wei, et al.
Veröffentlicht: (2024)
von: Ye, Wei, et al.
Veröffentlicht: (2024)
Towards Cultural Bridge by Bahnaric-Vietnamese Translation Using Transfer Learning of Sequence-To-Sequence Pre-training Language Model
von: Dat, Phan Tran Minh, et al.
Veröffentlicht: (2025)
von: Dat, Phan Tran Minh, et al.
Veröffentlicht: (2025)
Sequence Representation
von: Anonymous Author
Veröffentlicht: (2026)
von: Anonymous Author
Veröffentlicht: (2026)
DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD
von: Qi, Xianbiao, et al.
Veröffentlicht: (2025)
von: Qi, Xianbiao, et al.
Veröffentlicht: (2025)
The Full‐Length Genomic Sequence of the HLA‐DPA1*02:103 Allele Was Identified Using Oxford Nanopore Sequencing
von: Gang Li, et al.
Veröffentlicht: (2025)
von: Gang Li, et al.
Veröffentlicht: (2025)
Pre-trained Models Succeed in Medical Imaging with Representation Similarity Degradation
von: Zu, Wenqiang, et al.
Veröffentlicht: (2025)
von: Zu, Wenqiang, et al.
Veröffentlicht: (2025)
UniVid: Unifying Vision Tasks with Pre-trained Video Generation Models
von: Chen, Lan, et al.
Veröffentlicht: (2025)
von: Chen, Lan, et al.
Veröffentlicht: (2025)
Beyond the Sequence: Statistics-Driven Pre-training for Stabilizing Sequential Recommendation Model
von: Wang, Sirui, et al.
Veröffentlicht: (2024)
von: Wang, Sirui, et al.
Veröffentlicht: (2024)
Can LLM-Generated Text Empower Surgical Vision-Language Pre-training?
von: Che, Chengan, et al.
Veröffentlicht: (2026)
von: Che, Chengan, et al.
Veröffentlicht: (2026)
Caption First, VQA Second: Knowledge Density, Not Task Format, Drives Multimodal Scaling
von: Zou, Hongjian, et al.
Veröffentlicht: (2026)
von: Zou, Hongjian, et al.
Veröffentlicht: (2026)
Transforming Chatbot Text: A Sequence-to-Sequence Approach
von: Reddy, Natesh, et al.
Veröffentlicht: (2025)
von: Reddy, Natesh, et al.
Veröffentlicht: (2025)
COPA: Efficient Vision-Language Pre-training Through Collaborative Object- and Patch-Text Alignment
von: Jiang, Chaoya, et al.
Veröffentlicht: (2023)
von: Jiang, Chaoya, et al.
Veröffentlicht: (2023)
Revisiting Continual Semantic Segmentation with Pre-trained Vision Models
von: Zhang, Duzhen, et al.
Veröffentlicht: (2025)
von: Zhang, Duzhen, et al.
Veröffentlicht: (2025)
TiMix: Text-aware Image Mixing for Effective Vision-Language Pre-training
von: Jiang, Chaoya, et al.
Veröffentlicht: (2023)
von: Jiang, Chaoya, et al.
Veröffentlicht: (2023)
UltraSeP: Sequence-aware Pre-training for Echocardiography Probe Movement Guidance
von: Jiang, Haojun, et al.
Veröffentlicht: (2024)
von: Jiang, Haojun, et al.
Veröffentlicht: (2024)
Exploring Pre-trained Text-to-Video Diffusion Models for Referring Video Object Segmentation
von: Zhu, Zixin, et al.
Veröffentlicht: (2024)
von: Zhu, Zixin, et al.
Veröffentlicht: (2024)
TRIPS: Efficient Vision-and-Language Pre-training with Text-Relevant Image Patch Selection
von: Jiang, Chaoya, et al.
Veröffentlicht: (2023)
von: Jiang, Chaoya, et al.
Veröffentlicht: (2023)
Boosting Vision Semantic Density with Anatomy Normality Modeling for Medical Vision-language Pre-training
von: Cao, Weiwei, et al.
Veröffentlicht: (2025)
von: Cao, Weiwei, et al.
Veröffentlicht: (2025)
SimpleGPT: Improving GPT via A Simple Normalization Strategy
von: Chen, Marco, et al.
Veröffentlicht: (2026)
von: Chen, Marco, et al.
Veröffentlicht: (2026)
Delving into Muon and Beyond: Deep Analysis and Extensions
von: Qi, Xianbiao, et al.
Veröffentlicht: (2026)
von: Qi, Xianbiao, et al.
Veröffentlicht: (2026)
Sequence-to-Sequence Neural Diarization with Automatic Speaker Detection and Representation
von: Cheng, Ming, et al.
Veröffentlicht: (2024)
von: Cheng, Ming, et al.
Veröffentlicht: (2024)
The Cinema of Mia Hansen-Løve
von: Ince, Kate
Veröffentlicht: (2023)
von: Ince, Kate
Veröffentlicht: (2023)
A Novel Graph-Sequence Learning Model for Inductive Text Classification
von: Wang, Zuo, et al.
Veröffentlicht: (2025)
von: Wang, Zuo, et al.
Veröffentlicht: (2025)
Vision Model Pre-training on Interleaved Image-Text Data via Latent Compression Learning
von: Yang, Chenyu, et al.
Veröffentlicht: (2024)
von: Yang, Chenyu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
VTAgent: Agentic Keyframe Anchoring for Evidence-Aware Video TextVQA
von: He, Haibin, et al.
Veröffentlicht: (2026) -
PASH at TREC 2021 Deep Learning Track: Generative Enhanced Model for Multi-stage Ranking
von: Qiao, Yixuan, et al.
Veröffentlicht: (2022) -
Gather and Trace: Rethinking Video TextVQA from an Instance-oriented Perspective
von: Zhang, Yan, et al.
Veröffentlicht: (2025) -
Track the Answer: Extending TextVQA from Image to Video with Spatio-Temporal Clues
von: Zhang, Yan, et al.
Veröffentlicht: (2024) -
SFA: Scan, Focus, and Amplify toward Guidance-aware Answering for Video TextVQA
von: He, Haibin, et al.
Veröffentlicht: (2025)