SEA-Vision: A Multilingual Benchmark for Comprehensive Document and Scene Text Understanding in Southeast Asia
Fuente:
arXiv
Saved in:
| Main Authors: | Yue, Pengfei, Zhao, Xingran, Chen, Juntao, Hou, Peng, Longchao, Wang, Lin, Jianghang, Zhang, Shengchuan, Zeng, Anxiang, Cao, Liujuan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Referring Industrial Anomaly Segmentation
by: Yue, Pengfei, et al.
Published: (2026)
by: Yue, Pengfei, et al.
Published: (2026)
Active-SAOOD: Active Sparsely Annotated Oriented Object Detection in Remote Sensing Images
by: Lin, Yu, et al.
Published: (2026)
by: Lin, Yu, et al.
Published: (2026)
What You Perceive Is What You Conceive: A Cognition-Inspired Framework for Open Vocabulary Image Segmentation
by: Lin, Jianghang, et al.
Published: (2025)
by: Lin, Jianghang, et al.
Published: (2025)
Pseudo-Label Quality Decoupling and Correction for Semi-Supervised Instance Segmentation
by: Lin, Jianghang, et al.
Published: (2025)
by: Lin, Jianghang, et al.
Published: (2025)
GOI: Find 3D Gaussians of Interest with an Optimizable Open-vocabulary Semantic-space Hyperplane
by: Qu, Yansong, et al.
Published: (2024)
by: Qu, Yansong, et al.
Published: (2024)
Training-Free Hierarchical Scene Understanding for Gaussian Splatting with Superpoint Graphs
by: Dai, Shaohui, et al.
Published: (2025)
by: Dai, Shaohui, et al.
Published: (2025)
SEA-Guard: Culturally Grounded Multilingual Safeguard for Southeast Asia
by: Tasawong, Panuthep, et al.
Published: (2026)
by: Tasawong, Panuthep, et al.
Published: (2026)
S$^2$Teacher: Step-by-step Teacher for Sparsely Annotated Oriented Object Detection
by: Lin, Yu, et al.
Published: (2025)
by: Lin, Yu, et al.
Published: (2025)
Generate Aligned Anomaly: Region-Guided Few-Shot Anomaly Image-Mask Pair Synthesis for Industrial Inspection
by: Lu, Yilin, et al.
Published: (2025)
by: Lu, Yilin, et al.
Published: (2025)
Understanding What Is Not Said:Referring Remote Sensing Image Segmentation with Scarce Expressions
by: Ye, Kai, et al.
Published: (2025)
by: Ye, Kai, et al.
Published: (2025)
Exploring Semantic Consistency and Style Diversity for Domain Generalized Semantic Segmentation
by: Niu, Hongwei, et al.
Published: (2024)
by: Niu, Hongwei, et al.
Published: (2024)
SCOUT: Semi-supervised Camouflaged Object Detection by Utilizing Text and Adaptive Data Selection
by: Yan, Weiqi, et al.
Published: (2025)
by: Yan, Weiqi, et al.
Published: (2025)
HieraVid: Hierarchical Token Pruning for Fast Video Large Language Models
by: Guo, Yansong, et al.
Published: (2026)
by: Guo, Yansong, et al.
Published: (2026)
FlashWorld: High-quality 3D Scene Generation within Seconds
by: Li, Xinyang, et al.
Published: (2025)
by: Li, Xinyang, et al.
Published: (2025)
Director3D: Real-world Camera Trajectory and 3D Scene Generation from Text
by: Li, Xinyang, et al.
Published: (2024)
by: Li, Xinyang, et al.
Published: (2024)
Discover, Segment, and Select: A Progressive Mechanism for Zero-shot Camouflaged Object Segmentation
by: Yang, Yilong, et al.
Published: (2026)
by: Yang, Yilong, et al.
Published: (2026)
SynergyAmodal: Deocclude Anything with Text Control
by: Li, Xinyang, et al.
Published: (2025)
by: Li, Xinyang, et al.
Published: (2025)
EOV-Seg: Efficient Open-Vocabulary Panoptic Segmentation
by: Niu, Hongwei, et al.
Published: (2024)
by: Niu, Hongwei, et al.
Published: (2024)
Depth-Guided Semi-Supervised Instance Segmentation
by: Chen, Xin, et al.
Published: (2024)
by: Chen, Xin, et al.
Published: (2024)
Evolving, Not Training: Zero-Shot Reasoning Segmentation via Evolutionary Prompting
by: Ye, Kai, et al.
Published: (2025)
by: Ye, Kai, et al.
Published: (2025)
AnomalyPainter: Vision-Language-Diffusion Synergy for Zero-Shot Realistic and Diverse Industrial Anomaly Synthesis
by: Lai, Zhangyu, et al.
Published: (2025)
by: Lai, Zhangyu, et al.
Published: (2025)
Crowdsource, Crawl, or Generate? Creating SEA-VL, a Multicultural Vision-Language Dataset for Southeast Asia
by: Cahyawijaya, Samuel, et al.
Published: (2025)
by: Cahyawijaya, Samuel, et al.
Published: (2025)
Breaking the Bias: Recalibrating the Attention of Industrial Anomaly Detection
by: Chen, Xin, et al.
Published: (2024)
by: Chen, Xin, et al.
Published: (2024)
HUWSOD: Holistic Self-training for Unified Weakly Supervised Object Detection
by: Cao, Liujuan, et al.
Published: (2024)
by: Cao, Liujuan, et al.
Published: (2024)
Bharat Scene Text: A Novel Comprehensive Dataset and Benchmark for Indian Language Scene Text Understanding
by: De, Anik, et al.
Published: (2025)
by: De, Anik, et al.
Published: (2025)
TextRefiner: Internal Visual Feature as Efficient Refiner for Vision-Language Models Prompt Tuning
by: Xie, Jingjing, et al.
Published: (2024)
by: Xie, Jingjing, et al.
Published: (2024)
SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia
by: Liu, Chaoqun, et al.
Published: (2025)
by: Liu, Chaoqun, et al.
Published: (2025)
Dual3D: Efficient and Consistent Text-to-3D Generation with Dual-mode Multi-view Latent Diffusion
by: Li, Xinyang, et al.
Published: (2024)
by: Li, Xinyang, et al.
Published: (2024)
HRSAM: Efficient Interactive Segmentation in High-Resolution Images
by: Huang, You, et al.
Published: (2024)
by: Huang, You, et al.
Published: (2024)
Inter2Former: Dynamic Hybrid Attention for Efficient High-Precision Interactive
by: Huang, You, et al.
Published: (2025)
by: Huang, You, et al.
Published: (2025)
UCOD-DPL: Unsupervised Camouflaged Object Detection via Dynamic Pseudo-label Learning
by: Yan, Weiqi, et al.
Published: (2025)
by: Yan, Weiqi, et al.
Published: (2025)
Evolving High-Quality Rendering and Reconstruction in a Unified Framework with Contribution-Adaptive Regularization
by: Shen, You, et al.
Published: (2025)
by: Shen, You, et al.
Published: (2025)
FocSAM: Delving Deeply into Focused Objects in Segmenting Anything
by: Huang, You, et al.
Published: (2024)
by: Huang, You, et al.
Published: (2024)
Filipino Benchmarks for Measuring Sexist and Homophobic Bias in Multilingual Language Models from Southeast Asia
by: Gamboa, Lance Calvin Lim, et al.
Published: (2024)
by: Gamboa, Lance Calvin Lim, et al.
Published: (2024)
HW-MLVQA: Elucidating Multilingual Handwritten Document Understanding with a Comprehensive VQA Benchmark
by: Pal, Aniket, et al.
Published: (2025)
by: Pal, Aniket, et al.
Published: (2025)
Transformer-empowered Multi-modal Item Embedding for Enhanced Image Search in E-Commerce
by: Liu, Chang, et al.
Published: (2023)
by: Liu, Chang, et al.
Published: (2023)
JaWildText: A Benchmark for Vision-Language Models on Japanese Scene Text Understanding
by: Maeda, Koki, et al.
Published: (2026)
by: Maeda, Koki, et al.
Published: (2026)
Learning from Medical Entity Trees: An Entity-Centric Medical Data Engineering Framework for MLLMs
by: Lin, Jianghang, et al.
Published: (2026)
by: Lin, Jianghang, et al.
Published: (2026)
SceneVTG++: Controllable Multilingual Visual Text Generation in the Wild
by: Liu, Jiawei, et al.
Published: (2025)
by: Liu, Jiawei, et al.
Published: (2025)
NeRF-DetS: Enhanced Adaptive Spatial-wise Sampling and View-wise Fusion Strategies for NeRF-based Indoor Multi-view 3D Object Detection
by: Huang, Chi, et al.
Published: (2024)
by: Huang, Chi, et al.
Published: (2024)
Similar Items
-
Referring Industrial Anomaly Segmentation
by: Yue, Pengfei, et al.
Published: (2026) -
Active-SAOOD: Active Sparsely Annotated Oriented Object Detection in Remote Sensing Images
by: Lin, Yu, et al.
Published: (2026) -
What You Perceive Is What You Conceive: A Cognition-Inspired Framework for Open Vocabulary Image Segmentation
by: Lin, Jianghang, et al.
Published: (2025) -
Pseudo-Label Quality Decoupling and Correction for Semi-Supervised Instance Segmentation
by: Lin, Jianghang, et al.
Published: (2025) -
GOI: Find 3D Gaussians of Interest with an Optimizable Open-vocabulary Semantic-space Hyperplane
by: Qu, Yansong, et al.
Published: (2024)