StrucTexTv3: An Efficient Vision-Language Model for Text-rich Image Perception, Comprehension, and Beyond
Fuente:
arXiv
Saved in:
| Main Authors: | Lyu, Pengyuan, Li, Yulin, Zhou, Hao, Ma, Weihong, Wan, Xingyu, Xie, Qunyi, Wu, Liang, Zhang, Chengquan, Yao, Kun, Ding, Errui, Wang, Jingdong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Unified Multi-granularity Text Detection with Interactive Attention
by: Wan, Xingyu, et al.
Published: (2024)
by: Wan, Xingyu, et al.
Published: (2024)
MMTIT-Bench: A Multilingual and Multi-Scenario Benchmark with Cognition-Perception-Reasoning Guided Text-Image Machine Translation
by: Li, Gengluo, et al.
Published: (2026)
by: Li, Gengluo, et al.
Published: (2026)
TexRO: Generating Delicate Textures of 3D Models by Recursive Optimization
by: Wu, Jinbo, et al.
Published: (2024)
by: Wu, Jinbo, et al.
Published: (2024)
Prune Redundancy, Preserve Essence: Vision Token Compression in VLMs via Synergistic Importance-Diversity
by: Fang, Zhengyao, et al.
Published: (2026)
by: Fang, Zhengyao, et al.
Published: (2026)
Recognition-Synergistic Scene Text Editing
by: Fang, Zhengyao, et al.
Published: (2025)
by: Fang, Zhengyao, et al.
Published: (2025)
FullAnno: A Data Engine for Enhancing Image Comprehension of MLLMs
by: Hao, Jing, et al.
Published: (2024)
by: Hao, Jing, et al.
Published: (2024)
MicroViTv2: Beyond the FLOPS for Edge Energy-Friendly Vision Transformers
by: Setyawan, Novendra, et al.
Published: (2026)
by: Setyawan, Novendra, et al.
Published: (2026)
WeCromCL: Weakly Supervised Cross-Modality Contrastive Learning for Transcription-only Supervised Text Spotting
by: Wu, Jingjing, et al.
Published: (2024)
by: Wu, Jingjing, et al.
Published: (2024)
TopoSD: Topology-Enhanced Lane Segment Perception with SDMap Prior
by: Yang, Sen, et al.
Published: (2024)
by: Yang, Sen, et al.
Published: (2024)
Towards Real-World Document Parsing via Realistic Scene Synthesis and Document-Aware Training
by: Li, Gengluo, et al.
Published: (2026)
by: Li, Gengluo, et al.
Published: (2026)
Interpretable Face Anti-Spoofing: Enhancing Generalization with Multimodal Large Language Models
by: Zhang, Guosheng, et al.
Published: (2025)
by: Zhang, Guosheng, et al.
Published: (2025)
MS-DETR: Efficient DETR Training with Mixed Supervision
by: Zhao, Chuyang, et al.
Published: (2024)
by: Zhao, Chuyang, et al.
Published: (2024)
StrucText-Eval: Evaluating Large Language Model's Reasoning Ability in Structure-Rich Text
by: Gu, Zhouhong, et al.
Published: (2024)
by: Gu, Zhouhong, et al.
Published: (2024)
FiTv2: Scalable and Improved Flexible Vision Transformer for Diffusion Model
by: Wang, ZiDong, et al.
Published: (2024)
by: Wang, ZiDong, et al.
Published: (2024)
TexEditor: Structure-Preserving Text-Driven Texture Editing
by: Zhao, Bo, et al.
Published: (2026)
by: Zhao, Bo, et al.
Published: (2026)
Skim then Focus: Integrating Contextual and Fine-grained Views for Repetitive Action Counting
by: Zhao, Zhengqi, et al.
Published: (2024)
by: Zhao, Zhengqi, et al.
Published: (2024)
OVLW-DETR: Open-Vocabulary Light-Weighted Detection Transformer
by: Wang, Yu, et al.
Published: (2024)
by: Wang, Yu, et al.
Published: (2024)
TexGaussian: Generating High-quality PBR Material via Octree-based 3D Gaussian Splatting
by: Xiong, Bojun, et al.
Published: (2024)
by: Xiong, Bojun, et al.
Published: (2024)
GenesisTex2: Stable, Consistent and High-Quality Text-to-Texture Generation
by: Lu, Jiawei, et al.
Published: (2024)
by: Lu, Jiawei, et al.
Published: (2024)
Add-SD: Rational Generation without Manual Reference
by: Yang, Lingfeng, et al.
Published: (2024)
by: Yang, Lingfeng, et al.
Published: (2024)
VDG: Vision-Only Dynamic Gaussian for Driving Simulation
by: Li, Hao, et al.
Published: (2024)
by: Li, Hao, et al.
Published: (2024)
Struc-EMB: The Potential of Structure-Aware Encoding in Language Embeddings
by: Liu, Shikun, et al.
Published: (2025)
by: Liu, Shikun, et al.
Published: (2025)
CODER: Coupled Diversity-Sensitive Momentum Contrastive Learning for Image-Text Retrieval
by: Wang, Haoran, et al.
Published: (2022)
by: Wang, Haoran, et al.
Published: (2022)
ALoRE: Efficient Visual Adaptation via Aggregating Low Rank Experts
by: Du, Sinan, et al.
Published: (2024)
by: Du, Sinan, et al.
Published: (2024)
StrucSum: Graph-Structured Reasoning for Long Document Extractive Summarization with LLMs
by: Yuan, Haohan, et al.
Published: (2025)
by: Yuan, Haohan, et al.
Published: (2025)
MGMapNet: Multi-Granularity Representation Learning for End-to-End Vectorized HD Map Construction
by: Yang, Jing, et al.
Published: (2024)
by: Yang, Jing, et al.
Published: (2024)
MonoFormer: One Transformer for Both Diffusion and Autoregression
by: Zhao, Chuyang, et al.
Published: (2024)
by: Zhao, Chuyang, et al.
Published: (2024)
Weak H‐Bonds of Fluorophenyl Synthons Enabling Efficient and Entropically Stabilized Supramolecular Electro‐Optic Dendrimers
by: Di Zhang, et al.
Published: (2025)
by: Di Zhang, et al.
Published: (2025)
Struc-Bench: Are Large Language Models Really Good at Generating Complex Structured Data?
by: Tang, Xiangru, et al.
Published: (2023)
by: Tang, Xiangru, et al.
Published: (2023)
3DRS: MLLMs Need 3D-Aware Representation Supervision for Scene Understanding
by: Huang, Xiaohu, et al.
Published: (2025)
by: Huang, Xiaohu, et al.
Published: (2025)
TexTailor: Customized Text-aligned Texturing via Effective Resampling
by: Lee, Suin, et al.
Published: (2025)
by: Lee, Suin, et al.
Published: (2025)
TexPro: Text-guided PBR Texturing with Procedural Material Modeling
by: Dang, Ziqiang, et al.
Published: (2024)
by: Dang, Ziqiang, et al.
Published: (2024)
MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech
by: Chen, Huakang, et al.
Published: (2026)
by: Chen, Huakang, et al.
Published: (2026)
StrucADT: Generating Structure-controlled 3D Point Clouds with Adjacency Diffusion Transformer
by: Shu, Zhenyu, et al.
Published: (2025)
by: Shu, Zhenyu, et al.
Published: (2025)
Rubber Tube–Based Triboelectric Nanogenerator for Simultaneous Energy Harvesting and Real‐Time Health Monitoring in Taekwondo Athletes
by: Chengquan Piao, et al.
Published: (2026)
by: Chengquan Piao, et al.
Published: (2026)
TexAVi: Generating Stereoscopic VR Video Clips from Text Descriptions
by: Srihari, Vriksha, et al.
Published: (2025)
by: Srihari, Vriksha, et al.
Published: (2025)
TexLiDAR: Automated Text Understanding for Panoramic LiDAR Data
by: Cohen, Naor, et al.
Published: (2025)
by: Cohen, Naor, et al.
Published: (2025)
Perception Before Reasoning: Two-Stage Reinforcement Learning for Visual Reasoning in Vision-Language Models
by: Chen, Yan, et al.
Published: (2025)
by: Chen, Yan, et al.
Published: (2025)
Cross from Left to Right Brain: Adaptive Text Dreamer for Vision-and-Language Navigation
by: Zhang, Pingrui, et al.
Published: (2025)
by: Zhang, Pingrui, et al.
Published: (2025)
Gradient-based Sampling for Class Imbalanced Semi-supervised Object Detection
by: Li, Jiaming, et al.
Published: (2024)
by: Li, Jiaming, et al.
Published: (2024)
Similar Items
-
Towards Unified Multi-granularity Text Detection with Interactive Attention
by: Wan, Xingyu, et al.
Published: (2024) -
MMTIT-Bench: A Multilingual and Multi-Scenario Benchmark with Cognition-Perception-Reasoning Guided Text-Image Machine Translation
by: Li, Gengluo, et al.
Published: (2026) -
TexRO: Generating Delicate Textures of 3D Models by Recursive Optimization
by: Wu, Jinbo, et al.
Published: (2024) -
Prune Redundancy, Preserve Essence: Vision Token Compression in VLMs via Synergistic Importance-Diversity
by: Fang, Zhengyao, et al.
Published: (2026) -
Recognition-Synergistic Scene Text Editing
by: Fang, Zhengyao, et al.
Published: (2025)