InstructTable: Improving Table Structure Recognition Through Instructions
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Boming, Wang, Zining, Guo, Zhentao, Liu, Jianqiang, Duan, Chen, Gu, Yu, zhou, Kai, Yan, Pengfei |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PositionOCR: Augmenting Positional Awareness in Multi-Modal Models via Hybrid Specialist Integration
by: Duan, Chen, et al.
Published: (2026)
by: Duan, Chen, et al.
Published: (2026)
InstructOCR: Instruction Boosting Scene Text Spotting
by: Duan, Chen, et al.
Published: (2024)
by: Duan, Chen, et al.
Published: (2024)
Multimodal Large Language Models for Text-rich Image Understanding: A Comprehensive Review
by: Fu, Pei, et al.
Published: (2025)
by: Fu, Pei, et al.
Published: (2025)
Instruct-CLIP: Improving Instruction-Guided Image Editing with Automated Data Refinement Using Contrastive Learning
by: Chen, Sherry X., et al.
Published: (2025)
by: Chen, Sherry X., et al.
Published: (2025)
UniTabNet: Bridging Vision and Language Models for Enhanced Table Structure Recognition
by: Zhang, Zhenrong, et al.
Published: (2024)
by: Zhang, Zhenrong, et al.
Published: (2024)
Marten: Visual Question Answering with Mask Generation for Multi-modal Document Understanding
by: Wang, Zining, et al.
Published: (2025)
by: Wang, Zining, et al.
Published: (2025)
InstructSAM: A Training-Free Framework for Instruction-Oriented Remote Sensing Object Recognition
by: Zheng, Yijie, et al.
Published: (2025)
by: Zheng, Yijie, et al.
Published: (2025)
InstructVEdit: A Holistic Approach for Instructional Video Editing
by: Zhang, Chi, et al.
Published: (2025)
by: Zhang, Chi, et al.
Published: (2025)
InstructEngine: Instruction-driven Text-to-Image Alignment
by: Lu, Xingyu, et al.
Published: (2025)
by: Lu, Xingyu, et al.
Published: (2025)
InstructRestore: Region-Customized Image Restoration with Human Instructions
by: Liu, Shuaizheng, et al.
Published: (2025)
by: Liu, Shuaizheng, et al.
Published: (2025)
Uncertainty-Instructed Structure Injection for Generalizable HD Map Construction
by: Liu, Xiaolu, et al.
Published: (2025)
by: Liu, Xiaolu, et al.
Published: (2025)
OmniParser: A Unified Framework for Text Spotting, Key Information Extraction and Table Recognition
by: Wan, Jianqiang, et al.
Published: (2024)
by: Wan, Jianqiang, et al.
Published: (2024)
Instruct-ReID: A Multi-purpose Person Re-identification Task with Instructions
by: He, Weizhen, et al.
Published: (2023)
by: He, Weizhen, et al.
Published: (2023)
InstructUDrag: Joint Text Instructions and Object Dragging for Interactive Image Editing
by: Yu, Haoran, et al.
Published: (2025)
by: Yu, Haoran, et al.
Published: (2025)
InstructSAM: Segment Any Instance with Any Instructions
by: Yuan, Yuqian, et al.
Published: (2026)
by: Yuan, Yuqian, et al.
Published: (2026)
InstructBrush: Learning Attention-based Instruction Optimization for Image Editing
by: Zhao, Ruoyu, et al.
Published: (2024)
by: Zhao, Ruoyu, et al.
Published: (2024)
InstructDET: Diversifying Referring Object Detection with Generalized Instructions
by: Dang, Ronghao, et al.
Published: (2023)
by: Dang, Ronghao, et al.
Published: (2023)
Uncertainty Quantification in Table Structure Recognition
by: Ajayi, Kehinde, et al.
Published: (2024)
by: Ajayi, Kehinde, et al.
Published: (2024)
Efficient Multi-branch Segmentation Network for Situation Awareness in Autonomous Navigation
by: Zhou, Guan-Cheng, et al.
Published: (2024)
by: Zhou, Guan-Cheng, et al.
Published: (2024)
InstructAttribute: Fine-grained Object Attributes editing with Instruction
by: Yin, Xingxi, et al.
Published: (2025)
by: Yin, Xingxi, et al.
Published: (2025)
Instruction-Guided Scene Text Recognition
by: Du, Yongkun, et al.
Published: (2024)
by: Du, Yongkun, et al.
Published: (2024)
Instruct-ReID++: Towards Universal Purpose Instruction-Guided Person Re-identification
by: He, Weizhen, et al.
Published: (2024)
by: He, Weizhen, et al.
Published: (2024)
SPRINT: Script-agnostic Structure Recognition in Tables
by: Kudale, Dhruv, et al.
Published: (2025)
by: Kudale, Dhruv, et al.
Published: (2025)
InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation
by: Yang, Shuai, et al.
Published: (2025)
by: Yang, Shuai, et al.
Published: (2025)
Instruct-Imagen: Image Generation with Multi-modal Instruction
by: Hu, Hexiang, et al.
Published: (2024)
by: Hu, Hexiang, et al.
Published: (2024)
VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search
by: Jia, Yiming, et al.
Published: (2025)
by: Jia, Yiming, et al.
Published: (2025)
LATTE: Improving Latex Recognition for Tables and Formulae with Iterative Refinement
by: Jiang, Nan, et al.
Published: (2024)
by: Jiang, Nan, et al.
Published: (2024)
TIGER: Text-Instructed 3D Gaussian Retrieval and Coherent Editing
by: Xu, Teng, et al.
Published: (2024)
by: Xu, Teng, et al.
Published: (2024)
TC-OCR: TableCraft OCR for Efficient Detection & Recognition of Table Structure & Content
by: Anand, Avinash, et al.
Published: (2024)
by: Anand, Avinash, et al.
Published: (2024)
InstructCV: Instruction-Tuned Text-to-Image Diffusion Models as Vision Generalists
by: Gan, Yulu, et al.
Published: (2023)
by: Gan, Yulu, et al.
Published: (2023)
InstructTA: Instruction-Tuned Targeted Attack for Large Vision-Language Models
by: Wang, Xunguang, et al.
Published: (2023)
by: Wang, Xunguang, et al.
Published: (2023)
InstructAV2AV: Instruction-Guided Audio-Video Joint Editing
by: Zheng, Haojie, et al.
Published: (2026)
by: Zheng, Haojie, et al.
Published: (2026)
LORE++: Logical Location Regression Network for Table Structure Recognition with Pre-training
by: Long, Rujiao, et al.
Published: (2024)
by: Long, Rujiao, et al.
Published: (2024)
InstructRL4Pix: Training Diffusion for Image Editing by Reinforcement Learning
by: Li, Tiancheng, et al.
Published: (2024)
by: Li, Tiancheng, et al.
Published: (2024)
InstructHumans: Editing Animated 3D Human Textures with Instructions
by: Zhu, Jiayin, et al.
Published: (2024)
by: Zhu, Jiayin, et al.
Published: (2024)
Self-Supervised Pre-Training for Table Structure Recognition Transformer
by: Peng, ShengYun, et al.
Published: (2024)
by: Peng, ShengYun, et al.
Published: (2024)
TFLOP: Table Structure Recognition Framework with Layout Pointer Mechanism
by: Khang, Minsoo, et al.
Published: (2025)
by: Khang, Minsoo, et al.
Published: (2025)
VICI: VLM-Instructed Cross-view Image-localisation
by: Zhang, Xiaohan, et al.
Published: (2025)
by: Zhang, Xiaohan, et al.
Published: (2025)
Instruct2See: Learning to Remove Any Obstructions Across Distributions
by: Li, Junhang, et al.
Published: (2025)
by: Li, Junhang, et al.
Published: (2025)
InstructAny2Pix: Flexible Visual Editing via Multimodal Instruction Following
by: Li, Shufan, et al.
Published: (2023)
by: Li, Shufan, et al.
Published: (2023)
Similar Items
-
PositionOCR: Augmenting Positional Awareness in Multi-Modal Models via Hybrid Specialist Integration
by: Duan, Chen, et al.
Published: (2026) -
InstructOCR: Instruction Boosting Scene Text Spotting
by: Duan, Chen, et al.
Published: (2024) -
Multimodal Large Language Models for Text-rich Image Understanding: A Comprehensive Review
by: Fu, Pei, et al.
Published: (2025) -
Instruct-CLIP: Improving Instruction-Guided Image Editing with Automated Data Refinement Using Contrastive Learning
by: Chen, Sherry X., et al.
Published: (2025) -
UniTabNet: Bridging Vision and Language Models for Enhanced Table Structure Recognition
by: Zhang, Zhenrong, et al.
Published: (2024)