VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization
Fuente:
arXiv
Guardado en:
| Autores principales: | Nguyen, Son, Nguyen, Giang, Dao, Hung, Do, Thao, Kim, Daeyoung |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Improving Zero-Shot Object-Level Change Detection by Incorporating Visual Correspondence
por: Nguyen, Hung Huy, et al.
Publicado: (2025)
por: Nguyen, Hung Huy, et al.
Publicado: (2025)
ZSPAPrune: Zero-Shot Prompt-Aware Token Pruning for Vision-Language Models
por: Zhang, Pu, et al.
Publicado: (2025)
por: Zhang, Pu, et al.
Publicado: (2025)
LungCRCT: Causal Representation based Lung CT Processing for Lung Cancer Treatment
por: Kim, Daeyoung
Publicado: (2026)
por: Kim, Daeyoung
Publicado: (2026)
LightHCG: a Lightweight yet powerful HSIC Disentanglement based Causal Glaucoma Detection Model framework
por: Kim, Daeyoung
Publicado: (2025)
por: Kim, Daeyoung
Publicado: (2025)
Zoom-shot: Fast and Efficient Unsupervised Zero-Shot Transfer of CLIP to Vision Encoders with Multimodal Loss
por: Shipard, Jordan, et al.
Publicado: (2024)
por: Shipard, Jordan, et al.
Publicado: (2024)
See then Tell: Enhancing Key Information Extraction with Vision Grounding
por: Liu, Shuhang, et al.
Publicado: (2024)
por: Liu, Shuhang, et al.
Publicado: (2024)
Vision-Aware Text Features in Referring Image Segmentation: From Object Understanding to Context Understanding
por: Nguyen-Truong, Hai, et al.
Publicado: (2024)
por: Nguyen-Truong, Hai, et al.
Publicado: (2024)
Universal Multi-Domain Translation via Diffusion Routers
por: Kieu, Duc, et al.
Publicado: (2025)
por: Kieu, Duc, et al.
Publicado: (2025)
Binary Verification for Zero-Shot Vision
por: Hu, Rongbin, et al.
Publicado: (2025)
por: Hu, Rongbin, et al.
Publicado: (2025)
Frequency-Aware Token Reduction for Efficient Vision Transformer
por: Lee, Dong-Jae, et al.
Publicado: (2025)
por: Lee, Dong-Jae, et al.
Publicado: (2025)
Zero-Shot Vision-and-Language Navigation with Collision Mitigation in Continuous Environment
por: Jeong, Seongjun, et al.
Publicado: (2024)
por: Jeong, Seongjun, et al.
Publicado: (2024)
LeafNet: A Large-Scale Dataset and Comprehensive Benchmark for Foundational Vision-Language Understanding of Plant Diseases
por: Quoc, Khang Nguyen, et al.
Publicado: (2026)
por: Quoc, Khang Nguyen, et al.
Publicado: (2026)
LightX3ECG: A Lightweight and eXplainable Deep Learning System for 3-lead Electrocardiogram Classification
por: Le, Khiem H., et al.
Publicado: (2022)
por: Le, Khiem H., et al.
Publicado: (2022)
SwiftTry: Fast and Consistent Video Virtual Try-On with Diffusion Models
por: Nguyen, Hung, et al.
Publicado: (2024)
por: Nguyen, Hung, et al.
Publicado: (2024)
KGAlign: Joint Semantic-Structural Knowledge Encoding for Multimodal Fake News Detection
por: La, Tuan-Vinh, et al.
Publicado: (2025)
por: La, Tuan-Vinh, et al.
Publicado: (2025)
YoChameleon: Personalized Vision and Language Generation
por: Nguyen, Thao, et al.
Publicado: (2025)
por: Nguyen, Thao, et al.
Publicado: (2025)
Rotation-Aligned Key Channel Pruning for Efficient Vision-Language Model Inference
por: Kang, Beomseok, et al.
Publicado: (2026)
por: Kang, Beomseok, et al.
Publicado: (2026)
Zero-Shot Industrial Anomaly Segmentation with Image-Aware Prompt Generation
por: Park, SoYoung, et al.
Publicado: (2025)
por: Park, SoYoung, et al.
Publicado: (2025)
CSD-VAR: Content-Style Decomposition in Visual Autoregressive Models
por: Nguyen, Quang-Binh, et al.
Publicado: (2025)
por: Nguyen, Quang-Binh, et al.
Publicado: (2025)
Bidirectional Diffusion Bridge Models
por: Kieu, Duc, et al.
Publicado: (2025)
por: Kieu, Duc, et al.
Publicado: (2025)
PAT: Pixel-wise Adaptive Training for Long-tailed Segmentation
por: Do, Khoi, et al.
Publicado: (2024)
por: Do, Khoi, et al.
Publicado: (2024)
KRAST: Knowledge-Augmented Robotic Action Recognition with Structured Text for Vision-Language Models
por: Nguyen, Son Hai, et al.
Publicado: (2025)
por: Nguyen, Son Hai, et al.
Publicado: (2025)
PromptGuard: An Orchestrated Prompting Framework for Principled Synthetic Text Generation for Vulnerable Populations using LLMs with Enhanced Safety, Fairness, and Controllability
por: Vu, Tung, et al.
Publicado: (2025)
por: Vu, Tung, et al.
Publicado: (2025)
Anomaly-Aware Vision-Language Adapters for Zero-Shot Anomaly Detection
por: Aqeel, Muhammad, et al.
Publicado: (2026)
por: Aqeel, Muhammad, et al.
Publicado: (2026)
LangXAI: Integrating Large Vision Models for Generating Textual Explanations to Enhance Explainability in Visual Perception Tasks
por: Nguyen, Truong Thanh Hung, et al.
Publicado: (2024)
por: Nguyen, Truong Thanh Hung, et al.
Publicado: (2024)
Spatio-Temporal Context Prompting for Zero-Shot Action Detection
por: Huang, Wei-Jhe, et al.
Publicado: (2024)
por: Huang, Wei-Jhe, et al.
Publicado: (2024)
A Guideline-Aware AI Agent for Zero-Shot Target Volume Auto-Delineation
por: Kim, Yoon Jo, et al.
Publicado: (2026)
por: Kim, Yoon Jo, et al.
Publicado: (2026)
DuFal: Dual-Frequency-Aware Learning for High-Fidelity Extremely Sparse-view CBCT Reconstruction
por: Van, Cuong Tran, et al.
Publicado: (2026)
por: Van, Cuong Tran, et al.
Publicado: (2026)
STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models
por: Nguyen-Nhu, Tinh-Anh, et al.
Publicado: (2025)
por: Nguyen-Nhu, Tinh-Anh, et al.
Publicado: (2025)
SPATIOROUTE: Dynamic Prompt Routing for Zero-Shot Spatial Reasoning
por: Chunhachatrachai, Pawat, et al.
Publicado: (2026)
por: Chunhachatrachai, Pawat, et al.
Publicado: (2026)
SketchVLM: Vision language models can annotate images to explain thoughts and guide users
por: Collins, Brandon, et al.
Publicado: (2026)
por: Collins, Brandon, et al.
Publicado: (2026)
MasHeNe: A Benchmark for Head and Neck CT Mass Segmentation using Window-Enhanced Mamba with Frequency-Domain Integration
por: Dao, Thao Thi Phuong, et al.
Publicado: (2025)
por: Dao, Thao Thi Phuong, et al.
Publicado: (2025)
Enhancing the Fairness and Performance of Edge Cameras with Explainable AI
por: Nguyen, Truong Thanh Hung, et al.
Publicado: (2024)
por: Nguyen, Truong Thanh Hung, et al.
Publicado: (2024)
AC-MAMBASEG: An adaptive convolution and Mamba-based architecture for enhanced skin lesion segmentation
por: Nguyen, Viet-Thanh, et al.
Publicado: (2024)
por: Nguyen, Viet-Thanh, et al.
Publicado: (2024)
TinyVLM: Zero-Shot Object Detection on Microcontrollers via Vision-Language Distillation with Matryoshka Embeddings
por: Wilson, Bibin
Publicado: (2026)
por: Wilson, Bibin
Publicado: (2026)
DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization
por: Nguyen, Ngoc-Son, et al.
Publicado: (2026)
por: Nguyen, Ngoc-Son, et al.
Publicado: (2026)
Context-Aware Pseudo-Label Scoring for Zero-Shot Video Summarization
por: Wu, Yuanli, et al.
Publicado: (2025)
por: Wu, Yuanli, et al.
Publicado: (2025)
SalientFusion: Context-Aware Compositional Zero-Shot Food Recognition
por: Song, Jiajun, et al.
Publicado: (2025)
por: Song, Jiajun, et al.
Publicado: (2025)
GRAZE: Grounded Refinement and Motion-Aware Zero-Shot Event Localization
por: Zaidi, Syed Ahsan Masud, et al.
Publicado: (2026)
por: Zaidi, Syed Ahsan Masud, et al.
Publicado: (2026)
LightZeroNav: Zero-Shot Vision Language Navigation in Continuous Environments Based on Lightweight VLMs
por: Luo, Kun, et al.
Publicado: (2026)
por: Luo, Kun, et al.
Publicado: (2026)
Ejemplares similares
-
Improving Zero-Shot Object-Level Change Detection by Incorporating Visual Correspondence
por: Nguyen, Hung Huy, et al.
Publicado: (2025) -
ZSPAPrune: Zero-Shot Prompt-Aware Token Pruning for Vision-Language Models
por: Zhang, Pu, et al.
Publicado: (2025) -
LungCRCT: Causal Representation based Lung CT Processing for Lung Cancer Treatment
por: Kim, Daeyoung
Publicado: (2026) -
LightHCG: a Lightweight yet powerful HSIC Disentanglement based Causal Glaucoma Detection Model framework
por: Kim, Daeyoung
Publicado: (2025) -
Zoom-shot: Fast and Efficient Unsupervised Zero-Shot Transfer of CLIP to Vision Encoders with Multimodal Loss
por: Shipard, Jordan, et al.
Publicado: (2024)