DRISHTIKON: Visual Grounding at Multiple Granularities in Documents
Fuente:
arXiv
Saved in:
| Main Authors: | Kasuba, Badri Vishal, Chaudhuri, Parag, Ramakrishnan, Ganesh |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PLATTER: A Page-Level Handwritten Text Recognition System for Indic Scripts
by: Kasuba, Badri Vishal, et al.
Published: (2025)
by: Kasuba, Badri Vishal, et al.
Published: (2025)
SPRINT: Script-agnostic Structure Recognition in Tables
by: Kudale, Dhruv, et al.
Published: (2025)
by: Kudale, Dhruv, et al.
Published: (2025)
TEXTRON: Weakly Supervised Multilingual Text Detection through Data Programming
by: Kudale, Dhruv, et al.
Published: (2024)
by: Kudale, Dhruv, et al.
Published: (2024)
Can AI Assistance Aid in the Grading of Handwritten Answer Sheets?
by: Sil, Pritam, et al.
Published: (2024)
by: Sil, Pritam, et al.
Published: (2024)
Uncertainty-Aware Subset Selection for Robust Visual Explainability under Distribution Shifts
by: Gupta, Madhav, et al.
Published: (2025)
by: Gupta, Madhav, et al.
Published: (2025)
Seeing to Ground: Visual Attention for Hallucination-Resilient MDLLMs
by: Narnaware, Vishal, et al.
Published: (2026)
by: Narnaware, Vishal, et al.
Published: (2026)
GroundingSuite: Measuring Complex Multi-Granular Pixel Grounding
by: Hu, Rui, et al.
Published: (2025)
by: Hu, Rui, et al.
Published: (2025)
A Multi-Granularity Retrieval Framework for Visually-Rich Documents
by: Xu, Mingjun, et al.
Published: (2025)
by: Xu, Mingjun, et al.
Published: (2025)
Learning Disentangled Representation in Object-Centric Models for Visual Dynamics Prediction via Transformers
by: Gandhi, Sanket, et al.
Published: (2024)
by: Gandhi, Sanket, et al.
Published: (2024)
PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination
by: Dai, Ming, et al.
Published: (2025)
by: Dai, Ming, et al.
Published: (2025)
TruthLens: Visual Grounding for Universal DeepFake Reasoning
by: Kundu, Rohit, et al.
Published: (2025)
by: Kundu, Rohit, et al.
Published: (2025)
Early Exit and Multi Stage Knowledge Distillation in VLMs for Video Summarization
by: Khan, Anas Anwarul Haq, et al.
Published: (2025)
by: Khan, Anas Anwarul Haq, et al.
Published: (2025)
IQA-Spider: Unifying Multi-Granularity Image Quality Assessment with Reasoning, Grounding and Referring
by: Peng, Xinge, et al.
Published: (2026)
by: Peng, Xinge, et al.
Published: (2026)
Modality Agnostic Efficient Long Range Encoder
by: Parag, Toufiq, et al.
Published: (2025)
by: Parag, Toufiq, et al.
Published: (2025)
Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding
by: Xiao, Han, et al.
Published: (2025)
by: Xiao, Han, et al.
Published: (2025)
INSITE: labelling medical images using submodular functions and semi-supervised data programming
by: Gautam, Akshat, et al.
Published: (2024)
by: Gautam, Akshat, et al.
Published: (2024)
RemoteVAR: Autoregressive Visual Modeling for Remote Sensing Change Detection
by: Korkmaz, Yilmaz, et al.
Published: (2026)
by: Korkmaz, Yilmaz, et al.
Published: (2026)
Boosting Medical Visual Understanding From Multi-Granular Language Learning
by: Li, Zihan, et al.
Published: (2025)
by: Li, Zihan, et al.
Published: (2025)
MG-LLaVA: Towards Multi-Granularity Visual Instruction Tuning
by: Zhao, Xiangyu, et al.
Published: (2024)
by: Zhao, Xiangyu, et al.
Published: (2024)
DOGR: Towards Versatile Visual Document Grounding and Referring
by: Zhou, Yinan, et al.
Published: (2024)
by: Zhou, Yinan, et al.
Published: (2024)
Attention Grounded Enhancement for Visual Document Retrieval
by: Cui, Wanqing, et al.
Published: (2025)
by: Cui, Wanqing, et al.
Published: (2025)
AWRaCLe: All-Weather Image Restoration using Visual In-Context Learning
by: Rajagopalan, Sudarshan, et al.
Published: (2024)
by: Rajagopalan, Sudarshan, et al.
Published: (2024)
Stress Tests REVEAL Fragile Temporal and Visual Grounding in Video-Language Models
by: T V, Sethuraman, et al.
Published: (2026)
by: T V, Sethuraman, et al.
Published: (2026)
Multiple Consistent 2D-3D Mappings for Robust Zero-Shot 3D Visual Grounding
by: Yin, Yufei, et al.
Published: (2026)
by: Yin, Yufei, et al.
Published: (2026)
MG-MotionLLM: A Unified Framework for Motion Comprehension and Generation across Multiple Granularities
by: Wu, Bizhu, et al.
Published: (2025)
by: Wu, Bizhu, et al.
Published: (2025)
On the Role of Visual Grounding in VQA
by: Reich, Daniel, et al.
Published: (2024)
by: Reich, Daniel, et al.
Published: (2024)
ChangingGrounding: 3D Visual Grounding in Changing Scenes
by: Hu, Miao, et al.
Published: (2025)
by: Hu, Miao, et al.
Published: (2025)
Direct Visual Grounding by Directing Attention of Visual Tokens
by: Esmaeilkhani, Parsa, et al.
Published: (2025)
by: Esmaeilkhani, Parsa, et al.
Published: (2025)
PhysLab: A Benchmark Dataset for Multi-Granularity Visual Parsing of Physics Experiments
by: Zou, Minghao, et al.
Published: (2025)
by: Zou, Minghao, et al.
Published: (2025)
Towards Unified Referring Expression Segmentation Across Omni-Level Visual Target Granularities
by: Liu, Jing, et al.
Published: (2025)
by: Liu, Jing, et al.
Published: (2025)
Granulon: Awakening Pixel-Level Visual Encoders with Adaptive Multi-Granularity Semantics for MLLM
by: Mao, Junyuan, et al.
Published: (2026)
by: Mao, Junyuan, et al.
Published: (2026)
RadioFormer: A Multiple-Granularity Radio Map Estimation Transformer with 1\textpertenthousand Spatial Sampling
by: Fang, Zheng, et al.
Published: (2025)
by: Fang, Zheng, et al.
Published: (2025)
ResVG: Enhancing Relation and Semantic Understanding in Multiple Instances for Visual Grounding
by: Zheng, Minghang, et al.
Published: (2024)
by: Zheng, Minghang, et al.
Published: (2024)
VGGT: Visual Geometry Grounded Transformer
by: Wang, Jianyuan, et al.
Published: (2025)
by: Wang, Jianyuan, et al.
Published: (2025)
Grounded Reinforcement Learning for Visual Reasoning
by: Sarch, Gabriel, et al.
Published: (2025)
by: Sarch, Gabriel, et al.
Published: (2025)
Visual Intention Grounding for Egocentric Assistants
by: Sun, Pengzhan, et al.
Published: (2025)
by: Sun, Pengzhan, et al.
Published: (2025)
Quantized Visual Geometry Grounded Transformer
by: Feng, Weilun, et al.
Published: (2025)
by: Feng, Weilun, et al.
Published: (2025)
Towards Visual Grounding: A Survey
by: Xiao, Linhui, et al.
Published: (2024)
by: Xiao, Linhui, et al.
Published: (2024)
Context-Infused Visual Grounding for Art
by: Khan, Selina, et al.
Published: (2024)
by: Khan, Selina, et al.
Published: (2024)
MVP: Multiple View Prediction Improves GUI Grounding
by: Zhang, Yunzhu, et al.
Published: (2025)
by: Zhang, Yunzhu, et al.
Published: (2025)
Similar Items
-
PLATTER: A Page-Level Handwritten Text Recognition System for Indic Scripts
by: Kasuba, Badri Vishal, et al.
Published: (2025) -
SPRINT: Script-agnostic Structure Recognition in Tables
by: Kudale, Dhruv, et al.
Published: (2025) -
TEXTRON: Weakly Supervised Multilingual Text Detection through Data Programming
by: Kudale, Dhruv, et al.
Published: (2024) -
Can AI Assistance Aid in the Grading of Handwritten Answer Sheets?
by: Sil, Pritam, et al.
Published: (2024) -
Uncertainty-Aware Subset Selection for Robust Visual Explainability under Distribution Shifts
by: Gupta, Madhav, et al.
Published: (2025)