Unifying Scientific Communication: Fine-Grained Correspondence Across Scientific Media
Fuente:
arXiv
Saved in:
| Main Authors: | M, Megha Mariam K., Balasubramanian, Vineeth N., Jawahar, C. V. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Attend to what I say: Highlighting relevant content on slides
by: M, Megha Mariam K, et al.
Published: (2026)
by: M, Megha Mariam K, et al.
Published: (2026)
PhyEduVideo: A Benchmark for Evaluating Text-to-Video Models for Physics Education
by: M, Megha Mariam K., et al.
Published: (2026)
by: M, Megha Mariam K., et al.
Published: (2026)
Can Reasons Help Improve Pedestrian Intent Estimation? A Cross-Modal Approach
by: Khindkar, Vaishnavi, et al.
Published: (2024)
by: Khindkar, Vaishnavi, et al.
Published: (2024)
C2FDrone: Coarse-to-Fine Drone-to-Drone Detection using Vision Transformer Networks
by: Rebbapragada, Sairam VC, et al.
Published: (2024)
by: Rebbapragada, Sairam VC, et al.
Published: (2024)
Source-Free Domain Adaptation by Optimizing Batch-Wise Cosine Similarity
by: Pathak, Harsharaj, et al.
Published: (2026)
by: Pathak, Harsharaj, et al.
Published: (2026)
Efficient Vocabulary-Free Fine-Grained Visual Recognition in the Age of Multimodal LLMs
by: Kuchibhotla, Hari Chandana, et al.
Published: (2025)
by: Kuchibhotla, Hari Chandana, et al.
Published: (2025)
Multiple Instance Learning for Glioma Diagnosis using Hematoxylin and Eosin Whole Slide Images: An Indian Cohort Study
by: Chauhan, Ekansh, et al.
Published: (2024)
by: Chauhan, Ekansh, et al.
Published: (2024)
$\oslash$ Source Models Leak What They Shouldn't $\nrightarrow$: Unlearning Zero-Shot Transfer in Domain Adaptation Through Adversarial Optimization
by: Devalapally, Arnav, et al.
Published: (2026)
by: Devalapally, Arnav, et al.
Published: (2026)
Precise Event Spotting in Sports Videos: Solving Long-Range Dependency and Class Imbalance
by: Santra, Sanchayan, et al.
Published: (2025)
by: Santra, Sanchayan, et al.
Published: (2025)
iSHIFT: Lightweight Slow-Fast GUI Agent with Adaptive Perception
by: Mehrotra, Sarthak, et al.
Published: (2025)
by: Mehrotra, Sarthak, et al.
Published: (2025)
LogicCBMs: Logic-Enhanced Concept-Based Learning
by: Vemuri, Deepika SN, et al.
Published: (2025)
by: Vemuri, Deepika SN, et al.
Published: (2025)
An AI-Powered Autonomous Underwater System for Sea Exploration and Scientific Research
by: Almazrouei, Hamad, et al.
Published: (2025)
by: Almazrouei, Hamad, et al.
Published: (2025)
UniCorrn: Unified Correspondence Transformer Across 2D and 3D
by: Goswami, Prajnan, et al.
Published: (2026)
by: Goswami, Prajnan, et al.
Published: (2026)
Open-Set Object Detection By Aligning Known Class Representations
by: Sarkar, Hiran, et al.
Published: (2024)
by: Sarkar, Hiran, et al.
Published: (2024)
On Evaluation of Vision Datasets and Models using Human Competency Frameworks
by: Ramachandran, Rahul, et al.
Published: (2024)
by: Ramachandran, Rahul, et al.
Published: (2024)
Advancing Ante-Hoc Explainable Models through Generative Adversarial Networks
by: Garg, Tanmay, et al.
Published: (2024)
by: Garg, Tanmay, et al.
Published: (2024)
Understanding Task Transfer in Vision-Language Models
by: Sachdeva, Bhuvan, et al.
Published: (2025)
by: Sachdeva, Bhuvan, et al.
Published: (2025)
Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation
by: Choi, Jiho, et al.
Published: (2025)
by: Choi, Jiho, et al.
Published: (2025)
Can Better Text Semantics in Prompt Tuning Improve VLM Generalization?
by: Kuchibhotla, Hari Chandana, et al.
Published: (2024)
by: Kuchibhotla, Hari Chandana, et al.
Published: (2024)
IndicSTR12: A Dataset for Indic Scene Text Recognition
by: Lunia, Harsh, et al.
Published: (2024)
by: Lunia, Harsh, et al.
Published: (2024)
Advancing Question Answering on Handwritten Documents: A State-of-the-Art Recognition-Based Model for HW-SQuAD
by: Pal, Aniket, et al.
Published: (2024)
by: Pal, Aniket, et al.
Published: (2024)
Source-free Video Domain Adaptation by Learning from Noisy Labels
by: Dasgupta, Avijit, et al.
Published: (2023)
by: Dasgupta, Avijit, et al.
Published: (2023)
NoTeS-Bank: Benchmarking Neural Transcription and Search for Scientific Notes Understanding
by: Pal, Aniket, et al.
Published: (2025)
by: Pal, Aniket, et al.
Published: (2025)
Self-Supervised Spatial Correspondence Across Modalities
by: Shrivastava, Ayush, et al.
Published: (2025)
by: Shrivastava, Ayush, et al.
Published: (2025)
Prompt2LVideos: Exploring Prompts for Understanding Long-Form Multimodal Videos
by: Jahagirdar, Soumya Shamarao, et al.
Published: (2025)
by: Jahagirdar, Soumya Shamarao, et al.
Published: (2025)
Enhancing Scientific Visual Question Answering via Vision-Caption aware Supervised Fine-Tuning
by: Kapuriya, Janak, et al.
Published: (2025)
by: Kapuriya, Janak, et al.
Published: (2025)
Real-Time Human Reconstruction and Animation using Feed-Forward Gaussian Splatting
by: Chatterjee, Devdoot, et al.
Published: (2026)
by: Chatterjee, Devdoot, et al.
Published: (2026)
Chain-of-Thought Degrades Visual Spatial Reasoning Capabilities of Multimodal LLMs
by: Kancheti, Sai Srinivas, et al.
Published: (2026)
by: Kancheti, Sai Srinivas, et al.
Published: (2026)
How Does India Cook Biryani?
by: Goel, Shubham, et al.
Published: (2026)
by: Goel, Shubham, et al.
Published: (2026)
EW-DETR: Evolving World Object Detection via Incremental Low-Rank DEtection TRansformer
by: Monga, Munish, et al.
Published: (2026)
by: Monga, Munish, et al.
Published: (2026)
HW-MLVQA: Elucidating Multilingual Handwritten Document Understanding with a Comprehensive VQA Benchmark
by: Pal, Aniket, et al.
Published: (2025)
by: Pal, Aniket, et al.
Published: (2025)
Benchmarking Scientific Image Forgery Detectors
by: Cardenuto, João P., et al.
Published: (2021)
by: Cardenuto, João P., et al.
Published: (2021)
Toward Unified Fine-Grained Vehicle Classification and Automatic License Plate Recognition
by: Lima, Gabriel E., et al.
Published: (2026)
by: Lima, Gabriel E., et al.
Published: (2026)
Faithful GRPO: Improving Visual Spatial Reasoning in Multimodal Language Models via Constrained Policy Optimization
by: Kancheti, Sai Srinivas, et al.
Published: (2026)
by: Kancheti, Sai Srinivas, et al.
Published: (2026)
Mind's Eye: A Benchmark of Visual Abstraction, Transformation and Composition for Multimodal LLMs
by: Sinha, Rohit, et al.
Published: (2026)
by: Sinha, Rohit, et al.
Published: (2026)
A Fine-Grained Attention and Geometric Correspondence Model for Musculoskeletal Risk Classification in Athletes Using Multimodal Visual and Skeletal Features
by: Rahman, Md. Abdur, et al.
Published: (2025)
by: Rahman, Md. Abdur, et al.
Published: (2025)
Foundation Model Priors Enhance Object Focus in Feature Space for Source-Free Object Detection
by: VCR, Sairam, et al.
Published: (2025)
by: VCR, Sairam, et al.
Published: (2025)
PEDESTRIANQA: A Benchmark for Vision-Language Models on Pedestrian Intention and Trajectory Prediction
by: Mishra, Naman, et al.
Published: (2026)
by: Mishra, Naman, et al.
Published: (2026)
Semantic Correspondence: Unified Benchmarking and a Strong Baseline
by: Zhang, Kaiyan, et al.
Published: (2025)
by: Zhang, Kaiyan, et al.
Published: (2025)
Fiducial Focus Augmentation for Facial Landmark Detection
by: Kar, Purbayan, et al.
Published: (2024)
by: Kar, Purbayan, et al.
Published: (2024)
Similar Items
-
Attend to what I say: Highlighting relevant content on slides
by: M, Megha Mariam K, et al.
Published: (2026) -
PhyEduVideo: A Benchmark for Evaluating Text-to-Video Models for Physics Education
by: M, Megha Mariam K., et al.
Published: (2026) -
Can Reasons Help Improve Pedestrian Intent Estimation? A Cross-Modal Approach
by: Khindkar, Vaishnavi, et al.
Published: (2024) -
C2FDrone: Coarse-to-Fine Drone-to-Drone Detection using Vision Transformer Networks
by: Rebbapragada, Sairam VC, et al.
Published: (2024) -
Source-Free Domain Adaptation by Optimizing Batch-Wise Cosine Similarity
by: Pathak, Harsharaj, et al.
Published: (2026)