Attend to what I say: Highlighting relevant content on slides
Fuente:
arXiv
Saved in:
| Main Authors: | M, Megha Mariam K, Jawahar, C. V. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unifying Scientific Communication: Fine-Grained Correspondence Across Scientific Media
by: M, Megha Mariam K., et al.
Published: (2026)
by: M, Megha Mariam K., et al.
Published: (2026)
PhyEduVideo: A Benchmark for Evaluating Text-to-Video Models for Physics Education
by: M, Megha Mariam K., et al.
Published: (2026)
by: M, Megha Mariam K., et al.
Published: (2026)
Multiple Instance Learning for Glioma Diagnosis using Hematoxylin and Eosin Whole Slide Images: An Indian Cohort Study
by: Chauhan, Ekansh, et al.
Published: (2024)
by: Chauhan, Ekansh, et al.
Published: (2024)
Attend what matters: Leveraging vision foundational models for breast cancer classification using mammograms
by: Sanghvi, Samyak, et al.
Published: (2026)
by: Sanghvi, Samyak, et al.
Published: (2026)
IndicSTR12: A Dataset for Indic Scene Text Recognition
by: Lunia, Harsh, et al.
Published: (2024)
by: Lunia, Harsh, et al.
Published: (2024)
Advancing Question Answering on Handwritten Documents: A State-of-the-Art Recognition-Based Model for HW-SQuAD
by: Pal, Aniket, et al.
Published: (2024)
by: Pal, Aniket, et al.
Published: (2024)
Source-free Video Domain Adaptation by Learning from Noisy Labels
by: Dasgupta, Avijit, et al.
Published: (2023)
by: Dasgupta, Avijit, et al.
Published: (2023)
Benford's law: what does it say on adversarial images?
by: Zago, João G., et al.
Published: (2021)
by: Zago, João G., et al.
Published: (2021)
Prompt2LVideos: Exploring Prompts for Understanding Long-Form Multimodal Videos
by: Jahagirdar, Soumya Shamarao, et al.
Published: (2025)
by: Jahagirdar, Soumya Shamarao, et al.
Published: (2025)
Real-Time Human Reconstruction and Animation using Feed-Forward Gaussian Splatting
by: Chatterjee, Devdoot, et al.
Published: (2026)
by: Chatterjee, Devdoot, et al.
Published: (2026)
How Does India Cook Biryani?
by: Goel, Shubham, et al.
Published: (2026)
by: Goel, Shubham, et al.
Published: (2026)
EW-DETR: Evolving World Object Detection via Incremental Low-Rank DEtection TRansformer
by: Monga, Munish, et al.
Published: (2026)
by: Monga, Munish, et al.
Published: (2026)
HW-MLVQA: Elucidating Multilingual Handwritten Document Understanding with a Comprehensive VQA Benchmark
by: Pal, Aniket, et al.
Published: (2025)
by: Pal, Aniket, et al.
Published: (2025)
Attend to Not Attended: Structure-then-Detail Token Merging for Post-training DiT Acceleration
by: Fang, Haipeng, et al.
Published: (2025)
by: Fang, Haipeng, et al.
Published: (2025)
PEDESTRIANQA: A Benchmark for Vision-Language Models on Pedestrian Intention and Trajectory Prediction
by: Mishra, Naman, et al.
Published: (2026)
by: Mishra, Naman, et al.
Published: (2026)
Locality-Attending Vision Transformer
by: Hajimiri, Sina, et al.
Published: (2026)
by: Hajimiri, Sina, et al.
Published: (2026)
Face Time Traveller : Travel Through Ages Without Losing Identity
by: Kar, Purbayan, et al.
Published: (2026)
by: Kar, Purbayan, et al.
Published: (2026)
Reading Between the Lanes: Text VideoQA on the Road
by: Tom, George, et al.
Published: (2023)
by: Tom, George, et al.
Published: (2023)
Can Reasons Help Improve Pedestrian Intent Estimation? A Cross-Modal Approach
by: Khindkar, Vaishnavi, et al.
Published: (2024)
by: Khindkar, Vaishnavi, et al.
Published: (2024)
Towards Deployable OCR models for Indic languages
by: Mathew, Minesh, et al.
Published: (2022)
by: Mathew, Minesh, et al.
Published: (2022)
HighlightMe: Detecting Highlights from Human-Centric Videos
by: Bhattacharya, Uttaran, et al.
Published: (2021)
by: Bhattacharya, Uttaran, et al.
Published: (2021)
GeneralAD: Anomaly Detection Across Domains by Attending to Distorted Features
by: Sträter, Luc P. J., et al.
Published: (2024)
by: Sträter, Luc P. J., et al.
Published: (2024)
DriveSafe: A Framework for Risk Detection and Safety Suggestions in Driving Scenarios
by: Artham, Sainithin, et al.
Published: (2026)
by: Artham, Sainithin, et al.
Published: (2026)
Gameplay Highlights Generation
by: Edithal, Vignesh, et al.
Published: (2025)
by: Edithal, Vignesh, et al.
Published: (2025)
Pedestrian Intention and Trajectory Prediction in Unstructured Traffic Using IDD-PeD
by: Bokkasam, Ruthvik, et al.
Published: (2025)
by: Bokkasam, Ruthvik, et al.
Published: (2025)
Attend and Enrich: Enhanced Visual Prompt for Zero-Shot Learning
by: Liu, Man, et al.
Published: (2024)
by: Liu, Man, et al.
Published: (2024)
Reasoning to Attend: Try to Understand How <SEG> Token Works
by: Qian, Rui, et al.
Published: (2024)
by: Qian, Rui, et al.
Published: (2024)
AI-Generated Lecture Slides for Improving Slide Element Detection and Retrieval
by: Maniyar, Suyash, et al.
Published: (2025)
by: Maniyar, Suyash, et al.
Published: (2025)
Attend Before Attention: Efficient and Scalable Video Understanding via Autoregressive Gazing
by: Shi, Baifeng, et al.
Published: (2026)
by: Shi, Baifeng, et al.
Published: (2026)
Where MLLMs Attend and What They Rely On: Explaining Autoregressive Token Generation
by: Chen, Ruoyu, et al.
Published: (2025)
by: Chen, Ruoyu, et al.
Published: (2025)
Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers
by: Ren, Sucheng, et al.
Published: (2025)
by: Ren, Sucheng, et al.
Published: (2025)
Perturb, Attend, Detect and Localize (PADL): Robust Proactive Image Defense
by: Bartolucci, Filippo, et al.
Published: (2024)
by: Bartolucci, Filippo, et al.
Published: (2024)
IDD-X: A Multi-View Dataset for Ego-relative Important Object Localization and Explanation in Dense and Unstructured Traffic
by: Parikh, Chirag, et al.
Published: (2024)
by: Parikh, Chirag, et al.
Published: (2024)
SVD-ViT: Does SVD Make Vision Transformers Attend More to the Foreground?
by: Murata, Haruhiko, et al.
Published: (2026)
by: Murata, Haruhiko, et al.
Published: (2026)
Diffuse, Attend, and Segment: Unsupervised Zero-Shot Segmentation using Stable Diffusion
by: Tian, Junjiao, et al.
Published: (2023)
by: Tian, Junjiao, et al.
Published: (2023)
KFFocus: Highlighting Keyframes for Enhanced Video Understanding
by: Nie, Ming, et al.
Published: (2025)
by: Nie, Ming, et al.
Published: (2025)
Attend to Evidence: Evidence-Anchored Spatial Attention Supervision for Multimodal RLVR
by: Hu, Ruina, et al.
Published: (2026)
by: Hu, Ruina, et al.
Published: (2026)
A Dataset for Semantic Segmentation in the Presence of Unknowns
by: Laskar, Zakaria, et al.
Published: (2025)
by: Laskar, Zakaria, et al.
Published: (2025)
Attend-and-Refine: Interactive keypoint estimation and quantitative cervical vertebrae analysis for bone age assessment
by: Kim, Jinhee, et al.
Published: (2025)
by: Kim, Jinhee, et al.
Published: (2025)
Show Me What I Like: Detecting User-Specific Video Highlights Using Content-Based Multi-Head Attention
by: Bhattacharya, Uttaran, et al.
Published: (2022)
by: Bhattacharya, Uttaran, et al.
Published: (2022)
Similar Items
-
Unifying Scientific Communication: Fine-Grained Correspondence Across Scientific Media
by: M, Megha Mariam K., et al.
Published: (2026) -
PhyEduVideo: A Benchmark for Evaluating Text-to-Video Models for Physics Education
by: M, Megha Mariam K., et al.
Published: (2026) -
Multiple Instance Learning for Glioma Diagnosis using Hematoxylin and Eosin Whole Slide Images: An Indian Cohort Study
by: Chauhan, Ekansh, et al.
Published: (2024) -
Attend what matters: Leveraging vision foundational models for breast cancer classification using mammograms
by: Sanghvi, Samyak, et al.
Published: (2026) -
IndicSTR12: A Dataset for Indic Scene Text Recognition
by: Lunia, Harsh, et al.
Published: (2024)