Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions
Fuente:
arXiv
Saved in:
| Main Authors: | Zong, Chang, Li, Bin, Zhou, Shoujun, Wan, Jian, Zhang, Lei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multimodal Integration Challenges in Emotionally Expressive Child Avatars for Training Applications
by: Salehi, Pegah, et al.
Published: (2025)
by: Salehi, Pegah, et al.
Published: (2025)
FreeAskWorld: An Interactive and Closed-Loop Simulator for Human-Centric Embodied AI
by: Peng, Yuhang, et al.
Published: (2025)
by: Peng, Yuhang, et al.
Published: (2025)
Inclusive AI for Group Interactions: Predicting Gaze-Direction Behaviors in People with Intellectual and Developmental Disabilities
by: Huang, Giulia, et al.
Published: (2026)
by: Huang, Giulia, et al.
Published: (2026)
TETRIS: Towards Exploring the Robustness of Interactive Segmentation
by: Moskalenko, Andrey, et al.
Published: (2024)
by: Moskalenko, Andrey, et al.
Published: (2024)
Attire-Based Anomaly Detection in Restricted Areas Using YOLOv8 for Enhanced CCTV Security
by: B, Abdul Aziz A., et al.
Published: (2024)
by: B, Abdul Aziz A., et al.
Published: (2024)
VitalLens 2.0: High-Fidelity rPPG for Heart Rate Variability Estimation from Face Video
by: Rouast, Philipp V.
Published: (2025)
by: Rouast, Philipp V.
Published: (2025)
BREPS: Bounding-Box Robustness Evaluation of Promptable Segmentation
by: Moskalenko, Andrey, et al.
Published: (2026)
by: Moskalenko, Andrey, et al.
Published: (2026)
In-Depth Analysis of Emotion Recognition through Knowledge-Based Large Language Models
by: Han, Bin, et al.
Published: (2024)
by: Han, Bin, et al.
Published: (2024)
Less Detail, Better Answers: Degradation-Driven Prompting for VQA
by: Han, Haoxuan, et al.
Published: (2026)
by: Han, Haoxuan, et al.
Published: (2026)
Embodied Cognition Augmented End2End Autonomous Driving
by: Niu, Ling, et al.
Published: (2025)
by: Niu, Ling, et al.
Published: (2025)
Motion Consistency Loss for Monocular Visual Odometry with Attention-Based Deep Learning
by: Françani, André O., et al.
Published: (2024)
by: Françani, André O., et al.
Published: (2024)
Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis
by: Salehi, Pegah, et al.
Published: (2024)
by: Salehi, Pegah, et al.
Published: (2024)
Human Motion Capture from Loose and Sparse Inertial Sensors with Garment-aware Diffusion Models
by: Ilic, Andela, et al.
Published: (2025)
by: Ilic, Andela, et al.
Published: (2025)
EgoPoser: Robust Real-Time Egocentric Pose Estimation from Sparse and Intermittent Observations Everywhere
by: Jiang, Jiaxi, et al.
Published: (2023)
by: Jiang, Jiaxi, et al.
Published: (2023)
Group Inertial Poser: Multi-Person Pose and Global Translation from Sparse Inertial Sensors and Ultra-Wideband Ranging
by: Xue, Ying, et al.
Published: (2025)
by: Xue, Ying, et al.
Published: (2025)
Z-Order Transformer for Feed-Forward Gaussian Splatting
by: Wang, Can, et al.
Published: (2026)
by: Wang, Can, et al.
Published: (2026)
Fast Data Aware Neural Architecture Search via Supernet Accelerated Evaluation
by: Njor, Emil, et al.
Published: (2025)
by: Njor, Emil, et al.
Published: (2025)
EV-CLIP: Efficient Visual Prompt Adaptation for CLIP in Few-shot Action Recognition under Visual Challenges
by: Jon, Hyo Jin, et al.
Published: (2026)
by: Jon, Hyo Jin, et al.
Published: (2026)
DancingBox: A Lightweight MoCap System for Character Animation from Physical Proxies
by: Yuan, Haocheng, et al.
Published: (2026)
by: Yuan, Haocheng, et al.
Published: (2026)
LRVS-Fashion: Extending Visual Search with Referring Instructions
by: Lepage, Simon, et al.
Published: (2023)
by: Lepage, Simon, et al.
Published: (2023)
Depth Priors in Removal Neural Radiance Fields
by: Guo, Zhihao, et al.
Published: (2024)
by: Guo, Zhihao, et al.
Published: (2024)
How Well Do Vision-Language Models Understand Sequential Driving Scenes? A Sensitivity Study
by: Brusnicki, Roberto, et al.
Published: (2026)
by: Brusnicki, Roberto, et al.
Published: (2026)
Nonverbal Immediacy Analysis in Education: A Multimodal Computational Model
by: Petković, Uroš, et al.
Published: (2024)
by: Petković, Uroš, et al.
Published: (2024)
Robust Perspective Correction for Real-World Crack Evolution Tracking in Image-Based Structural Health Monitoring
by: Sun, Xinxin, et al.
Published: (2025)
by: Sun, Xinxin, et al.
Published: (2025)
PAT-VCM: Plug-and-Play Auxiliary Tokens for Video Coding for Machines
by: Jiang, Wei, et al.
Published: (2026)
by: Jiang, Wei, et al.
Published: (2026)
Interactive Image Selection and Training for Brain Tumor Segmentation Network
by: Cerqueira, Matheus A., et al.
Published: (2024)
by: Cerqueira, Matheus A., et al.
Published: (2024)
MvBody: Multi-View-Based Hybrid Transformer Using Optical 3D Body Scan for Explainable Cesarean Section Prediction
by: Cheng, Ruting, et al.
Published: (2025)
by: Cheng, Ruting, et al.
Published: (2025)
MotionFollower: Editing Video Motion via Lightweight Score-Guided Diffusion
by: Tu, Shuyuan, et al.
Published: (2024)
by: Tu, Shuyuan, et al.
Published: (2024)
Performance Decay in Deepfake Detection: The Limitations of Training on Outdated Data
by: Richings, Jack, et al.
Published: (2025)
by: Richings, Jack, et al.
Published: (2025)
Facial Surgery Preview Based on the Orthognathic Treatment Prediction
by: Han, Huijun, et al.
Published: (2024)
by: Han, Huijun, et al.
Published: (2024)
SSD-Poser: Avatar Pose Estimation with State Space Duality from Sparse Observations
by: Zhao, Shuting, et al.
Published: (2025)
by: Zhao, Shuting, et al.
Published: (2025)
A Cost-Effective Eye-Tracker for Early Detection of Mild Cognitive Impairment
by: Greco, Danilo, et al.
Published: (2024)
by: Greco, Danilo, et al.
Published: (2024)
Spatial-ViLT: Enhancing Visual Spatial Reasoning through Multi-Task Learning
by: Islam, Chashi Mahiul, et al.
Published: (2025)
by: Islam, Chashi Mahiul, et al.
Published: (2025)
TALON: Test-time Adaptive Learning for On-the-Fly Category Discovery
by: Wu, Yanan, et al.
Published: (2026)
by: Wu, Yanan, et al.
Published: (2026)
EatGAN: An Edge-Attention Guided Generative Adversarial Network for Single Image Super-Resolution
by: Rao, Penghao, et al.
Published: (2025)
by: Rao, Penghao, et al.
Published: (2025)
TWIG: Two-Step Image Generation using Segmentation Masks in Diffusion Models
by: Rakib, Mazharul Islam, et al.
Published: (2025)
by: Rakib, Mazharul Islam, et al.
Published: (2025)
Isolated Sign Language Recognition with Segmentation and Pose Estimation
by: Perkins, Daniel, et al.
Published: (2025)
by: Perkins, Daniel, et al.
Published: (2025)
Sequence Transferability and Task Order Selection in Continual Learning
by: Nguyen, Thinh, et al.
Published: (2025)
by: Nguyen, Thinh, et al.
Published: (2025)
Towards Human Cognition Level-based Experiment Design for Counterfactual Explanations (XAI)
by: Suffian, Muhammad, et al.
Published: (2022)
by: Suffian, Muhammad, et al.
Published: (2022)
Socratic Planner: Self-QA-Based Zero-Shot Planning for Embodied Instruction Following
by: Shin, Suyeon, et al.
Published: (2024)
by: Shin, Suyeon, et al.
Published: (2024)
Similar Items
-
Multimodal Integration Challenges in Emotionally Expressive Child Avatars for Training Applications
by: Salehi, Pegah, et al.
Published: (2025) -
FreeAskWorld: An Interactive and Closed-Loop Simulator for Human-Centric Embodied AI
by: Peng, Yuhang, et al.
Published: (2025) -
Inclusive AI for Group Interactions: Predicting Gaze-Direction Behaviors in People with Intellectual and Developmental Disabilities
by: Huang, Giulia, et al.
Published: (2026) -
TETRIS: Towards Exploring the Robustness of Interactive Segmentation
by: Moskalenko, Andrey, et al.
Published: (2024) -
Attire-Based Anomaly Detection in Restricted Areas Using YOLOv8 for Enhanced CCTV Security
by: B, Abdul Aziz A., et al.
Published: (2024)