CAPE: A CLIP-Aware Pointing Ensemble of Complementary Heatmap Cues for Embodied Reference Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Eyiokur, Fevziye Irem, Yaman, Dogucan, Ekenel, Hazım Kemal, Waibel, Alexander |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Multimodal Depth-Aware Method For Embodied Reference Understanding
by: Eyiokur, Fevziye Irem, et al.
Published: (2025)
by: Eyiokur, Fevziye Irem, et al.
Published: (2025)
Assessing Identity Leakage in Talking Face Generation: Metrics and Evaluation Framework
by: Yaman, Dogucan, et al.
Published: (2025)
by: Yaman, Dogucan, et al.
Published: (2025)
Audio-driven Talking Face Generation with Stabilized Synchronization Loss
by: Yaman, Dogucan, et al.
Published: (2023)
by: Yaman, Dogucan, et al.
Published: (2023)
Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation
by: Yaman, Dogucan, et al.
Published: (2025)
by: Yaman, Dogucan, et al.
Published: (2025)
Audio-Visual Speech Representation Expert for Enhanced Talking Face Video Generation and Evaluation
by: Yaman, Dogucan, et al.
Published: (2024)
by: Yaman, Dogucan, et al.
Published: (2024)
Shared Latent Representation for Joint Text-to-Audio-Visual Synthesis
by: Yaman, Dogucan, et al.
Published: (2025)
by: Yaman, Dogucan, et al.
Published: (2025)
Analyzing the Effect of Combined Degradations on Face Recognition
by: Sarıtaş, Erdi, et al.
Published: (2024)
by: Sarıtaş, Erdi, et al.
Published: (2024)
Analyzing the Feature Extractor Networks for Face Image Synthesis
by: Sarıtaş, Erdi, et al.
Published: (2024)
by: Sarıtaş, Erdi, et al.
Published: (2024)
Improved MambdaBDA Framework for Robust Building Damage Assessment Across Disaster Domains
by: Gençoğlu, Alp Eren, et al.
Published: (2026)
by: Gençoğlu, Alp Eren, et al.
Published: (2026)
Assessing the Use of Face Swapping Methods as Face Anonymizers in Videos
by: Muştu, Mustafa İzzet, et al.
Published: (2025)
by: Muştu, Mustafa İzzet, et al.
Published: (2025)
Impact of Surface Reflections in Maritime Obstacle Detection
by: Yalçın, Samed, et al.
Published: (2024)
by: Yalçın, Samed, et al.
Published: (2024)
Yolo-Key-6D: Single Stage Monocular 6D Pose Estimation with Keypoint Enhancements
by: Çetiner, Kemal Alperen, et al.
Published: (2026)
by: Çetiner, Kemal Alperen, et al.
Published: (2026)
Impact of Face Alignment on Face Image Quality
by: Onaran, Eren, et al.
Published: (2024)
by: Onaran, Eren, et al.
Published: (2024)
In-Bed Pose Estimation: A Review
by: Yazıcı, Ziya Ata, et al.
Published: (2024)
by: Yazıcı, Ziya Ata, et al.
Published: (2024)
Employing Vision-Language Models for Face Image Quality Assessment
by: Sarıtaş, Erdi, et al.
Published: (2026)
by: Sarıtaş, Erdi, et al.
Published: (2026)
On Applicability of Synthetic Datasets for Facial Expression Recognition
by: Azmoudeh, Ali, et al.
Published: (2026)
by: Azmoudeh, Ali, et al.
Published: (2026)
GLIMS: Attention-Guided Lightweight Multi-Scale Hybrid Network for Volumetric Semantic Segmentation
by: Yazıcı, Ziya Ata, et al.
Published: (2024)
by: Yazıcı, Ziya Ata, et al.
Published: (2024)
Bias-Aware Face Mask Detection Dataset
by: Kantarcı, Alperen, et al.
Published: (2022)
by: Kantarcı, Alperen, et al.
Published: (2022)
Facial Attribute Based Text Guided Face Anonymization
by: Muştu, Mustafa İzzet, et al.
Published: (2025)
by: Muştu, Mustafa İzzet, et al.
Published: (2025)
Attention-Enhanced Hybrid Feature Aggregation Network for 3D Brain Tumor Segmentation
by: Yazıcı, Ziya Ata, et al.
Published: (2024)
by: Yazıcı, Ziya Ata, et al.
Published: (2024)
A Survey on Class-Agnostic Counting: Advancements from Reference-Based to Open-World Text-Guided Approaches
by: Ciampi, Luca, et al.
Published: (2025)
by: Ciampi, Luca, et al.
Published: (2025)
Titanic Calling: Low Bandwidth Video Conference from the Titanic Wreck
by: Eyiokur, Fevziye Irem, et al.
Published: (2024)
by: Eyiokur, Fevziye Irem, et al.
Published: (2024)
CAPE: CAM as a Probabilistic Ensemble for Enhanced DNN Interpretation
by: Chowdhury, Townim Faisal, et al.
Published: (2024)
by: Chowdhury, Townim Faisal, et al.
Published: (2024)
AD-DINO: Attention-Dynamic DINO for Distance-Aware Embodied Reference Understanding
by: Guo, Hao, et al.
Published: (2024)
by: Guo, Hao, et al.
Published: (2024)
Hues and Cues: Human vs. CLIP
by: Alabau-Bosque, Nuria, et al.
Published: (2025)
by: Alabau-Bosque, Nuria, et al.
Published: (2025)
CAPE: Connectivity-Aware Path Enforcement Loss for Curvilinear Structure Delineation
by: Esmaeilzadeh, Elyar, et al.
Published: (2025)
by: Esmaeilzadeh, Elyar, et al.
Published: (2025)
Refining CNN-based Heatmap Regression with Gradient-based Corner Points for Electrode Localization
by: Wu, Lin
Published: (2024)
by: Wu, Lin
Published: (2024)
CLIP-BEVFormer: Enhancing Multi-View Image-Based BEV Detector with Ground Truth Flow
by: Pan, Chenbin, et al.
Published: (2024)
by: Pan, Chenbin, et al.
Published: (2024)
FA^{3}-CLIP: Frequency-Aware Cues Fusion and Attack-Agnostic Prompt Learning for Unified Face Attack Detection
by: Li, Yongze, et al.
Published: (2025)
by: Li, Yongze, et al.
Published: (2025)
Deformable-Heatmap-Segmentation for Automobile Visual Perception
by: Jin, Hongyu
Published: (2024)
by: Jin, Hongyu
Published: (2024)
Synthetic Image Detection with CLIP: Understanding and Assessing Predictive Cues
by: Willi, Marco, et al.
Published: (2026)
by: Willi, Marco, et al.
Published: (2026)
NeuroCLIP: Neuromorphic Data Understanding by CLIP and SNN
by: Guo, Yufei, et al.
Published: (2023)
by: Guo, Yufei, et al.
Published: (2023)
CueBench: Advancing Unified Understanding of Context-Aware Video Anomalies in Real-World
by: Yu, Yating, et al.
Published: (2025)
by: Yu, Yating, et al.
Published: (2025)
OCDet: Object Center Detection via Bounding Box-Aware Heatmap Prediction on Edge Devices with NPUs
by: Xin, Chen, et al.
Published: (2024)
by: Xin, Chen, et al.
Published: (2024)
Ultrasound-CLIP: Semantic-Aware Contrastive Pre-training for Ultrasound Image-Text Understanding
by: Jin, Jiayun, et al.
Published: (2026)
by: Jin, Jiayun, et al.
Published: (2026)
Salt & Pepper Heatmaps: Diffusion-informed Landmark Detection Strategy
by: Wyatt, Julian, et al.
Published: (2024)
by: Wyatt, Julian, et al.
Published: (2024)
Heatmap Pooling Network for Action Recognition from RGB Videos
by: Liu, Mengyuan, et al.
Published: (2025)
by: Liu, Mengyuan, et al.
Published: (2025)
Embodied Understanding of Driving Scenarios
by: Zhou, Yunsong, et al.
Published: (2024)
by: Zhou, Yunsong, et al.
Published: (2024)
MedP-CLIP: Medical CLIP with Region-Aware Prompt Integration
by: Peng, Jiahui, et al.
Published: (2026)
by: Peng, Jiahui, et al.
Published: (2026)
BadCLIP: Trigger-Aware Prompt Learning for Backdoor Attacks on CLIP
by: Bai, Jiawang, et al.
Published: (2023)
by: Bai, Jiawang, et al.
Published: (2023)
Similar Items
-
A Multimodal Depth-Aware Method For Embodied Reference Understanding
by: Eyiokur, Fevziye Irem, et al.
Published: (2025) -
Assessing Identity Leakage in Talking Face Generation: Metrics and Evaluation Framework
by: Yaman, Dogucan, et al.
Published: (2025) -
Audio-driven Talking Face Generation with Stabilized Synchronization Loss
by: Yaman, Dogucan, et al.
Published: (2023) -
Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation
by: Yaman, Dogucan, et al.
Published: (2025) -
Audio-Visual Speech Representation Expert for Enhanced Talking Face Video Generation and Evaluation
by: Yaman, Dogucan, et al.
Published: (2024)