Guardado en:
| Autores principales: | Wulff, Theodor, Abawi, Fares, Allgeuer, Philipp, Wermter, Stefan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2504.05913 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Unified Dynamic Scanpath Predictors Outperform Individually Trained Neural Models
por: Abawi, Fares, et al.
Publicado: (2024)
por: Abawi, Fares, et al.
Publicado: (2024)
Unconstrained Open Vocabulary Image Classification: Zero-Shot Transfer from Text to Image via CLIP Inversion
por: Allgeuer, Philipp, et al.
Publicado: (2024)
por: Allgeuer, Philipp, et al.
Publicado: (2024)
Human Impression of Humanoid Robots Mirroring Social Cues
por: Fu, Di, et al.
Publicado: (2024)
por: Fu, Di, et al.
Publicado: (2024)
Wrapyfi: A Python Wrapper for Integrating Robots, Sensors, and Applications across Multiple Middleware
por: Abawi, Fares, et al.
Publicado: (2023)
por: Abawi, Fares, et al.
Publicado: (2023)
Pointing-Guided Target Estimation via Transformer-Based Attention
por: Müller, Luca, et al.
Publicado: (2025)
por: Müller, Luca, et al.
Publicado: (2025)
SalFormer360: a transformer-based saliency estimation model for 360-degree videos
por: Wahba, Mahmoud Z. A., et al.
Publicado: (2026)
por: Wahba, Mahmoud Z. A., et al.
Publicado: (2026)
Segmenting the motion components of a video: A long-term unsupervised model
por: Meunier, Etienne, et al.
Publicado: (2023)
por: Meunier, Etienne, et al.
Publicado: (2023)
Noise-Free Explanation for Driving Action Prediction
por: Zhu, Hongbo, et al.
Publicado: (2024)
por: Zhu, Hongbo, et al.
Publicado: (2024)
Prompt-to-Gesture: Measuring the Capabilities of Image-to-Video Deictic Gesture Generation
por: Ali, Hassan, et al.
Publicado: (2026)
por: Ali, Hassan, et al.
Publicado: (2026)
A method for estimating roadway billboard salience
por: Haladova, Zuzana Berger, et al.
Publicado: (2025)
por: Haladova, Zuzana Berger, et al.
Publicado: (2025)
Opti-CAM: Optimizing saliency maps for interpretability
por: Zhang, Hanwei, et al.
Publicado: (2023)
por: Zhang, Hanwei, et al.
Publicado: (2023)
Improving saliency models' predictions of the next fixation with humans' intrinsic cost of gaze shifts
por: Kadner, Florian, et al.
Publicado: (2022)
por: Kadner, Florian, et al.
Publicado: (2022)
Dream-to-Recon: Monocular 3D Reconstruction with Diffusion-Depth Distillation from Single Images
por: Wulff, Philipp, et al.
Publicado: (2025)
por: Wulff, Philipp, et al.
Publicado: (2025)
Concept-Based Explanations in Computer Vision: Where Are We and Where Could We Go?
por: Lee, Jae Hee, et al.
Publicado: (2024)
por: Lee, Jae Hee, et al.
Publicado: (2024)
Bias Leaves a Gradient Trail: Label-Free Bias Identification via Gradient Probes on Concept Decompositions
por: Vitry, Thomas, et al.
Publicado: (2026)
por: Vitry, Thomas, et al.
Publicado: (2026)
Bridging visual saliency and large language models for explainable deep learning in medical imaging
por: Nguezet, Paul Valery, et al.
Publicado: (2026)
por: Nguezet, Paul Valery, et al.
Publicado: (2026)
Gradient-based multi-focus image fusion with focus-aware saliency enhancement
por: Li, Haoyu, et al.
Publicado: (2025)
por: Li, Haoyu, et al.
Publicado: (2025)
StateVLM: A State-Aware Vision-Language Model for Robotic Affordance Reasoning
por: Sun, Xiaowen, et al.
Publicado: (2026)
por: Sun, Xiaowen, et al.
Publicado: (2026)
Online pre-training with long-form videos
por: Kato, Itsuki, et al.
Publicado: (2024)
por: Kato, Itsuki, et al.
Publicado: (2024)
Towards Learning a Generalizable 3D Scene Representation from 2D Observations
por: Gromniak, Martin, et al.
Publicado: (2026)
por: Gromniak, Martin, et al.
Publicado: (2026)
Read Between the Layers: Leveraging Multi-Layer Representations for Rehearsal-Free Continual Learning with Pre-Trained Models
por: Ahrens, Kyra, et al.
Publicado: (2023)
por: Ahrens, Kyra, et al.
Publicado: (2023)
AMEGO: Active Memory from long EGOcentric videos
por: Goletto, Gabriele, et al.
Publicado: (2024)
por: Goletto, Gabriele, et al.
Publicado: (2024)
Koala: Key frame-conditioned long video-LLM
por: Tan, Reuben, et al.
Publicado: (2024)
por: Tan, Reuben, et al.
Publicado: (2024)
Comparing Apples to Oranges: LLM-powered Multimodal Intention Prediction in an Object Categorization Task
por: Ali, Hassan, et al.
Publicado: (2024)
por: Ali, Hassan, et al.
Publicado: (2024)
From Neural Activations to Concepts: A Survey on Explaining Concepts in Neural Networks
por: Lee, Jae Hee, et al.
Publicado: (2023)
por: Lee, Jae Hee, et al.
Publicado: (2023)
FabuLight-ASD: Unveiling Speech Activity via Body Language
por: Carneiro, Hugo, et al.
Publicado: (2024)
por: Carneiro, Hugo, et al.
Publicado: (2024)
Is ImageNet worth 1 video? Learning strong image encoders from 1 long unlabelled video
por: Venkataramanan, Shashanka, et al.
Publicado: (2023)
por: Venkataramanan, Shashanka, et al.
Publicado: (2023)
Recent Advances in Medical Imaging Segmentation: A Survey
por: Bougourzi, Fares, et al.
Publicado: (2025)
por: Bougourzi, Fares, et al.
Publicado: (2025)
MDS-ViTNet: Improving saliency prediction for Eye-Tracking with Vision Transformer
por: Ignat, Polezhaev, et al.
Publicado: (2024)
por: Ignat, Polezhaev, et al.
Publicado: (2024)
Evaluating saliency scores in point clouds of natural environments by learning surface anomalies
por: Arav, Reuma, et al.
Publicado: (2024)
por: Arav, Reuma, et al.
Publicado: (2024)
MIXER: Mixed Hyperspherical Random Embedding Neural Network for Texture Recognition
por: Fares, Ricardo T., et al.
Publicado: (2025)
por: Fares, Ricardo T., et al.
Publicado: (2025)
An HMM-based framework for identity-aware long-term multi-object tracking from sparse and uncertain identification: use case on long-term tracking in livestock
por: Bibinbe, Anne Marthe Sophie Ngo, et al.
Publicado: (2025)
por: Bibinbe, Anne Marthe Sophie Ngo, et al.
Publicado: (2025)
Taming generative video models for zero-shot optical flow extraction
por: Kim, Seungwoo, et al.
Publicado: (2025)
por: Kim, Seungwoo, et al.
Publicado: (2025)
Decoding Matters: Efficient Mamba-Based Decoder with Distribution-Aware Deep Supervision for Medical Image Segmentation
por: Bougourzi, Fares, et al.
Publicado: (2026)
por: Bougourzi, Fares, et al.
Publicado: (2026)
Extremely Fine-Grained Visual Classification over Resembling Glyphs in the Wild
por: Bougourzi, Fares, et al.
Publicado: (2024)
por: Bougourzi, Fares, et al.
Publicado: (2024)
VALD: Multi-Stage Vision Attack Detection for Efficient LVLM Defense
por: Kadvil, Nadav, et al.
Publicado: (2026)
por: Kadvil, Nadav, et al.
Publicado: (2026)
Instance-level quantitative saliency in multiple sclerosis lesion segmentation
por: Spagnolo, Federico, et al.
Publicado: (2024)
por: Spagnolo, Federico, et al.
Publicado: (2024)
FALCONEye: Finding Answers and Localizing Content in ONE-hour-long videos with multi-modal LLMs
por: Plou, Carlos, et al.
Publicado: (2025)
por: Plou, Carlos, et al.
Publicado: (2025)
Toward Unified Practices in Trajectory Prediction Research on Bird's-Eye-View Datasets
por: Westny, Theodor, et al.
Publicado: (2024)
por: Westny, Theodor, et al.
Publicado: (2024)
Learning semantical dynamics and spatiotemporal collaboration for human pose estimation in video
por: Feng, Runyang, et al.
Publicado: (2025)
por: Feng, Runyang, et al.
Publicado: (2025)
Ejemplares similares
-
Unified Dynamic Scanpath Predictors Outperform Individually Trained Neural Models
por: Abawi, Fares, et al.
Publicado: (2024) -
Unconstrained Open Vocabulary Image Classification: Zero-Shot Transfer from Text to Image via CLIP Inversion
por: Allgeuer, Philipp, et al.
Publicado: (2024) -
Human Impression of Humanoid Robots Mirroring Social Cues
por: Fu, Di, et al.
Publicado: (2024) -
Wrapyfi: A Python Wrapper for Integrating Robots, Sensors, and Applications across Multiple Middleware
por: Abawi, Fares, et al.
Publicado: (2023) -
Pointing-Guided Target Estimation via Transformer-Based Attention
por: Müller, Luca, et al.
Publicado: (2025)