Attentive Feature Aggregation or: How Policies Learn to Stop Worrying about Robustness and Attend to Task-Relevant Visual Cues
Fuente:
arXiv
Saved in:
| Main Authors: | Tsagkas, Nikolaos, Sochopoulos, Andreas, Danier, Duolikun, Vijayakumar, Sethu, Kouris, Alexandros, Mac Aodha, Oisin, Lu, Chris Xiaoxuan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Temporal Trap: Entanglement in Pre-Trained Visual Representations for Visuomotor Policy Learning
by: Tsagkas, Nikolaos, et al.
Published: (2025)
by: Tsagkas, Nikolaos, et al.
Published: (2025)
DepthCues: Evaluating Monocular Depth Perception in Large Vision Models
by: Danier, Duolikun, et al.
Published: (2024)
by: Danier, Duolikun, et al.
Published: (2024)
Fast Flow-based Visuomotor Policies via Conditional Optimal Transport Couplings
by: Sochopoulos, Andreas, et al.
Published: (2025)
by: Sochopoulos, Andreas, et al.
Published: (2025)
Click to Grasp: Zero-Shot Precise Manipulation via Visual Diffusion Descriptors
by: Tsagkas, Nikolaos, et al.
Published: (2024)
by: Tsagkas, Nikolaos, et al.
Published: (2024)
Learning Deep Dynamical Systems using Stable Neural ODEs
by: Sochopoulos, Andreas, et al.
Published: (2024)
by: Sochopoulos, Andreas, et al.
Published: (2024)
View-Consistent Diffusion Representations for 3D-Consistent Video Generation
by: Danier, Duolikun, et al.
Published: (2025)
by: Danier, Duolikun, et al.
Published: (2025)
Enhancing Tactile-based Reinforcement Learning for Robotic Control
by: Miller, Elle, et al.
Published: (2025)
by: Miller, Elle, et al.
Published: (2025)
Human-in-the-loop Optimisation in Robot-assisted Gait Training
by: Christou, Andreas, et al.
Published: (2025)
by: Christou, Andreas, et al.
Published: (2025)
roto 2.0: The Robot Tactile Olympiad
by: Miller, Elle, et al.
Published: (2026)
by: Miller, Elle, et al.
Published: (2026)
SAOR: Single-View Articulated Object Reconstruction
by: Aygün, Mehmet, et al.
Published: (2023)
by: Aygün, Mehmet, et al.
Published: (2023)
Interpretable Text-Guided Image Clustering via Iterative Search
by: Zhao, Bingchen, et al.
Published: (2025)
by: Zhao, Bingchen, et al.
Published: (2025)
BVI-VFI: A Video Quality Database for Video Frame Interpolation
by: Danier, Duolikun, et al.
Published: (2022)
by: Danier, Duolikun, et al.
Published: (2022)
Enhancing Deformable Convolution based Video Frame Interpolation with Coarse-to-fine 3D CNN
by: Danier, Duolikun, et al.
Published: (2022)
by: Danier, Duolikun, et al.
Published: (2022)
ST-MFNet: A Spatio-Temporal Multi-Flow Network for Frame Interpolation
by: Danier, Duolikun, et al.
Published: (2021)
by: Danier, Duolikun, et al.
Published: (2021)
LDMVFI: Video Frame Interpolation with Latent Diffusion Models
by: Danier, Duolikun, et al.
Published: (2023)
by: Danier, Duolikun, et al.
Published: (2023)
A Subjective Quality Study for Video Frame Interpolation
by: Danier, Duolikun, et al.
Published: (2022)
by: Danier, Duolikun, et al.
Published: (2022)
Learning Precise Affordances from Egocentric Videos for Robotic Manipulation
by: Li, Gen, et al.
Published: (2024)
by: Li, Gen, et al.
Published: (2024)
Self-Supervised Multimodal Learning: A Survey
by: Zong, Yongshuo, et al.
Published: (2023)
by: Zong, Yongshuo, et al.
Published: (2023)
Improving Semantic Correspondence with Viewpoint-Guided Spherical Maps
by: Mariotti, Octave, et al.
Published: (2023)
by: Mariotti, Octave, et al.
Published: (2023)
Representational Similarity via Interpretable Visual Concepts
by: Kondapaneni, Neehar, et al.
Published: (2025)
by: Kondapaneni, Neehar, et al.
Published: (2025)
Representational Difference Explanations
by: Kondapaneni, Neehar, et al.
Published: (2025)
by: Kondapaneni, Neehar, et al.
Published: (2025)
RankDVQA: Deep VQA based on Ranking-inspired Hybrid Training
by: Feng, Chen, et al.
Published: (2022)
by: Feng, Chen, et al.
Published: (2022)
Poke and Strike: Learning Task-Informed Exploration Policies
by: Aoyama, Marina Y., et al.
Published: (2025)
by: Aoyama, Marina Y., et al.
Published: (2025)
Pseudo-Equilibria, or: How to Stop Worrying About Crypto and Just Analyze the Game
by: Psomas, Alexandros, et al.
Published: (2025)
by: Psomas, Alexandros, et al.
Published: (2025)
MoonSeg3R: Monocular Online Zero-Shot Segment Anything in 3D with Reconstructive Foundation Priors
by: Du, Zhipeng, et al.
Published: (2025)
by: Du, Zhipeng, et al.
Published: (2025)
Less is More: Discovering Concise Network Explanations
by: Kondapaneni, Neehar, et al.
Published: (2024)
by: Kondapaneni, Neehar, et al.
Published: (2024)
Generating Binary Species Range Maps
by: Dorm, Filip, et al.
Published: (2024)
by: Dorm, Filip, et al.
Published: (2024)
Labeled Data Selection for Category Discovery
by: Zhao, Bingchen, et al.
Published: (2024)
by: Zhao, Bingchen, et al.
Published: (2024)
MotionPhysics: Learnable Motion Distillation for Text-Guided Simulation
by: Wang, Miaowei, et al.
Published: (2026)
by: Wang, Miaowei, et al.
Published: (2026)
CleverBirds: A Multiple-Choice Benchmark for Fine-grained Human Knowledge Tracing
by: Bossemeyer, Leonie, et al.
Published: (2025)
by: Bossemeyer, Leonie, et al.
Published: (2025)
Enhancing 2D Representation Learning with a 3D Prior
by: Aygün, Mehmet, et al.
Published: (2024)
by: Aygün, Mehmet, et al.
Published: (2024)
Jamais Vu: Exposing the Generalization Gap in Supervised Semantic Correspondence
by: Mariotti, Octave, et al.
Published: (2025)
by: Mariotti, Octave, et al.
Published: (2025)
VesselSDF: Distance Field Priors for Vascular Network Reconstruction
by: Esposito, Salvatore, et al.
Published: (2025)
by: Esposito, Salvatore, et al.
Published: (2025)
MVAD: A Multiple Visual Artifact Detector for Video Streaming
by: Feng, Chen, et al.
Published: (2024)
by: Feng, Chen, et al.
Published: (2024)
BVI-Artefact: An Artefact Detection Benchmark Dataset for Streamed Videos
by: Feng, Chen, et al.
Published: (2023)
by: Feng, Chen, et al.
Published: (2023)
Sample-efficient Integration of New Modalities into Large Language Models
by: İnce, Osman Batur, et al.
Published: (2025)
by: İnce, Osman Batur, et al.
Published: (2025)
WildSAT: Learning Satellite Image Representations from Wildlife Observations
by: Daroya, Rangel, et al.
Published: (2024)
by: Daroya, Rangel, et al.
Published: (2024)
How I Learned to Stop Worrying and Love Manga!
by: Samantha Archibald Mora, et al.
Published: (2023)
by: Samantha Archibald Mora, et al.
Published: (2023)
Few-shot transfer of tool-use skills using human demonstrations with proximity and tactile sensing
by: Aoyama, Marina Y., et al.
Published: (2025)
by: Aoyama, Marina Y., et al.
Published: (2025)
acoupi: An Open-Source Python Framework for Deploying Bioacoustic AI Models on Edge Devices
by: Vuilliomenet, Aude, et al.
Published: (2025)
by: Vuilliomenet, Aude, et al.
Published: (2025)
Similar Items
-
The Temporal Trap: Entanglement in Pre-Trained Visual Representations for Visuomotor Policy Learning
by: Tsagkas, Nikolaos, et al.
Published: (2025) -
DepthCues: Evaluating Monocular Depth Perception in Large Vision Models
by: Danier, Duolikun, et al.
Published: (2024) -
Fast Flow-based Visuomotor Policies via Conditional Optimal Transport Couplings
by: Sochopoulos, Andreas, et al.
Published: (2025) -
Click to Grasp: Zero-Shot Precise Manipulation via Visual Diffusion Descriptors
by: Tsagkas, Nikolaos, et al.
Published: (2024) -
Learning Deep Dynamical Systems using Stable Neural ODEs
by: Sochopoulos, Andreas, et al.
Published: (2024)