FALCON: Future-Aware Learning with Contextual Object-Centric Pretraining for UAV Action Recognition
Fuente:
arXiv
Saved in:
| Main Authors: | Xian, Ruiqi, Wu, Xiyang, Guan, Tianrui, Wang, Xijun, Gong, Boqing, Manocha, Dinesh |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SCP: Soft Conditional Prompt Learning for Aerial Video Action Recognition
by: Wang, Xijun, et al.
Published: (2023)
by: Wang, Xijun, et al.
Published: (2023)
AGL-NET: Aerial-Ground Cross-Modal Global Localization with Varying Scales
by: Guan, Tianrui, et al.
Published: (2024)
by: Guan, Tianrui, et al.
Published: (2024)
LOC-ZSON: Language-driven Object-Centric Zero-Shot Object Retrieval and Navigation
by: Guan, Tianrui, et al.
Published: (2024)
by: Guan, Tianrui, et al.
Published: (2024)
Paired-CSLiDAR: Height-Stratified Registration for Cross-Source Aerial-Ground LiDAR Pose Refinement
by: Hoover, Montana, et al.
Published: (2026)
by: Hoover, Montana, et al.
Published: (2026)
DAVE: Diverse Atomic Visual Elements Dataset with High Representation of Vulnerable Road Users in Complex and Unpredictable Environments
by: Wang, Xijun, et al.
Published: (2024)
by: Wang, Xijun, et al.
Published: (2024)
Recasting Generic Pretrained Vision Transformers As Object-Centric Scene Encoders For Manipulation Policies
by: Qian, Jianing, et al.
Published: (2024)
by: Qian, Jianing, et al.
Published: (2024)
Personalized Embodied Navigation for Portable Object Finding
by: Dorbala, Vishnu Sashank, et al.
Published: (2024)
by: Dorbala, Vishnu Sashank, et al.
Published: (2024)
Improving Zero-Shot ObjectNav with Generative Communication
by: Dorbala, Vishnu Sashank, et al.
Published: (2024)
by: Dorbala, Vishnu Sashank, et al.
Published: (2024)
HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models
by: Guan, Tianrui, et al.
Published: (2023)
by: Guan, Tianrui, et al.
Published: (2023)
SLAT-Phys: Fast Material Property Field Prediction from Structured 3D Latents
by: Das, Rocktim Jyoti, et al.
Published: (2026)
by: Das, Rocktim Jyoti, et al.
Published: (2026)
Object-Centric Action-Enhanced Representations for Robot Visuo-Motor Policy Learning
by: Giannakakis, Nikos, et al.
Published: (2025)
by: Giannakakis, Nikos, et al.
Published: (2025)
Bi-VLM: Pushing Ultra-Low Precision Post-Training Quantization Boundaries in Vision-Language Models
by: Wang, Xijun, et al.
Published: (2025)
by: Wang, Xijun, et al.
Published: (2025)
MotionHint: Self-Supervised Monocular Visual Odometry with Motion Constraints
by: Wang, Cong, et al.
Published: (2021)
by: Wang, Cong, et al.
Published: (2021)
TK-Planes: Tiered K-Planes with High Dimensional Feature Vectors for Dynamic UAV-based Scenes
by: Maxey, Christopher, et al.
Published: (2024)
by: Maxey, Christopher, et al.
Published: (2024)
Learning Action-Conditional and Object-Centric Gaussian Splatting World Models for Rigid Objects
by: Kreber, Jens U., et al.
Published: (2026)
by: Kreber, Jens U., et al.
Published: (2026)
OCRA: Object-Centric Learning with 3D and Tactile Priors for Human-to-Robot Action Transfer
by: Wang, Kuanning, et al.
Published: (2026)
by: Wang, Kuanning, et al.
Published: (2026)
AutoHallusion: Automatic Generation of Hallucination Benchmarks for Vision-Language Models
by: Wu, Xiyang, et al.
Published: (2024)
by: Wu, Xiyang, et al.
Published: (2024)
PhysGS: Bayesian-Inferred Gaussian Splatting for Physical Property Estimation
by: Chopra, Samarth, et al.
Published: (2025)
by: Chopra, Samarth, et al.
Published: (2025)
Splatblox: Traversability-Aware Gaussian Splatting for Outdoor Robot Navigation
by: Chopra, Samarth, et al.
Published: (2025)
by: Chopra, Samarth, et al.
Published: (2025)
CSCPR: Cross-Source-Context Indoor RGB-D Place Recognition
by: Liang, Jing, et al.
Published: (2024)
by: Liang, Jing, et al.
Published: (2024)
On the Vulnerability of LLM/VLM-Controlled Robotics
by: Wu, Xiyang, et al.
Published: (2024)
by: Wu, Xiyang, et al.
Published: (2024)
UAV-ON: A Benchmark for Open-World Object Goal Navigation with Aerial Agents
by: Xiao, Jianqiang, et al.
Published: (2025)
by: Xiao, Jianqiang, et al.
Published: (2025)
TAG: Target-Agnostic Guidance for Stable Object-Centric Inference in Vision-Language-Action Models
by: Zhou, Jiaying, et al.
Published: (2026)
by: Zhou, Jiaying, et al.
Published: (2026)
Differentiable Frequency-based Disentanglement for Aerial Video Action Recognition
by: Kothandaraman, Divya, et al.
Published: (2022)
by: Kothandaraman, Divya, et al.
Published: (2022)
CLAP: Contrastive Latent Action Pretraining for Learning Vision-Language-Action Models from Human Videos
by: Zhang, Chubin, et al.
Published: (2026)
by: Zhang, Chubin, et al.
Published: (2026)
AerialMind: Towards Referring Multi-Object Tracking in UAV Scenarios
by: Chen, Chenglizhao, et al.
Published: (2025)
by: Chen, Chenglizhao, et al.
Published: (2025)
Disentangled Object-Centric Image Representation for Robotic Manipulation
by: Emukpere, David, et al.
Published: (2025)
by: Emukpere, David, et al.
Published: (2025)
Latent Action Pretraining Through World Modeling
by: Tharwat, Bahey, et al.
Published: (2025)
by: Tharwat, Bahey, et al.
Published: (2025)
FuncGrasp: Learning Object-Centric Neural Grasp Functions from Single Annotated Example Object
by: Chen, Hanzhi, et al.
Published: (2024)
by: Chen, Hanzhi, et al.
Published: (2024)
STEP: Spatial Temporal Graph Convolutional Networks for Emotion Perception from Gaits
by: Bhattacharya, Uttaran, et al.
Published: (2019)
by: Bhattacharya, Uttaran, et al.
Published: (2019)
Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy
by: Zhang, Tianyi, et al.
Published: (2025)
by: Zhang, Tianyi, et al.
Published: (2025)
Object-Centric Instruction Augmentation for Robotic Manipulation
by: Wen, Junjie, et al.
Published: (2024)
by: Wen, Junjie, et al.
Published: (2024)
Research on Robot Path Planning Based on Reinforcement Learning
by: Ruiqi, Wang
Published: (2024)
by: Ruiqi, Wang
Published: (2024)
Composing Pre-Trained Object-Centric Representations for Robotics From "What" and "Where" Foundation Models
by: Shi, Junyao, et al.
Published: (2024)
by: Shi, Junyao, et al.
Published: (2024)
4D-ROLLS: 4D Radar Occupancy Learning via LiDAR Supervision
by: Liu, Ruihan, et al.
Published: (2025)
by: Liu, Ruihan, et al.
Published: (2025)
3D Feature Distillation with Object-Centric Priors
by: Tziafas, Georgios, et al.
Published: (2024)
by: Tziafas, Georgios, et al.
Published: (2024)
Can Transformers Capture Spatial Relations between Objects?
by: Wen, Chuan, et al.
Published: (2024)
by: Wen, Chuan, et al.
Published: (2024)
Hand-Object Interaction Pretraining from Videos
by: Singh, Himanshu Gaurav, et al.
Published: (2024)
by: Singh, Himanshu Gaurav, et al.
Published: (2024)
Visual Place Recognition for Large-Scale UAV Applications
by: Papapetros, Ioannis Tsampikos, et al.
Published: (2025)
by: Papapetros, Ioannis Tsampikos, et al.
Published: (2025)
UAV4D: Dynamic Neural Rendering of Human-Centric UAV Imagery using Gaussian Splatting
by: Choi, Jaehoon, et al.
Published: (2025)
by: Choi, Jaehoon, et al.
Published: (2025)
Similar Items
-
SCP: Soft Conditional Prompt Learning for Aerial Video Action Recognition
by: Wang, Xijun, et al.
Published: (2023) -
AGL-NET: Aerial-Ground Cross-Modal Global Localization with Varying Scales
by: Guan, Tianrui, et al.
Published: (2024) -
LOC-ZSON: Language-driven Object-Centric Zero-Shot Object Retrieval and Navigation
by: Guan, Tianrui, et al.
Published: (2024) -
Paired-CSLiDAR: Height-Stratified Registration for Cross-Source Aerial-Ground LiDAR Pose Refinement
by: Hoover, Montana, et al.
Published: (2026) -
DAVE: Diverse Atomic Visual Elements Dataset with High Representation of Vulnerable Road Users in Complex and Unpredictable Environments
by: Wang, Xijun, et al.
Published: (2024)