Edge-Optimized Multimodal Learning for UAV Video Understanding via BLIP-2
Fuente:
arXiv
Saved in:
| Main Authors: | Feng, Yizhan, Snoussi, Hichem, Teng, Jing, Liu, Jian, Wang, Yuyang, Cherouat, Abel, Wang, Tian |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Large Language Models to Enhance Multi-task Drone Operations in Simulated Environments
by: Feng, Yizhan, et al.
Published: (2026)
by: Feng, Yizhan, et al.
Published: (2026)
Hybrid Distillation with CoT Guidance for Edge-Drone Control Code Generation
by: Feng, Yizhan, et al.
Published: (2026)
by: Feng, Yizhan, et al.
Published: (2026)
Distillation-based fabric anomaly detection
by: Thomine, Simon, et al.
Published: (2024)
by: Thomine, Simon, et al.
Published: (2024)
CSE: Surface Anomaly Detection with Contrastively Selected Embedding
by: Thomine, Simon, et al.
Published: (2024)
by: Thomine, Simon, et al.
Published: (2024)
Enhanced UAV Path Planning Using the Tangent Intersection Guidance (TIG) Algorithm
by: Cheriet, Hichem, et al.
Published: (2025)
by: Cheriet, Hichem, et al.
Published: (2025)
Multimodal Perception System for Real Open Environment
by: Sha, Yuyang
Published: (2024)
by: Sha, Yuyang
Published: (2024)
FineCog-Nav: Integrating Fine-grained Cognitive Modules for Zero-shot Multimodal UAV Navigation
by: Shao, Dian, et al.
Published: (2026)
by: Shao, Dian, et al.
Published: (2026)
DriveBLIP2: Attention-Guided Explanation Generation for Complex Driving Scenarios
by: Ling, Shihong, et al.
Published: (2025)
by: Ling, Shihong, et al.
Published: (2025)
Continuous Marine Tracking via Autonomous UAV Handoff
by: Kim, Heegyeong, et al.
Published: (2025)
by: Kim, Heegyeong, et al.
Published: (2025)
Simultaneous Tactile-Visual Perception for Learning Multimodal Robot Manipulation
by: Li, Yuyang, et al.
Published: (2025)
by: Li, Yuyang, et al.
Published: (2025)
UAV-Flow Colosseo: A Real-World Benchmark for Flying-on-a-Word UAV Imitation Learning
by: Wang, Xiangyu, et al.
Published: (2025)
by: Wang, Xiangyu, et al.
Published: (2025)
Exploring the best way for UAV visual localization under Low-altitude Multi-view Observation Condition: a Benchmark
by: Ye, Yibin, et al.
Published: (2025)
by: Ye, Yibin, et al.
Published: (2025)
Decoupling Ego-Motion from Target Dynamics via Dual-Interval Motion Cues for UAV Detection
by: Wang, Liuyang, et al.
Published: (2026)
by: Wang, Liuyang, et al.
Published: (2026)
Robo-MUTUAL: Robotic Multimodal Task Specification via Unimodal Learning
by: Li, Jianxiong, et al.
Published: (2024)
by: Li, Jianxiong, et al.
Published: (2024)
InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation
by: Yang, Shuai, et al.
Published: (2025)
by: Yang, Shuai, et al.
Published: (2025)
UAV-Track VLA: Embodied Aerial Tracking via Vision-Language-Action Models
by: Zhang, Qiyao, et al.
Published: (2026)
by: Zhang, Qiyao, et al.
Published: (2026)
ManipTrans: Efficient Dexterous Bimanual Manipulation Transfer via Residual Learning
by: Li, Kailin, et al.
Published: (2025)
by: Li, Kailin, et al.
Published: (2025)
EgoMimic: Scaling Imitation Learning via Egocentric Video
by: Kareer, Simar, et al.
Published: (2024)
by: Kareer, Simar, et al.
Published: (2024)
RoboMirror: Understand Before You Imitate for Video to Humanoid Locomotion
by: Li, Zhe, et al.
Published: (2025)
by: Li, Zhe, et al.
Published: (2025)
Acoustic Field Video for Multimodal Scene Understanding
by: Kim, Daehwa, et al.
Published: (2026)
by: Kim, Daehwa, et al.
Published: (2026)
Detector-Augmented SAMURAI for Long-Duration Drone Tracking
by: Lenhard, Tamara R., et al.
Published: (2026)
by: Lenhard, Tamara R., et al.
Published: (2026)
CityWalker: Learning Embodied Urban Navigation from Web-Scale Videos
by: Liu, Xinhao, et al.
Published: (2024)
by: Liu, Xinhao, et al.
Published: (2024)
Binding Touch to Everything: Learning Unified Multimodal Tactile Representations
by: Yang, Fengyu, et al.
Published: (2024)
by: Yang, Fengyu, et al.
Published: (2024)
Action Images: End-to-End Policy Learning via Multiview Video Generation
by: Zhen, Haoyu, et al.
Published: (2026)
by: Zhen, Haoyu, et al.
Published: (2026)
Active Human Pose Estimation via an Autonomous UAV Agent
by: Chen, Jingxi, et al.
Published: (2024)
by: Chen, Jingxi, et al.
Published: (2024)
Robot Learning from Human Videos: A Survey
by: Ma, Junyi, et al.
Published: (2026)
by: Ma, Junyi, et al.
Published: (2026)
BridgeV2W: Bridging Video Generation Models to Embodied World Models via Embodiment Masks
by: Chen, Yixiang, et al.
Published: (2026)
by: Chen, Yixiang, et al.
Published: (2026)
Towards Realistic UAV Vision-Language Navigation: Platform, Benchmark, and Methodology
by: Wang, Xiangyu, et al.
Published: (2024)
by: Wang, Xiangyu, et al.
Published: (2024)
ReliOcc: Towards Reliable Semantic Occupancy Prediction via Uncertainty Learning
by: Wang, Song, et al.
Published: (2024)
by: Wang, Song, et al.
Published: (2024)
LSGS-Loc: Towards Robust 3DGS-Based Visual Localization for Large-Scale UAV Scenarios
by: Zhang, Xiang, et al.
Published: (2026)
by: Zhang, Xiang, et al.
Published: (2026)
Tracking Meets Large Multimodal Models for Driving Scenario Understanding
by: Ishaq, Ayesha, et al.
Published: (2025)
by: Ishaq, Ayesha, et al.
Published: (2025)
Target-Oriented Object Grasping via Multimodal Human Guidance
by: Xie, Pengwei, et al.
Published: (2024)
by: Xie, Pengwei, et al.
Published: (2024)
Embodied Scene Understanding for Vision Language Models via MetaVQA
by: Wang, Weizhen, et al.
Published: (2025)
by: Wang, Weizhen, et al.
Published: (2025)
OpenGaussian: Towards Point-Level 3D Gaussian-based Open Vocabulary Understanding
by: Wu, Yanmin, et al.
Published: (2024)
by: Wu, Yanmin, et al.
Published: (2024)
CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning
by: Yang, Jiange, et al.
Published: (2025)
by: Yang, Jiange, et al.
Published: (2025)
Understanding Particles From Video: Property Estimation of Granular Materials via Visuo-Haptic Learning
by: Zhang, Zeqing, et al.
Published: (2024)
by: Zhang, Zeqing, et al.
Published: (2024)
OmniColor: A Global Camera Pose Optimization Approach of LiDAR-360Camera Fusion for Colorizing Point Clouds
by: Liu, Bonan, et al.
Published: (2024)
by: Liu, Bonan, et al.
Published: (2024)
Vision-Language Model for Accurate Crater Detection
by: Bauer, Patrick, et al.
Published: (2026)
by: Bauer, Patrick, et al.
Published: (2026)
Affordance-R1: Reinforcement Learning for Generalizable Affordance Reasoning in Multimodal Large Language Model
by: Wang, Hanqing, et al.
Published: (2025)
by: Wang, Hanqing, et al.
Published: (2025)
Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction
by: Fan, Chenyou, et al.
Published: (2025)
by: Fan, Chenyou, et al.
Published: (2025)
Similar Items
-
Large Language Models to Enhance Multi-task Drone Operations in Simulated Environments
by: Feng, Yizhan, et al.
Published: (2026) -
Hybrid Distillation with CoT Guidance for Edge-Drone Control Code Generation
by: Feng, Yizhan, et al.
Published: (2026) -
Distillation-based fabric anomaly detection
by: Thomine, Simon, et al.
Published: (2024) -
CSE: Surface Anomaly Detection with Contrastively Selected Embedding
by: Thomine, Simon, et al.
Published: (2024) -
Enhanced UAV Path Planning Using the Tangent Intersection Guidance (TIG) Algorithm
by: Cheriet, Hichem, et al.
Published: (2025)