See, Point, Fly: A Learning-Free VLM Framework for Universal Unmanned Aerial Navigation
Fuente:
arXiv
Saved in:
| Main Authors: | Hu, Chih Yao, Lin, Yang-Sen, Lee, Yuna, Su, Chih-Hai, Lee, Jie-Ying, Tsai, Shr-Ruei, Lin, Chin-Yang, Chen, Kuan-Wen, Ke, Tsung-Wei, Liu, Yu-Lun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BoostMVSNeRFs: Boosting MVS-based NeRFs to Generalizable View Synthesis in Large-scale Scenes
by: Su, Chih-Hai, et al.
Published: (2024)
by: Su, Chih-Hai, et al.
Published: (2024)
LightsOut: Diffusion-based Outpainting for Enhanced Lens Flare Removal
by: Tsai, Shr-Ruei, et al.
Published: (2025)
by: Tsai, Shr-Ruei, et al.
Published: (2025)
Skyfall-GS: Synthesizing Immersive 3D Urban Scenes from Satellite Imagery
by: Lee, Jie-Ying, et al.
Published: (2025)
by: Lee, Jie-Ying, et al.
Published: (2025)
PTT: Point-Trajectory Transformer for Efficient Temporal 3D Object Detection
by: Huang, Kuan-Chih, et al.
Published: (2023)
by: Huang, Kuan-Chih, et al.
Published: (2023)
Listen and Speak Fairly: A Study on Semantic Gender Bias in Speech Integrated Large Language Models
by: Lin, Yi-Cheng, et al.
Published: (2024)
by: Lin, Yi-Cheng, et al.
Published: (2024)
Weakly Supervised 3D Object Detection via Multi-Level Visual Guidance
by: Huang, Kuan-Chih, et al.
Published: (2023)
by: Huang, Kuan-Chih, et al.
Published: (2023)
Angle Robustness Unmanned Aerial Vehicle Navigation in GNSS-Denied Scenarios
by: Wang, Yuxin, et al.
Published: (2024)
by: Wang, Yuxin, et al.
Published: (2024)
Do Prompts Really Prompt? Exploring the Prompt Understanding Capability of Whisper
by: Yang, Chih-Kai, et al.
Published: (2024)
by: Yang, Chih-Kai, et al.
Published: (2024)
Digitisation of Impasto and Gloss in Oil Paintings via Spatially Varying Bidirectional Reflectance Distribution Function Acquisition
by: Chih Yang, et al.
Published: (2025)
by: Chih Yang, et al.
Published: (2025)
LTCXNet: Advancing Chest X-Ray Analysis with Solutions for Long-Tailed Multi-Label Classification and Fairness Challenges
by: Huang, Chin-Wei, et al.
Published: (2024)
by: Huang, Chin-Wei, et al.
Published: (2024)
Pairing Regularization for Mitigating Many-to-One Collapse in GANs
by: Lin, Kuan-Yu, et al.
Published: (2026)
by: Lin, Kuan-Yu, et al.
Published: (2026)
DeNVeR: Deformable Neural Vessel Representations for Unsupervised Video Vessel Segmentation
by: Wu, Chun-Hung, et al.
Published: (2024)
by: Wu, Chun-Hung, et al.
Published: (2024)
Investigating Zero-Shot Generalizability on Mandarin-English Code-Switched ASR and Speech-to-text Translation of Recent Foundation Models with Self-Supervision and Weak Supervision
by: Yang, Chih-Kai, et al.
Published: (2023)
by: Yang, Chih-Kai, et al.
Published: (2023)
Hide to See: Reasoning-prefix Masking for Visual-anchored Thinking in VLM Distillation
by: Yu, Seonghoon, et al.
Published: (2026)
by: Yu, Seonghoon, et al.
Published: (2026)
Zero Resource Code-switched Speech Benchmark Using Speech Utterance Pairs For Multiple Spoken Languages
by: Huang, Kuan-Po, et al.
Published: (2023)
by: Huang, Kuan-Po, et al.
Published: (2023)
Speech-Copilot: Leveraging Large Language Models for Speech Processing via Task Decomposition, Modularization, and Program Generation
by: Kuan, Chun-Yi, et al.
Published: (2024)
by: Kuan, Chun-Yi, et al.
Published: (2024)
A New Pipeline For Generating Instruction Dataset via RAG and Self Fine-Tuning
by: Song, Chih-Wei, et al.
Published: (2024)
by: Song, Chih-Wei, et al.
Published: (2024)
Progressive Alignment with VLM-LLM Feature to Augment Defect Classification for the ASE Dataset
by: Hsu, Chih-Chung, et al.
Published: (2024)
by: Hsu, Chih-Chung, et al.
Published: (2024)
Large Language Model as an Assignment Evaluator: Insights, Feedback, and Challenges in a 1000+ Student Course
by: Chiang, Cheng-Han, et al.
Published: (2024)
by: Chiang, Cheng-Han, et al.
Published: (2024)
Tree-of-Text: A Tree-based Prompting Framework for Table-to-Text Generation in the Sports Domain
by: Chiang, Shang-Hsuan, et al.
Published: (2026)
by: Chiang, Shang-Hsuan, et al.
Published: (2026)
FreeFly-Thinking : Aligning Chain-of-Thought Reasoning with Continuous UAV Navigation
by: Zhou, Jiaxu, et al.
Published: (2026)
by: Zhou, Jiaxu, et al.
Published: (2026)
Deep Transformer Network for Monocular Pose Estimation of Shipborne Unmanned Aerial Vehicle
by: Wickramasuriya, Maneesha, et al.
Published: (2024)
by: Wickramasuriya, Maneesha, et al.
Published: (2024)
UAV-MM3D: A Large-Scale Synthetic Benchmark for 3D Perception of Unmanned Aerial Vehicles with Multi-Modal Data
by: Zou, Longkun, et al.
Published: (2025)
by: Zou, Longkun, et al.
Published: (2025)
DenseSR: Image Shadow Removal as Dense Prediction
by: Lin, Yu-Fan, et al.
Published: (2025)
by: Lin, Yu-Fan, et al.
Published: (2025)
A Compendium of Autonomous Navigation using Object Detection and Tracking in Unmanned Aerial Vehicles
by: Arora, Mohit, et al.
Published: (2025)
by: Arora, Mohit, et al.
Published: (2025)
Scaling Agents for Computer Use
by: Gonzalez-Pumariega, Gonzalo, et al.
Published: (2025)
by: Gonzalez-Pumariega, Gonzalo, et al.
Published: (2025)
Team NYCU at Defactify4: Robust Detection and Source Identification of AI-Generated Images Using CNN and CLIP-Based Models
by: Yang, Tsan-Tsung, et al.
Published: (2025)
by: Yang, Tsan-Tsung, et al.
Published: (2025)
Retrieval-Augmented Language Model for Extreme Multi-Label Knowledge Graph Link Prediction
by: Lin, Yu-Hsiang, et al.
Published: (2024)
by: Lin, Yu-Hsiang, et al.
Published: (2024)
See Behind Walls in Real-time Using Aerial Drones and Augmented Reality
by: Yang, Sikai, et al.
Published: (2024)
by: Yang, Sikai, et al.
Published: (2024)
Virtual Guidance as a Mid-level Representation for Navigation with Augmented Reality
by: Yang, Hsuan-Kung, et al.
Published: (2023)
by: Yang, Hsuan-Kung, et al.
Published: (2023)
NeuCLIP: Efficient Large-Scale CLIP Training with Neural Normalizer Optimization
by: Wei, Xiyuan, et al.
Published: (2025)
by: Wei, Xiyuan, et al.
Published: (2025)
LookasideVLN: Direction-Aware Aerial Vision-and-Language Navigation
by: Ning, Yuwei, et al.
Published: (2026)
by: Ning, Yuwei, et al.
Published: (2026)
Spoken Stereoset: On Evaluating Social Bias Toward Speaker in Speech Large Language Models
by: Lin, Yi-Cheng, et al.
Published: (2024)
by: Lin, Yi-Cheng, et al.
Published: (2024)
Stream-DiffVSR: Low-Latency Streamable Video Super-Resolution via Auto-Regressive Diffusion
by: Shiu, Hau-Shiang, et al.
Published: (2025)
by: Shiu, Hau-Shiang, et al.
Published: (2025)
Hyacinth6B: A large language model for Traditional Chinese
by: Song, Chih-Wei, et al.
Published: (2024)
by: Song, Chih-Wei, et al.
Published: (2024)
An Efficient Additive Kolmogorov-Arnold Transformer for Point-Level Maize Localization in Unmanned Aerial Vehicle Imagery
by: Li, Fei, et al.
Published: (2026)
by: Li, Fei, et al.
Published: (2026)
Conditional Semi-Supervised Data Augmentation for Spam Message Detection with Low Resource Data
by: Nuha, Ulin, et al.
Published: (2024)
by: Nuha, Ulin, et al.
Published: (2024)
On-the-Fly Object-aware Representative Point Selection in Point Cloud
by: Zhang, Xiaoyu, et al.
Published: (2025)
by: Zhang, Xiaoyu, et al.
Published: (2025)
Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation
by: Wang, Zihan, et al.
Published: (2025)
by: Wang, Zihan, et al.
Published: (2025)
HIFuzz: Human Interaction Fuzzing for small Unmanned Aerial Vehicles
by: Chambers, Theodore, et al.
Published: (2023)
by: Chambers, Theodore, et al.
Published: (2023)
Similar Items
-
BoostMVSNeRFs: Boosting MVS-based NeRFs to Generalizable View Synthesis in Large-scale Scenes
by: Su, Chih-Hai, et al.
Published: (2024) -
LightsOut: Diffusion-based Outpainting for Enhanced Lens Flare Removal
by: Tsai, Shr-Ruei, et al.
Published: (2025) -
Skyfall-GS: Synthesizing Immersive 3D Urban Scenes from Satellite Imagery
by: Lee, Jie-Ying, et al.
Published: (2025) -
PTT: Point-Trajectory Transformer for Efficient Temporal 3D Object Detection
by: Huang, Kuan-Chih, et al.
Published: (2023) -
Listen and Speak Fairly: A Study on Semantic Gender Bias in Speech Integrated Large Language Models
by: Lin, Yi-Cheng, et al.
Published: (2024)