CLIPSwarm: Generating Drone Shows from Text Prompts with Vision-Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pueyo, Pablo, Montijano, Eduardo, Murillo, Ana C., Schwager, Mac |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Gen-Swarms: Adapting Deep Generative Models to Swarms of Drones
von: Plou, Carlos, et al.
Veröffentlicht: (2024)
von: Plou, Carlos, et al.
Veröffentlicht: (2024)
SketchPlan: Diffusion Based Drone Planning From Human Sketches
von: Norelius, Sixten, et al.
Veröffentlicht: (2025)
von: Norelius, Sixten, et al.
Veröffentlicht: (2025)
CineMPC: A Fully Autonomous Drone Cinematography System Incorporating Zoom, Focus, Pose, and Scene Composition
von: Pueyo, Pablo, et al.
Veröffentlicht: (2024)
von: Pueyo, Pablo, et al.
Veröffentlicht: (2024)
SpectralWaste Dataset: Multimodal Data for Waste Sorting Automation
von: Casao, Sara, et al.
Veröffentlicht: (2024)
von: Casao, Sara, et al.
Veröffentlicht: (2024)
VLN-Pilot: Large Vision-Language Model as an Autonomous Indoor Drone Operator
von: Dominguez-Dager, Bessie, et al.
Veröffentlicht: (2026)
von: Dominguez-Dager, Bessie, et al.
Veröffentlicht: (2026)
SIREN: Semantic, Initialization-Free Registration of Multi-Robot Gaussian Splatting Maps
von: Shorinwa, Ola, et al.
Veröffentlicht: (2025)
von: Shorinwa, Ola, et al.
Veröffentlicht: (2025)
SOUS VIDE: Cooking Visual Drone Navigation Policies in a Gaussian Splatting Vacuum
von: Low, JunEn, et al.
Veröffentlicht: (2024)
von: Low, JunEn, et al.
Veröffentlicht: (2024)
Coverage Optimization for Camera View Selection
von: Chen, Timothy, et al.
Veröffentlicht: (2026)
von: Chen, Timothy, et al.
Veröffentlicht: (2026)
Aion: Towards Hierarchical 4D Scene Graphs with Temporal Flow Dynamics
von: Catalano, Iacopo, et al.
Veröffentlicht: (2025)
von: Catalano, Iacopo, et al.
Veröffentlicht: (2025)
MAP-VLA: Memory-Augmented Prompting for Vision-Language-Action Model in Robotic Manipulation
von: Li, Runhao, et al.
Veröffentlicht: (2025)
von: Li, Runhao, et al.
Veröffentlicht: (2025)
Where to Perch in a Tree: Vision-Guidance for Tree-Grasping Drones
von: Dunnett, Alex, et al.
Veröffentlicht: (2026)
von: Dunnett, Alex, et al.
Veröffentlicht: (2026)
WeatherPrompt: Multi-modality Representation Learning for All-Weather Drone Visual Geo-Localization
von: Wen, Jiahao, et al.
Veröffentlicht: (2025)
von: Wen, Jiahao, et al.
Veröffentlicht: (2025)
OTTER: A Vision-Language-Action Model with Text-Aware Visual Feature Extraction
von: Huang, Huang, et al.
Veröffentlicht: (2025)
von: Huang, Huang, et al.
Veröffentlicht: (2025)
Touch-GS: Visual-Tactile Supervised 3D Gaussian Splatting
von: Swann, Aiden, et al.
Veröffentlicht: (2024)
von: Swann, Aiden, et al.
Veröffentlicht: (2024)
SynDroneVision: A Synthetic Dataset for Image-Based Drone Detection
von: Lenhard, Tamara R., et al.
Veröffentlicht: (2024)
von: Lenhard, Tamara R., et al.
Veröffentlicht: (2024)
Generalized Robot 3D Vision-Language Model with Fast Rendering and Pre-Training Vision-Language Alignment
von: Liu, Kangcheng, et al.
Veröffentlicht: (2023)
von: Liu, Kangcheng, et al.
Veröffentlicht: (2023)
Breaking Lock-In: Preserving Steerability under Low-Data VLA Post-Training
von: Huang, Suning, et al.
Veröffentlicht: (2026)
von: Huang, Suning, et al.
Veröffentlicht: (2026)
Seeing to Act, Prompting to Specify: A Bayesian Factorization of Vision Language Action Policy
von: Xu, Kechun, et al.
Veröffentlicht: (2025)
von: Xu, Kechun, et al.
Veröffentlicht: (2025)
Splat-MOVER: Multi-Stage, Open-Vocabulary Robotic Manipulation via Editable Gaussian Splatting
von: Shorinwa, Ola, et al.
Veröffentlicht: (2024)
von: Shorinwa, Ola, et al.
Veröffentlicht: (2024)
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models
von: Fang, Irving, et al.
Veröffentlicht: (2025)
von: Fang, Irving, et al.
Veröffentlicht: (2025)
PanoGen++: Domain-Adapted Text-Guided Panoramic Environment Generation for Vision-and-Language Navigation
von: Wang, Sen, et al.
Veröffentlicht: (2025)
von: Wang, Sen, et al.
Veröffentlicht: (2025)
Semantic-Aware Guided Drone Exploration for Language-Conditioned 3D Indoor Mapping
von: Vegesna, Nitin, et al.
Veröffentlicht: (2026)
von: Vegesna, Nitin, et al.
Veröffentlicht: (2026)
Self-Improving Vision-Language-Action Models with Data Generation via Residual RL
von: Xiao, Wenli, et al.
Veröffentlicht: (2025)
von: Xiao, Wenli, et al.
Veröffentlicht: (2025)
DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control
von: Wen, Junjie, et al.
Veröffentlicht: (2025)
von: Wen, Junjie, et al.
Veröffentlicht: (2025)
F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions
von: Lv, Qi, et al.
Veröffentlicht: (2025)
von: Lv, Qi, et al.
Veröffentlicht: (2025)
GeneralVLA: Generalizable Vision-Language-Action Models with Knowledge-Guided Trajectory Planning
von: Ma, Guoqing, et al.
Veröffentlicht: (2026)
von: Ma, Guoqing, et al.
Veröffentlicht: (2026)
Unified Vision-Language-Action Model
von: Wang, Yuqi, et al.
Veröffentlicht: (2025)
von: Wang, Yuqi, et al.
Veröffentlicht: (2025)
Bridging Text and Vision: A Multi-View Text-Vision Registration Approach for Cross-Modal Place Recognition
von: Shang, Tianyi, et al.
Veröffentlicht: (2025)
von: Shang, Tianyi, et al.
Veröffentlicht: (2025)
Video Individual Counting for Moving Drones
von: Fan, Yaowu, et al.
Veröffentlicht: (2025)
von: Fan, Yaowu, et al.
Veröffentlicht: (2025)
Multimodal Fusion and Vision-Language Models: A Survey for Robot Vision
von: Han, Xiaofeng, et al.
Veröffentlicht: (2025)
von: Han, Xiaofeng, et al.
Veröffentlicht: (2025)
AffordanceLLM: Grounding Affordance from Vision Language Models
von: Qian, Shengyi, et al.
Veröffentlicht: (2024)
von: Qian, Shengyi, et al.
Veröffentlicht: (2024)
Reasoning-VLA: A Fast and General Vision-Language-Action Reasoning Model for Autonomous Driving
von: Zhang, Dapeng, et al.
Veröffentlicht: (2025)
von: Zhang, Dapeng, et al.
Veröffentlicht: (2025)
ActDistill: General Action-Guided Self-Derived Distillation for Efficient Vision-Language-Action Models
von: Ye, Wencheng, et al.
Veröffentlicht: (2025)
von: Ye, Wencheng, et al.
Veröffentlicht: (2025)
DroneVis: Versatile Computer Vision Library for Drones
von: Heakl, Ahmed, et al.
Veröffentlicht: (2024)
von: Heakl, Ahmed, et al.
Veröffentlicht: (2024)
Decentralized Vehicle Coordination: The Berkeley DeepDrive Drone Dataset and Consensus-Based Models
von: Wu, Fangyu, et al.
Veröffentlicht: (2022)
von: Wu, Fangyu, et al.
Veröffentlicht: (2022)
Learning Camera Movement Control from Real-World Drone Videos
von: Hou, Yunzhong, et al.
Veröffentlicht: (2024)
von: Hou, Yunzhong, et al.
Veröffentlicht: (2024)
Exploring the Limits of Vision-Language-Action Manipulations in Cross-task Generalization
von: Zhou, Jiaming, et al.
Veröffentlicht: (2025)
von: Zhou, Jiaming, et al.
Veröffentlicht: (2025)
ScenarioControl: Vision-Language Controllable Vectorized Latent Scenario Generation
von: Gao, Lili, et al.
Veröffentlicht: (2026)
von: Gao, Lili, et al.
Veröffentlicht: (2026)
Zero-Shot 3D Visual Grounding from Vision-Language Models
von: Li, Rong, et al.
Veröffentlicht: (2025)
von: Li, Rong, et al.
Veröffentlicht: (2025)
Understanding the Impact of Geometric Foundation Models on Vision-Language-Action Models
von: Yang, Yurou, et al.
Veröffentlicht: (2026)
von: Yang, Yurou, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Gen-Swarms: Adapting Deep Generative Models to Swarms of Drones
von: Plou, Carlos, et al.
Veröffentlicht: (2024) -
SketchPlan: Diffusion Based Drone Planning From Human Sketches
von: Norelius, Sixten, et al.
Veröffentlicht: (2025) -
CineMPC: A Fully Autonomous Drone Cinematography System Incorporating Zoom, Focus, Pose, and Scene Composition
von: Pueyo, Pablo, et al.
Veröffentlicht: (2024) -
SpectralWaste Dataset: Multimodal Data for Waste Sorting Automation
von: Casao, Sara, et al.
Veröffentlicht: (2024) -
VLN-Pilot: Large Vision-Language Model as an Autonomous Indoor Drone Operator
von: Dominguez-Dager, Bessie, et al.
Veröffentlicht: (2026)