LLaRA: Supercharging Robot Learning Data for Vision-Language Policy
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Xiang, Mata, Cristina, Park, Jongwoo, Kahatapitiya, Kumara, Jang, Yoo Sung, Shang, Jinghuan, Ranasinghe, Kanchana, Burgert, Ryan, Cai, Mu, Lee, Yong Jae, Ryoo, Michael S. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Language Repository for Long Video Understanding
von: Kahatapitiya, Kumara, et al.
Veröffentlicht: (2024)
von: Kahatapitiya, Kumara, et al.
Veröffentlicht: (2024)
Too Many Frames, Not All Useful: Efficient Strategies for Long-Form Video QA
von: Park, Jongwoo, et al.
Veröffentlicht: (2024)
von: Park, Jongwoo, et al.
Veröffentlicht: (2024)
Understanding Long Videos with Multimodal Language Models
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2024)
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2024)
IVRA: Improving Visual-Token Relations for Robot Action Policy with Training-Free Hint-Based Guidance
von: Park, Jongwoo, et al.
Veröffentlicht: (2026)
von: Park, Jongwoo, et al.
Veröffentlicht: (2026)
Pixel Motion as Universal Representation for Robot Control
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2025)
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2025)
CoPT: Unsupervised Domain Adaptive Segmentation using Domain-Agnostic Text Embeddings
von: Mata, Cristina, et al.
Veröffentlicht: (2025)
von: Mata, Cristina, et al.
Veröffentlicht: (2025)
VicTR: Video-conditioned Text Representations for Activity Recognition
von: Kahatapitiya, Kumara, et al.
Veröffentlicht: (2023)
von: Kahatapitiya, Kumara, et al.
Veröffentlicht: (2023)
LLaRA: Large Language-Recommendation Assistant
von: Liao, Jiayi, et al.
Veröffentlicht: (2023)
von: Liao, Jiayi, et al.
Veröffentlicht: (2023)
Pixel Motion Diffusion is What We Need for Robot Control
von: Nguyen, E-Ro, et al.
Veröffentlicht: (2025)
von: Nguyen, E-Ro, et al.
Veröffentlicht: (2025)
Crossway Diffusion: Improving Diffusion-based Visuomotor Policy via Self-supervised Learning
von: Li, Xiang, et al.
Veröffentlicht: (2023)
von: Li, Xiang, et al.
Veröffentlicht: (2023)
LACE: Latent Visual Representation for Cross-Embodiment Learning
von: Jang, Yoo Sung, et al.
Veröffentlicht: (2026)
von: Jang, Yoo Sung, et al.
Veröffentlicht: (2026)
Learning to Localize Objects Improves Spatial Reasoning in Visual-LLMs
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2024)
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2024)
Adaptive Caching for Faster Video Generation with Diffusion Transformers
von: Kahatapitiya, Kumara, et al.
Veröffentlicht: (2024)
von: Kahatapitiya, Kumara, et al.
Veröffentlicht: (2024)
xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs
von: Ryoo, Michael S., et al.
Veröffentlicht: (2024)
von: Ryoo, Michael S., et al.
Veröffentlicht: (2024)
LLaVA-PruMerge: Adaptive Token Reduction for Efficient Large Multimodal Models
von: Shang, Yuzhang, et al.
Veröffentlicht: (2024)
von: Shang, Yuzhang, et al.
Veröffentlicht: (2024)
Future Optical Flow Prediction Improves Robot Control & Video Generation
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2026)
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2026)
LatentCRF: Continuous CRF for Efficient Latent Diffusion
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2024)
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2024)
Bayesian Principles Improve Prompt Learning In Vision-Language Models
von: Kim, Mingyu, et al.
Veröffentlicht: (2025)
von: Kim, Mingyu, et al.
Veröffentlicht: (2025)
Yo'LLaVA: Your Personalized Language and Vision Assistant
von: Nguyen, Thao, et al.
Veröffentlicht: (2024)
von: Nguyen, Thao, et al.
Veröffentlicht: (2024)
Hierarchical Text-to-Vision Self Supervised Alignment for Improved Histopathology Representation Learning
von: Watawana, Hasindri, et al.
Veröffentlicht: (2024)
von: Watawana, Hasindri, et al.
Veröffentlicht: (2024)
MotionV2V: Editing Motion in a Video
von: Burgert, Ryan, et al.
Veröffentlicht: (2025)
von: Burgert, Ryan, et al.
Veröffentlicht: (2025)
Robotic VLA Benefits from Joint Learning with Motion Image Diffusion
von: Fang, Yu, et al.
Veröffentlicht: (2025)
von: Fang, Yu, et al.
Veröffentlicht: (2025)
Object-Centric Diffusion for Efficient Video Editing
von: Kahatapitiya, Kumara, et al.
Veröffentlicht: (2024)
von: Kahatapitiya, Kumara, et al.
Veröffentlicht: (2024)
Image Translation with Kernel Prediction Networks for Semantic Segmentation
von: Mata, Cristina, et al.
Veröffentlicht: (2025)
von: Mata, Cristina, et al.
Veröffentlicht: (2025)
Theia: Distilling Diverse Vision Foundation Models for Robot Learning
von: Shang, Jinghuan, et al.
Veröffentlicht: (2024)
von: Shang, Jinghuan, et al.
Veröffentlicht: (2024)
ViP-LLaVA: Making Large Multimodal Models Understand Arbitrary Visual Prompts
von: Cai, Mu, et al.
Veröffentlicht: (2023)
von: Cai, Mu, et al.
Veröffentlicht: (2023)
X-LLaVA: Optimizing Bilingual Large Vision-Language Alignment
von: Shin, Dongjae, et al.
Veröffentlicht: (2024)
von: Shin, Dongjae, et al.
Veröffentlicht: (2024)
NAVIS: Concurrent Search and Update with Low Position-Seeking Overhead in On-SSD Graph-Based Vector Search
von: Song, Jaeyong, et al.
Veröffentlicht: (2026)
von: Song, Jaeyong, et al.
Veröffentlicht: (2026)
VisionLLaMA: A Unified LLaMA Backbone for Vision Tasks
von: Chu, Xiangxiang, et al.
Veröffentlicht: (2024)
von: Chu, Xiangxiang, et al.
Veröffentlicht: (2024)
Supercharging Bayesian Inference with Reliable AI-Informed Priors
von: Choi, Jongwoo, et al.
Veröffentlicht: (2026)
von: Choi, Jongwoo, et al.
Veröffentlicht: (2026)
First Logit Boosting: Visual Grounding Method to Mitigate Object Hallucination in Large Vision-Language Models
von: Ha, Jiwoo, et al.
Veröffentlicht: (2026)
von: Ha, Jiwoo, et al.
Veröffentlicht: (2026)
LoraMap: Harnessing the Power of LoRA Connections
von: Park, Hyeryun, et al.
Veröffentlicht: (2024)
von: Park, Hyeryun, et al.
Veröffentlicht: (2024)
CoLLaVO: Crayon Large Language and Vision mOdel
von: Lee, Byung-Kwan, et al.
Veröffentlicht: (2024)
von: Lee, Byung-Kwan, et al.
Veröffentlicht: (2024)
Vision as LoRA
von: Wang, Han, et al.
Veröffentlicht: (2025)
von: Wang, Han, et al.
Veröffentlicht: (2025)
Estimating Physical Information Consistency of Channel Data Augmentation for Remote Sensing Images
von: Burgert, Tom, et al.
Veröffentlicht: (2024)
von: Burgert, Tom, et al.
Veröffentlicht: (2024)
Sceniris: A Fast Procedural Scene Generation Framework
von: Shang, Jinghuan, et al.
Veröffentlicht: (2025)
von: Shang, Jinghuan, et al.
Veröffentlicht: (2025)
Continual Learning with Global Alignment
von: Bai, Xueying, et al.
Veröffentlicht: (2022)
von: Bai, Xueying, et al.
Veröffentlicht: (2022)
LC-Flow: Learning Local Continuous Optical Flow and Confidence from events
von: Jeon, Gunwoo, et al.
Veröffentlicht: (2026)
von: Jeon, Gunwoo, et al.
Veröffentlicht: (2026)
HeightLane: BEV Heightmap guided 3D Lane Detection
von: Park, Chaesong, et al.
Veröffentlicht: (2024)
von: Park, Chaesong, et al.
Veröffentlicht: (2024)
Supercharging Floorplan Localization with Semantic Rays
von: Grader, Yuval, et al.
Veröffentlicht: (2025)
von: Grader, Yuval, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Language Repository for Long Video Understanding
von: Kahatapitiya, Kumara, et al.
Veröffentlicht: (2024) -
Too Many Frames, Not All Useful: Efficient Strategies for Long-Form Video QA
von: Park, Jongwoo, et al.
Veröffentlicht: (2024) -
Understanding Long Videos with Multimodal Language Models
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2024) -
IVRA: Improving Visual-Token Relations for Robot Action Policy with Training-Free Hint-Based Guidance
von: Park, Jongwoo, et al.
Veröffentlicht: (2026) -
Pixel Motion as Universal Representation for Robot Control
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2025)