V-IRL: Grounding Virtual Intelligence in Real Life
Fuente:
arXiv
Salvato in:
| Autori principali: | Yang, Jihan, Ding, Runyu, Brown, Ellis, Qi, Xiaojuan, Xie, Saining |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
RegionPLC: Regional Point-Language Contrastive Learning for Open-World 3D Scene Understanding
di: Yang, Jihan, et al.
Pubblicazione: (2023)
di: Yang, Jihan, et al.
Pubblicazione: (2023)
Benchmark Designers Should "Train on the Test Set" to Expose Exploitable Non-Visual Shortcuts
di: Brown, Ellis, et al.
Pubblicazione: (2025)
di: Brown, Ellis, et al.
Pubblicazione: (2025)
Traveling Across Languages: Benchmarking Cross-Lingual Consistency in Multimodal LLMs
di: Wang, Hao, et al.
Pubblicazione: (2025)
di: Wang, Hao, et al.
Pubblicazione: (2025)
UniTok: A Unified Tokenizer for Visual Generation and Understanding
di: Ma, Chuofan, et al.
Pubblicazione: (2025)
di: Ma, Chuofan, et al.
Pubblicazione: (2025)
Visual IRL for Human-Like Robotic Manipulation
di: Asali, Ehsan, et al.
Pubblicazione: (2024)
di: Asali, Ehsan, et al.
Pubblicazione: (2024)
Can 3D Vision-Language Models Truly Understand Natural Language?
di: Deng, Weipeng, et al.
Pubblicazione: (2024)
di: Deng, Weipeng, et al.
Pubblicazione: (2024)
SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
di: Chu, Tianzhe, et al.
Pubblicazione: (2025)
di: Chu, Tianzhe, et al.
Pubblicazione: (2025)
IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model
di: Jiang, Anqing, et al.
Pubblicazione: (2025)
di: Jiang, Anqing, et al.
Pubblicazione: (2025)
Cambrian-P: Pose-Grounded Video Understanding
di: Yang, Jihan, et al.
Pubblicazione: (2026)
di: Yang, Jihan, et al.
Pubblicazione: (2026)
OmniGround: A Comprehensive Spatio-Temporal Grounding Benchmark for Real-World Complex Scenarios
di: Gao, Hong, et al.
Pubblicazione: (2025)
di: Gao, Hong, et al.
Pubblicazione: (2025)
VULCAN: Tool-Augmented Multi Agents for Iterative 3D Object Arrangement
di: Kuang, Zhengfei, et al.
Pubblicazione: (2025)
di: Kuang, Zhengfei, et al.
Pubblicazione: (2025)
SIMS-V: Simulated Instruction-Tuning for Spatial Video Understanding
di: Brown, Ellis, et al.
Pubblicazione: (2025)
di: Brown, Ellis, et al.
Pubblicazione: (2025)
Groma: Localized Visual Tokenization for Grounding Multimodal Large Language Models
di: Ma, Chuofan, et al.
Pubblicazione: (2024)
di: Ma, Chuofan, et al.
Pubblicazione: (2024)
Science-T2I: Addressing Scientific Illusions in Image Synthesis
di: Li, Jialuo, et al.
Pubblicazione: (2025)
di: Li, Jialuo, et al.
Pubblicazione: (2025)
ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning
di: Xu, Ziqiang, et al.
Pubblicazione: (2025)
di: Xu, Ziqiang, et al.
Pubblicazione: (2025)
BridgeEQA: Virtual Embodied Agents for Real Bridge Inspections
di: Varghese, Subin, et al.
Pubblicazione: (2025)
di: Varghese, Subin, et al.
Pubblicazione: (2025)
Texture-Preserving Diffusion Models for High-Fidelity Virtual Try-On
di: Yang, Xu, et al.
Pubblicazione: (2024)
di: Yang, Xu, et al.
Pubblicazione: (2024)
Data Pruning by Information Maximization
di: Tan, Haoru, et al.
Pubblicazione: (2025)
di: Tan, Haoru, et al.
Pubblicazione: (2025)
Detection of On-Ground Chestnuts Using Artificial Intelligence Toward Automated Picking
di: Fang, Kaixuan, et al.
Pubblicazione: (2026)
di: Fang, Kaixuan, et al.
Pubblicazione: (2026)
Scaling Text-to-Image Diffusion Transformers with Representation Autoencoders
di: Tong, Shengbang, et al.
Pubblicazione: (2026)
di: Tong, Shengbang, et al.
Pubblicazione: (2026)
Augmented Commonsense Knowledge for Remote Object Grounding
di: Mohammadi, Bahram, et al.
Pubblicazione: (2024)
di: Mohammadi, Bahram, et al.
Pubblicazione: (2024)
PhyT2V: LLM-Guided Iterative Self-Refinement for Physics-Grounded Text-to-Video Generation
di: Xue, Qiyao, et al.
Pubblicazione: (2024)
di: Xue, Qiyao, et al.
Pubblicazione: (2024)
Lesion-Aware Generative Artificial Intelligence for Virtual Contrast-Enhanced Mammography in Breast Cancer
di: Rofena, Aurora, et al.
Pubblicazione: (2025)
di: Rofena, Aurora, et al.
Pubblicazione: (2025)
VideoMiner: Iteratively Grounding Key Frames of Hour-Long Videos via Tree-based Group Relative Policy Optimization
di: Cao, Xinye, et al.
Pubblicazione: (2025)
di: Cao, Xinye, et al.
Pubblicazione: (2025)
Transition Matching Distillation for Fast Video Generation
di: Nie, Weili, et al.
Pubblicazione: (2026)
di: Nie, Weili, et al.
Pubblicazione: (2026)
DOGR: Towards Versatile Visual Document Grounding and Referring
di: Zhou, Yinan, et al.
Pubblicazione: (2024)
di: Zhou, Yinan, et al.
Pubblicazione: (2024)
Temporally-Grounded Language Generation: A Benchmark for Real-Time Vision-Language Models
di: Yu, Keunwoo Peter, et al.
Pubblicazione: (2025)
di: Yu, Keunwoo Peter, et al.
Pubblicazione: (2025)
AirV2X: Unified Air-Ground Vehicle-to-Everything Collaboration
di: Gao, Xiangbo, et al.
Pubblicazione: (2025)
di: Gao, Xiangbo, et al.
Pubblicazione: (2025)
Seeing the Trees for the Forest: Rethinking Weakly-Supervised Medical Visual Grounding
di: Huy, Ta Duc, et al.
Pubblicazione: (2025)
di: Huy, Ta Duc, et al.
Pubblicazione: (2025)
Mastering Negation: Boosting Grounding Models via Grouped Opposition-Based Learning
di: Yang, Zesheng, et al.
Pubblicazione: (2026)
di: Yang, Zesheng, et al.
Pubblicazione: (2026)
Grounding Intelligence in Movement
di: Segado, Melanie, et al.
Pubblicazione: (2025)
di: Segado, Melanie, et al.
Pubblicazione: (2025)
Component-Aware Structure-Preserving Style Transfer for Satellite Visual Sim2Real Data Construction
di: Xie, Zongwu, et al.
Pubblicazione: (2026)
di: Xie, Zongwu, et al.
Pubblicazione: (2026)
SeqTex: Generate Mesh Textures in Video Sequence
di: Yuan, Ze, et al.
Pubblicazione: (2025)
di: Yuan, Ze, et al.
Pubblicazione: (2025)
Pareto-Enhanced Portrait Generation: Vision-Aligned Text Supervision for Alignment, Realism, and Aesthetics
di: Wang, Yunlong, et al.
Pubblicazione: (2026)
di: Wang, Yunlong, et al.
Pubblicazione: (2026)
Phi-Ground Tech Report: Advancing Perception in GUI Grounding
di: Zhang, Miaosen, et al.
Pubblicazione: (2025)
di: Zhang, Miaosen, et al.
Pubblicazione: (2025)
Towards Visual Discrimination and Reasoning of Real-World Physical Dynamics: Physics-Grounded Anomaly Detection
di: Li, Wenqiao, et al.
Pubblicazione: (2025)
di: Li, Wenqiao, et al.
Pubblicazione: (2025)
DiT-VTON: Diffusion Transformer Framework for Unified Multi-Category Virtual Try-On and Virtual Try-All with Integrated Image Editing
di: Li, Qi, et al.
Pubblicazione: (2025)
di: Li, Qi, et al.
Pubblicazione: (2025)
V-Zen: Efficient GUI Understanding and Precise Grounding With A Novel Multimodal LLM
di: Rahman, Abdur, et al.
Pubblicazione: (2024)
di: Rahman, Abdur, et al.
Pubblicazione: (2024)
ESCA: Enabling Seamless Codec Avatar Execution through Algorithm and Hardware Co-Optimization for Virtual Reality
di: Zhu, Mingzhi, et al.
Pubblicazione: (2025)
di: Zhu, Mingzhi, et al.
Pubblicazione: (2025)
ToG-Bench: Task-Oriented Spatio-Temporal Grounding in Egocentric Videos
di: Xu, Qi'ao, et al.
Pubblicazione: (2025)
di: Xu, Qi'ao, et al.
Pubblicazione: (2025)
Documenti analoghi
-
RegionPLC: Regional Point-Language Contrastive Learning for Open-World 3D Scene Understanding
di: Yang, Jihan, et al.
Pubblicazione: (2023) -
Benchmark Designers Should "Train on the Test Set" to Expose Exploitable Non-Visual Shortcuts
di: Brown, Ellis, et al.
Pubblicazione: (2025) -
Traveling Across Languages: Benchmarking Cross-Lingual Consistency in Multimodal LLMs
di: Wang, Hao, et al.
Pubblicazione: (2025) -
UniTok: A Unified Tokenizer for Visual Generation and Understanding
di: Ma, Chuofan, et al.
Pubblicazione: (2025) -
Visual IRL for Human-Like Robotic Manipulation
di: Asali, Ehsan, et al.
Pubblicazione: (2024)