Arcadia: Toward a Full-Lifecycle Framework for Embodied Lifelong Learning
Fuente:
arXiv
Guardado en:
| Autores principales: | Gao, Minghe, Li, Juncheng, Lin, Yuze, Liu, Xuqi, Ji, Jiaming, Pan, Xiaoran, Xu, Zihan, Li, Xian, Li, Mingjie, Ji, Wei, Wei, Rong, Tang, Rui, Wang, Qizhou, Shen, Kai, Xiao, Jun, Wu, Qi, Tang, Siliang, Zhuang, Yueting |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Benchmarking Multimodal CoT Reward Model Stepwise by Visual Program
por: Gao, Minghe, et al.
Publicado: (2025)
por: Gao, Minghe, et al.
Publicado: (2025)
De-fine: Decomposing and Refining Visual Programs with Auto-Feedback
por: Gao, Minghe, et al.
Publicado: (2023)
por: Gao, Minghe, et al.
Publicado: (2023)
Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions
por: Li, Juncheng, et al.
Publicado: (2023)
por: Li, Juncheng, et al.
Publicado: (2023)
InstructVid2Vid: Controllable Video Editing with Natural Language Instructions
por: Qin, Bosheng, et al.
Publicado: (2023)
por: Qin, Bosheng, et al.
Publicado: (2023)
SpatialFusion: Endowing Unified Image Generation with Intrinsic 3D Geometric Awareness
por: Qiu, Haiyi, et al.
Publicado: (2026)
por: Qiu, Haiyi, et al.
Publicado: (2026)
What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities
por: Bu, Wendong, et al.
Publicado: (2025)
por: Bu, Wendong, et al.
Publicado: (2025)
KCM: KAN-Based Collaboration Models Enhance Pretrained Large Models
por: Dai, Guangyu, et al.
Publicado: (2025)
por: Dai, Guangyu, et al.
Publicado: (2025)
Generalist Virtual Agents: A Survey on Autonomous Agents Across Digital Platforms
por: Gao, Minghe, et al.
Publicado: (2024)
por: Gao, Minghe, et al.
Publicado: (2024)
T2S-GPT: Dynamic Vector Quantization for Autoregressive Sign Language Production from Text
por: Yin, Aoxiong, et al.
Publicado: (2024)
por: Yin, Aoxiong, et al.
Publicado: (2024)
Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining
por: Ge, Zhiqi, et al.
Publicado: (2024)
por: Ge, Zhiqi, et al.
Publicado: (2024)
Towards Physically Executable 3D Gaussian for Embodied Navigation
por: Miao, Bingchen, et al.
Publicado: (2025)
por: Miao, Bingchen, et al.
Publicado: (2025)
WorldGPT: Empowering LLM as Multimodal World Model
por: Ge, Zhiqi, et al.
Publicado: (2024)
por: Ge, Zhiqi, et al.
Publicado: (2024)
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training
por: Qiu, Haiyi, et al.
Publicado: (2024)
por: Qiu, Haiyi, et al.
Publicado: (2024)
Fact :Teaching MLLMs with Faithful, Concise and Transferable Rationales
por: Gao, Minghe, et al.
Publicado: (2024)
por: Gao, Minghe, et al.
Publicado: (2024)
GMFVAD: Using Grained Multi-modal Feature to Improve Video Anomaly Detection
por: Dai, Guangyu, et al.
Publicado: (2025)
por: Dai, Guangyu, et al.
Publicado: (2025)
HalluciDoctor: Mitigating Hallucinatory Toxicity in Visual Instruction Data
por: Yu, Qifan, et al.
Publicado: (2023)
por: Yu, Qifan, et al.
Publicado: (2023)
InstructSAM: Segment Any Instance with Any Instructions
por: Yuan, Yuqian, et al.
Publicado: (2026)
por: Yuan, Yuqian, et al.
Publicado: (2026)
Chart-HQA: A Benchmark for Hypothetical Question Answering in Charts
por: Chen, Xiangnan, et al.
Publicado: (2025)
por: Chen, Xiangnan, et al.
Publicado: (2025)
SOYO: A Tuning-Free Approach for Video Style Morphing via Style-Adaptive Interpolation in Diffusion Models
por: Zheng, Haoyu, et al.
Publicado: (2025)
por: Zheng, Haoyu, et al.
Publicado: (2025)
Auto-Encoding Morph-Tokens for Multimodal LLM
por: Pan, Kaihang, et al.
Publicado: (2024)
por: Pan, Kaihang, et al.
Publicado: (2024)
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL
por: Pan, Kaihang, et al.
Publicado: (2025)
por: Pan, Kaihang, et al.
Publicado: (2025)
Robust Modality-incomplete Anomaly Detection: A Modality-instructive Framework with Benchmark
por: Miao, Bingchen, et al.
Publicado: (2024)
por: Miao, Bingchen, et al.
Publicado: (2024)
Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning
por: Qian, Long, et al.
Publicado: (2024)
por: Qian, Long, et al.
Publicado: (2024)
AnyEdit: Mastering Unified High-Quality Image Editing for Any Idea
por: Yu, Qifan, et al.
Publicado: (2024)
por: Yu, Qifan, et al.
Publicado: (2024)
Bridging Local Details and Global Context in Text-Attributed Graphs
por: Wang, Yaoke, et al.
Publicado: (2024)
por: Wang, Yaoke, et al.
Publicado: (2024)
DuetRAG: Collaborative Retrieval-Augmented Generation
por: Jiao, Dian, et al.
Publicado: (2024)
por: Jiao, Dian, et al.
Publicado: (2024)
OmniMoGen: Unifying Human Motion Generation via Learning from Interleaved Text-Motion Instructions
por: Bu, Wendong, et al.
Publicado: (2025)
por: Bu, Wendong, et al.
Publicado: (2025)
Mastering Collaborative Multi-modal Data Selection: A Focus on Informativeness, Uniqueness, and Representativeness
por: Yu, Qifan, et al.
Publicado: (2024)
por: Yu, Qifan, et al.
Publicado: (2024)
LASER: Tuning-Free LLM-Driven Attention Control for Efficient Text-conditioned Image-to-Animation
por: Zheng, Haoyu, et al.
Publicado: (2024)
por: Zheng, Haoyu, et al.
Publicado: (2024)
Align$^2$LLaVA: Cascaded Human and Large Language Model Preference Alignment for Multi-modal Instruction Curation
por: Huang, Hongzhe, et al.
Publicado: (2024)
por: Huang, Hongzhe, et al.
Publicado: (2024)
CrossView Suite: Harnessing Cross-view Spatial Intelligence of MLLMs with Dataset, Model and Benchmark
por: Wang, Wei, et al.
Publicado: (2026)
por: Wang, Wei, et al.
Publicado: (2026)
Janus-Pro-R1: Advancing Collaborative Visual Comprehension and Generation via Reinforcement Learning
por: Pan, Kaihang, et al.
Publicado: (2025)
por: Pan, Kaihang, et al.
Publicado: (2025)
Fast Thinking for Large Language Models
por: Zheng, Haoyu, et al.
Publicado: (2025)
por: Zheng, Haoyu, et al.
Publicado: (2025)
MoA: Heterogeneous Mixture of Adapters for Parameter-Efficient Fine-Tuning of Large Language Models
por: Cao, Jie, et al.
Publicado: (2025)
por: Cao, Jie, et al.
Publicado: (2025)
Towards Meta-Cognitive Knowledge Editing for Multimodal LLMs
por: Fan, Zhaoyu, et al.
Publicado: (2025)
por: Fan, Zhaoyu, et al.
Publicado: (2025)
CORE: Code-based Inverse Self-Training Framework with Graph Expansion for Virtual Agents
por: Wang, Keyu, et al.
Publicado: (2026)
por: Wang, Keyu, et al.
Publicado: (2026)
Ask Questions with Double Hints: Visual Question Generation with Answer-awareness and Region-reference
por: Shen, Kai, et al.
Publicado: (2024)
por: Shen, Kai, et al.
Publicado: (2024)
IDEAL: Leveraging Infinite and Dynamic Characterizations of Large Language Models for Query-focused Summarization
por: Cao, Jie, et al.
Publicado: (2024)
por: Cao, Jie, et al.
Publicado: (2024)
MAKIMA: Tuning-free Multi-Attribute Open-domain Video Editing via Mask-Guided Attention Modulation
por: Zheng, Haoyu, et al.
Publicado: (2024)
por: Zheng, Haoyu, et al.
Publicado: (2024)
TeamLoRA: Boosting Low-Rank Adaptation with Expert Collaboration and Competition
por: Lin, Tianwei, et al.
Publicado: (2024)
por: Lin, Tianwei, et al.
Publicado: (2024)
Ejemplares similares
-
Benchmarking Multimodal CoT Reward Model Stepwise by Visual Program
por: Gao, Minghe, et al.
Publicado: (2025) -
De-fine: Decomposing and Refining Visual Programs with Auto-Feedback
por: Gao, Minghe, et al.
Publicado: (2023) -
Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions
por: Li, Juncheng, et al.
Publicado: (2023) -
InstructVid2Vid: Controllable Video Editing with Natural Language Instructions
por: Qin, Bosheng, et al.
Publicado: (2023) -
SpatialFusion: Endowing Unified Image Generation with Intrinsic 3D Geometric Awareness
por: Qiu, Haiyi, et al.
Publicado: (2026)