Guardado en:
| Autores principales: | Wang, Renhao, Geng, Haoran, Li, Tingle, Wang, Feishi, Anumanchipalli, Gopala, Darrell, Trevor, Li, Boyi, Abbeel, Pieter, Malik, Jitendra, Efros, Alexei A. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2507.02864 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Self-Supervised Audio-Visual Soundscape Stylization
por: Li, Tingle, et al.
Publicado: (2024)
por: Li, Tingle, et al.
Publicado: (2024)
Audio Texture Manipulation by Exemplar-Based Analogy
por: Cheng, Kan Jen, et al.
Publicado: (2025)
por: Cheng, Kan Jen, et al.
Publicado: (2025)
Prioritized Generative Replay
por: Wang, Renhao, et al.
Publicado: (2024)
por: Wang, Renhao, et al.
Publicado: (2024)
Interactive Task Planning with Language Models
por: Li, Boyi, et al.
Publicado: (2023)
por: Li, Boyi, et al.
Publicado: (2023)
Sounding that Object: Interactive Object-Aware Image to Audio Generation
por: Li, Tingle, et al.
Publicado: (2025)
por: Li, Tingle, et al.
Publicado: (2025)
ViTacFormer: Learning Cross-Modal Representation for Visuo-Tactile Dexterous Manipulation
por: Heng, Liang, et al.
Publicado: (2025)
por: Heng, Liang, et al.
Publicado: (2025)
Rodrigues Network for Learning Robot Actions
por: Zhang, Jialiang, et al.
Publicado: (2025)
por: Zhang, Jialiang, et al.
Publicado: (2025)
Synthesizing Moving People with 3D Control
por: Li, Boyi, et al.
Publicado: (2024)
por: Li, Boyi, et al.
Publicado: (2024)
D-REX: Differentiable Real-to-Sim-to-Real Engine for Learning Dexterous Grasping
por: Lou, Haozhe, et al.
Publicado: (2026)
por: Lou, Haozhe, et al.
Publicado: (2026)
DexGarmentLab: Dexterous Garment Manipulation Environment with Generalizable Policy
por: Wang, Yuran, et al.
Publicado: (2025)
por: Wang, Yuran, et al.
Publicado: (2025)
Rethinking Patch Dependence for Masked Autoencoders
por: Fu, Letian, et al.
Publicado: (2024)
por: Fu, Letian, et al.
Publicado: (2024)
SkillBlender: Towards Versatile Humanoid Whole-Body Loco-Manipulation via Skill Blending
por: Kuang, Yuxuan, et al.
Publicado: (2025)
por: Kuang, Yuxuan, et al.
Publicado: (2025)
Deep Sensorimotor Control by Imitating Predictive Models of Human Motion
por: Singh, Himanshu Gaurav, et al.
Publicado: (2025)
por: Singh, Himanshu Gaurav, et al.
Publicado: (2025)
DIPOLE: Fusing Vision and Geometry for Robust Visuomotor Generalization
por: Tang, Yikai, et al.
Publicado: (2025)
por: Tang, Yikai, et al.
Publicado: (2025)
Multi-Objective Learning for Diffusion Models: A Statistical Theory under Semi-Supervised Learning
por: Cheng, Ziheng, et al.
Publicado: (2026)
por: Cheng, Ziheng, et al.
Publicado: (2026)
StyleStream: Real-Time Zero-Shot Voice Style Conversion
por: Liu, Yisi, et al.
Publicado: (2026)
por: Liu, Yisi, et al.
Publicado: (2026)
End-to-end RL Improves Dexterous Grasping Policies
por: Singh, Ritvik, et al.
Publicado: (2025)
por: Singh, Ritvik, et al.
Publicado: (2025)
Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities
por: Lin, Guan-Ting, et al.
Publicado: (2025)
por: Lin, Guan-Ting, et al.
Publicado: (2025)
Learning Humanoid Locomotion over Challenging Terrain
por: Radosavovic, Ilija, et al.
Publicado: (2024)
por: Radosavovic, Ilija, et al.
Publicado: (2024)
Re-evaluating the Need for Multimodal Signals in Unsupervised Grammar Induction
por: Li, Boyi, et al.
Publicado: (2022)
por: Li, Boyi, et al.
Publicado: (2022)
Towards Hierarchical Spoken Language Dysfluency Modeling
por: Lian, Jiachen, et al.
Publicado: (2024)
por: Lian, Jiachen, et al.
Publicado: (2024)
Large Video Planner Enables Generalizable Robot Control
por: Chen, Boyuan, et al.
Publicado: (2025)
por: Chen, Boyuan, et al.
Publicado: (2025)
It's Never Too Late: Noise Optimization for Collapse Recovery in Trained Diffusion Models
por: Harrington, Anne, et al.
Publicado: (2025)
por: Harrington, Anne, et al.
Publicado: (2025)
Closing the Visual Sim-to-Real Gap with Object-Composable NeRFs
por: Mishra, Nikhil, et al.
Publicado: (2024)
por: Mishra, Nikhil, et al.
Publicado: (2024)
How to Peel with a Knife: Aligning Fine-Grained Manipulation with Human Preference
por: Lin, Toru, et al.
Publicado: (2026)
por: Lin, Toru, et al.
Publicado: (2026)
Twisting Lids Off with Two Hands
por: Lin, Toru, et al.
Publicado: (2024)
por: Lin, Toru, et al.
Publicado: (2024)
Visual Imitation Enables Contextual Humanoid Control
por: Allshire, Arthur, et al.
Publicado: (2025)
por: Allshire, Arthur, et al.
Publicado: (2025)
RT-VC: Real-Time Zero-Shot Voice Conversion with Speech Articulatory Coding
por: Liu, Yisi, et al.
Publicado: (2025)
por: Liu, Yisi, et al.
Publicado: (2025)
xT: Nested Tokenization for Larger Context in Large Images
por: Gupta, Ritwik, et al.
Publicado: (2024)
por: Gupta, Ritwik, et al.
Publicado: (2024)
LLM-grounded Diffusion: Enhancing Prompt Understanding of Text-to-Image Diffusion Models with Large Language Models
por: Lian, Long, et al.
Publicado: (2023)
por: Lian, Long, et al.
Publicado: (2023)
Learning Sim-to-Real Humanoid Locomotion in 15 Minutes
por: Seo, Younggyo, et al.
Publicado: (2025)
por: Seo, Younggyo, et al.
Publicado: (2025)
Schrodinger Audio-Visual Editor: Object-Level Audiovisual Removal
por: Xu, Weihan, et al.
Publicado: (2025)
por: Xu, Weihan, et al.
Publicado: (2025)
From Generated Human Videos to Physically Plausible Robot Trajectories
por: Ni, James, et al.
Publicado: (2025)
por: Ni, James, et al.
Publicado: (2025)
A Unified Framework for Model Editing
por: Gupta, Akshat, et al.
Publicado: (2024)
por: Gupta, Akshat, et al.
Publicado: (2024)
Is Bigger Edit Batch Size Always Better? -- An Empirical Study on Model Editing with Llama-3
por: Yoon, Junsang, et al.
Publicado: (2024)
por: Yoon, Junsang, et al.
Publicado: (2024)
Rebuilding ROME : Resolving Model Collapse during Sequential Model Editing
por: Gupta, Akshat, et al.
Publicado: (2024)
por: Gupta, Akshat, et al.
Publicado: (2024)
Geometric Interpretation of Layer Normalization and a Comparative Analysis with RMSNorm
por: Gupta, Akshat, et al.
Publicado: (2024)
por: Gupta, Akshat, et al.
Publicado: (2024)
Model Editing at Scale leads to Gradual and Catastrophic Forgetting
por: Gupta, Akshat, et al.
Publicado: (2024)
por: Gupta, Akshat, et al.
Publicado: (2024)
Self-Assessment Tests are Unreliable Measures of LLM Personality
por: Gupta, Akshat, et al.
Publicado: (2023)
por: Gupta, Akshat, et al.
Publicado: (2023)
Multimodal Segmentation for Vocal Tract Modeling
por: Jain, Rishi, et al.
Publicado: (2024)
por: Jain, Rishi, et al.
Publicado: (2024)
Ejemplares similares
-
Self-Supervised Audio-Visual Soundscape Stylization
por: Li, Tingle, et al.
Publicado: (2024) -
Audio Texture Manipulation by Exemplar-Based Analogy
por: Cheng, Kan Jen, et al.
Publicado: (2025) -
Prioritized Generative Replay
por: Wang, Renhao, et al.
Publicado: (2024) -
Interactive Task Planning with Language Models
por: Li, Boyi, et al.
Publicado: (2023) -
Sounding that Object: Interactive Object-Aware Image to Audio Generation
por: Li, Tingle, et al.
Publicado: (2025)