Language-free Compositional Action Generation via Decoupling Refinement
Fuente:
arXiv
Guardado en:
| Autores principales: | Liu, Xiao, Chen, Guangyi, Tang, Yansong, Wang, Guangrun, Zhang, Xiao-Ping, Lim, Ser-Nam |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
VideoMerge: Towards Training-free Long Video Generation
por: Zhang, Siyang, et al.
Publicado: (2025)
por: Zhang, Siyang, et al.
Publicado: (2025)
Towards Chunk-Wise Generation for Long Videos
por: Zhang, Siyang, et al.
Publicado: (2024)
por: Zhang, Siyang, et al.
Publicado: (2024)
Narrative Action Evaluation with Prompt-Guided Multimodal Interaction
por: Zhang, Shiyi, et al.
Publicado: (2024)
por: Zhang, Shiyi, et al.
Publicado: (2024)
DiReCT: Disentangled Regularization of Contrastive Trajectories for Physics-Refined Video Generation
por: Meyarian, Abolfazl, et al.
Publicado: (2026)
por: Meyarian, Abolfazl, et al.
Publicado: (2026)
Beyond Generation: Unlocking Universal Editing via Self-Supervised Fine-Tuning
por: Chen, Harold Haodong, et al.
Publicado: (2024)
por: Chen, Harold Haodong, et al.
Publicado: (2024)
Mitigating Dialogue Hallucination for Large Vision Language Models via Adversarial Instruction Tuning
por: Park, Dongmin, et al.
Publicado: (2024)
por: Park, Dongmin, et al.
Publicado: (2024)
VG-Refiner: Towards Tool-Refined Referring Grounded Reasoning via Agentic Reinforcement Learning
por: Wang, Yuji, et al.
Publicado: (2025)
por: Wang, Yuji, et al.
Publicado: (2025)
Scene Co-pilot: Procedural Text to Video Generation with Human in the Loop
por: Qian, Zhaofang, et al.
Publicado: (2024)
por: Qian, Zhaofang, et al.
Publicado: (2024)
Towards Unified 3D Object Detection via Algorithm and Data Unification
por: Li, Zhuoling, et al.
Publicado: (2024)
por: Li, Zhuoling, et al.
Publicado: (2024)
FSViewFusion: Few-Shots View Generation of Novel Objects
por: Hussain, Rukhshanda, et al.
Publicado: (2024)
por: Hussain, Rukhshanda, et al.
Publicado: (2024)
ATP-LLaVA: Adaptive Token Pruning for Large Vision Language Models
por: Ye, Xubing, et al.
Publicado: (2024)
por: Ye, Xubing, et al.
Publicado: (2024)
GS: Generative Segmentation via Label Diffusion
por: Chen, Yuhao, et al.
Publicado: (2025)
por: Chen, Yuhao, et al.
Publicado: (2025)
Temporal Action Localization with Cross Layer Task Decoupling and Refinement
por: Li, Qiang, et al.
Publicado: (2024)
por: Li, Qiang, et al.
Publicado: (2024)
Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation
por: Chen, Harold Haodong, et al.
Publicado: (2025)
por: Chen, Harold Haodong, et al.
Publicado: (2025)
BOOKAGENT: Orchestrating Safety-Aware Visual Narratives via Multi-Agent Cognitive Calibration
por: Gao, Bo, et al.
Publicado: (2026)
por: Gao, Bo, et al.
Publicado: (2026)
DreamMask: Boosting Open-vocabulary Panoptic Segmentation with Synthetic Data
por: Tu, Yuanpeng, et al.
Publicado: (2025)
por: Tu, Yuanpeng, et al.
Publicado: (2025)
Fast Encoding and Decoding for Implicit Video Representation
por: Chen, Hao, et al.
Publicado: (2024)
por: Chen, Hao, et al.
Publicado: (2024)
FDDet: Frequency-Decoupling for Boundary Refinement in Temporal Action Detection
por: Zhu, Xinnan, et al.
Publicado: (2025)
por: Zhu, Xinnan, et al.
Publicado: (2025)
Video Decomposition Prior: A Methodology to Decompose Videos into Layers
por: Shrivastava, Gaurav, et al.
Publicado: (2024)
por: Shrivastava, Gaurav, et al.
Publicado: (2024)
AirSketch: Generative Motion to Sketch
por: Lim, Hui Xian Grace, et al.
Publicado: (2024)
por: Lim, Hui Xian Grace, et al.
Publicado: (2024)
Composing Object Relations and Attributes for Image-Text Matching
por: Pham, Khoi, et al.
Publicado: (2024)
por: Pham, Khoi, et al.
Publicado: (2024)
Physical Autoregressive Model for Robotic Manipulation without Action Pretraining
por: Song, Zijian, et al.
Publicado: (2025)
por: Song, Zijian, et al.
Publicado: (2025)
Towards Self-Refinement of Vision-Language Models with Triangular Consistency
por: Deng, Yunlong, et al.
Publicado: (2025)
por: Deng, Yunlong, et al.
Publicado: (2025)
UML-CoT: Structured Reasoning and Planning with Unified Modeling Language for Robotic Room Cleaning
por: Chen, Hongyu, et al.
Publicado: (2025)
por: Chen, Hongyu, et al.
Publicado: (2025)
Zero-shot Synthetic Video Realism Enhancement via Structure-aware Denoising
por: Wang, Yifan, et al.
Publicado: (2025)
por: Wang, Yifan, et al.
Publicado: (2025)
What can Off-the-Shelves Large Multi-Modal Models do for Dynamic Scene Graph Generation?
por: Cui, Xuanming, et al.
Publicado: (2025)
por: Cui, Xuanming, et al.
Publicado: (2025)
DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness
por: Mohammadshirazi, Ahmad, et al.
Publicado: (2024)
por: Mohammadshirazi, Ahmad, et al.
Publicado: (2024)
Geometry aware 3D generation from in-the-wild images in ImageNet
por: Shen, Qijia, et al.
Publicado: (2024)
por: Shen, Qijia, et al.
Publicado: (2024)
IteRPrimE: Zero-shot Referring Image Segmentation with Iterative Grad-CAM Refinement and Primary Word Emphasis
por: Wang, Yuji, et al.
Publicado: (2025)
por: Wang, Yuji, et al.
Publicado: (2025)
Distilling Vision-Language Pretraining for Efficient Cross-Modal Retrieval
por: Jang, Young Kyun, et al.
Publicado: (2024)
por: Jang, Young Kyun, et al.
Publicado: (2024)
LARM: Large Auto-Regressive Model for Long-Horizon Embodied Intelligence
por: Li, Zhuoling, et al.
Publicado: (2024)
por: Li, Zhuoling, et al.
Publicado: (2024)
HumanGenesis: Agent-Based Geometric and Generative Modeling for Synthetic Human Dynamics
por: Li, Weiqi, et al.
Publicado: (2025)
por: Li, Weiqi, et al.
Publicado: (2025)
ACD: Direct Conditional Control for Video Diffusion Models via Attention Supervision
por: Li, Weiqi, et al.
Publicado: (2025)
por: Li, Weiqi, et al.
Publicado: (2025)
Towards Cross-modal Backward-compatible Representation Learning for Vision-Language Models
por: Jang, Young Kyun, et al.
Publicado: (2024)
por: Jang, Young Kyun, et al.
Publicado: (2024)
Niagara: Normal-Integrated Geometric Affine Field for Scene Reconstruction from a Single View
por: Wu, Xianzu, et al.
Publicado: (2025)
por: Wu, Xianzu, et al.
Publicado: (2025)
TAG: Target-Agnostic Guidance for Stable Object-Centric Inference in Vision-Language-Action Models
por: Zhou, Jiaying, et al.
Publicado: (2026)
por: Zhou, Jiaying, et al.
Publicado: (2026)
CLAP: Contrastive Latent Action Pretraining for Learning Vision-Language-Action Models from Human Videos
por: Zhang, Chubin, et al.
Publicado: (2026)
por: Zhang, Chubin, et al.
Publicado: (2026)
Compositional Zero-shot Learning via Progressive Language-based Observations
por: Li, Lin, et al.
Publicado: (2023)
por: Li, Lin, et al.
Publicado: (2023)
Visual Delta Generator with Large Multi-modal Models for Semi-supervised Composed Image Retrieval
por: Jang, Young Kyun, et al.
Publicado: (2024)
por: Jang, Young Kyun, et al.
Publicado: (2024)
AlignVid: Training-Free Attention Scaling for Semantic Fidelity in Text-Guided Image-to-Video Generation
por: Liu, Yexin, et al.
Publicado: (2025)
por: Liu, Yexin, et al.
Publicado: (2025)
Ejemplares similares
-
VideoMerge: Towards Training-free Long Video Generation
por: Zhang, Siyang, et al.
Publicado: (2025) -
Towards Chunk-Wise Generation for Long Videos
por: Zhang, Siyang, et al.
Publicado: (2024) -
Narrative Action Evaluation with Prompt-Guided Multimodal Interaction
por: Zhang, Shiyi, et al.
Publicado: (2024) -
DiReCT: Disentangled Regularization of Contrastive Trajectories for Physics-Refined Video Generation
por: Meyarian, Abolfazl, et al.
Publicado: (2026) -
Beyond Generation: Unlocking Universal Editing via Self-Supervised Fine-Tuning
por: Chen, Harold Haodong, et al.
Publicado: (2024)