iMotion-LLM: Instruction-Conditioned Trajectory Generation
Fuente:
arXiv
Guardado en:
| Autores principales: | Felemban, Abdulwahab, Hroub, Nussair, Ding, Jian, Abdelrahman, Eslam, Shen, Xiaoqian, Mohamed, Abduallah, Elhoseiny, Mohamed |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
por: Ataallah, Kirolos, et al.
Publicado: (2024)
por: Ataallah, Kirolos, et al.
Publicado: (2024)
StoryGPT-V: Large Language Models as Consistent Story Visualizers
por: Shen, Xiaoqian, et al.
Publicado: (2023)
por: Shen, Xiaoqian, et al.
Publicado: (2023)
FishNet++: Analyzing the capabilities of Multimodal Large Language Models in marine biology
por: Khan, Faizan Farooq, et al.
Publicado: (2025)
por: Khan, Faizan Farooq, et al.
Publicado: (2025)
CoT3DRef: Chain-of-Thoughts Data-Efficient 3D Visual Grounding
por: Abdelrahman, Eslam, et al.
Publicado: (2023)
por: Abdelrahman, Eslam, et al.
Publicado: (2023)
Kestrel: 3D Multimodal LLM for Part-Aware Grounded Description
por: Ahmed, Mahmoud, et al.
Publicado: (2024)
por: Ahmed, Mahmoud, et al.
Publicado: (2024)
Goldfish: Vision-Language Understanding of Arbitrarily Long Videos
por: Ataallah, Kirolos, et al.
Publicado: (2024)
por: Ataallah, Kirolos, et al.
Publicado: (2024)
How Well Can Vision Language Models See Image Details?
por: Gou, Chenhui, et al.
Publicado: (2024)
por: Gou, Chenhui, et al.
Publicado: (2024)
InfiniBench: A Benchmark for Large Multi-Modal Models in Long-Form Movies and TV Shows
por: Ataallah, Kirolos, et al.
Publicado: (2024)
por: Ataallah, Kirolos, et al.
Publicado: (2024)
Vgent: Graph-based Retrieval-Reasoning-Augmented Generation For Long Video Understanding
por: Shen, Xiaoqian, et al.
Publicado: (2025)
por: Shen, Xiaoqian, et al.
Publicado: (2025)
ToddlerDiffusion: Interactive Structured Image Generation with Cascaded Schrödinger Bridge
por: Abdelrahman, Eslam, et al.
Publicado: (2023)
por: Abdelrahman, Eslam, et al.
Publicado: (2023)
VRSBench: A Versatile Vision-Language Benchmark Dataset for Remote Sensing Image Understanding
por: Li, Xiang, et al.
Publicado: (2024)
por: Li, Xiang, et al.
Publicado: (2024)
Zoom-Zero: Reinforced Coarse-to-Fine Video Understanding via Temporal Zoom-in
por: Shen, Xiaoqian, et al.
Publicado: (2025)
por: Shen, Xiaoqian, et al.
Publicado: (2025)
WikiAutoGen: Towards Multi-Modal Wikipedia-Style Article Generation
por: Yang, Zhongyu, et al.
Publicado: (2025)
por: Yang, Zhongyu, et al.
Publicado: (2025)
A Dynamic Programming Framework for Discovering Count and Values of Multilevel Image Thresholding
por: Hegazy, Eslam, et al.
Publicado: (2026)
por: Hegazy, Eslam, et al.
Publicado: (2026)
INSTA-YOLO: Real-Time Instance Segmentation
por: Mohamed, Eslam, et al.
Publicado: (2021)
por: Mohamed, Eslam, et al.
Publicado: (2021)
Overcoming Generic Knowledge Loss with Selective Parameter Update
por: Zhang, Wenxuan, et al.
Publicado: (2023)
por: Zhang, Wenxuan, et al.
Publicado: (2023)
Neural Catalog: Scaling Species Recognition with Catalog of Life-Augmented Generation
por: Khan, Faizan Farooq, et al.
Publicado: (2025)
por: Khan, Faizan Farooq, et al.
Publicado: (2025)
Time Blindness: Why Video-Language Models Can't See What Humans Can?
por: Upadhyay, Ujjwal, et al.
Publicado: (2025)
por: Upadhyay, Ujjwal, et al.
Publicado: (2025)
Coordinating Multiple Conditions for Trajectory-Controlled Human Motion Generation
por: Cai, Deli, et al.
Publicado: (2026)
por: Cai, Deli, et al.
Publicado: (2026)
Efficient Self-supervised Vision Pretraining with Local Masked Reconstruction
por: Chen, Jun, et al.
Publicado: (2022)
por: Chen, Jun, et al.
Publicado: (2022)
ReefNet: A Large-Scale Dataset and Benchmark for Fine-Grained Coral Reef Recognition
por: Felemban, Abdulwahab, et al.
Publicado: (2025)
por: Felemban, Abdulwahab, et al.
Publicado: (2025)
The Devil is in the EOS: Sequence Training for Detailed Image Captioning
por: Mohamed, Abdelrahman, et al.
Publicado: (2025)
por: Mohamed, Abdelrahman, et al.
Publicado: (2025)
3DCoMPaT200: Language-Grounded Compositional Understanding of Parts and Materials of 3D Shapes
por: Ahmed, Mahmoud, et al.
Publicado: (2025)
por: Ahmed, Mahmoud, et al.
Publicado: (2025)
Controlling Structured Output Representations from Attributes using Conditional Generative Models
por: Debbagh, Mohamed
Publicado: (2023)
por: Debbagh, Mohamed
Publicado: (2023)
From Statics to Dynamics: Physics-Aware Image Editing with Latent Transition Priors
por: Zhao, Liangbing, et al.
Publicado: (2026)
por: Zhao, Liangbing, et al.
Publicado: (2026)
AI Art Neural Constellation: Revealing the Collective and Contrastive State of AI-Generated and Human Art
por: Khan, Faizan Farooq, et al.
Publicado: (2024)
por: Khan, Faizan Farooq, et al.
Publicado: (2024)
iLRM: An Iterative Large 3D Reconstruction Model
por: Kang, Gyeongjin, et al.
Publicado: (2025)
por: Kang, Gyeongjin, et al.
Publicado: (2025)
Self-Consistency for LLM-Based Motion Trajectory Generation and Verification
por: Ma, Jiaju, et al.
Publicado: (2026)
por: Ma, Jiaju, et al.
Publicado: (2026)
Domain-Aware Continual Zero-Shot Learning
por: Yi, Kai, et al.
Publicado: (2021)
por: Yi, Kai, et al.
Publicado: (2021)
IKMo: Image-Keyframed Motion Generation with Trajectory-Pose Conditioned Motion Diffusion Model
por: Zhao, Yang, et al.
Publicado: (2025)
por: Zhao, Yang, et al.
Publicado: (2025)
Category-level Text-to-Image Retrieval Improved: Bridging the Domain Gap with Diffusion Models and Vision Encoders
por: Khan, Faizan Farooq, et al.
Publicado: (2025)
por: Khan, Faizan Farooq, et al.
Publicado: (2025)
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding
por: Kim, Younggun, et al.
Publicado: (2025)
por: Kim, Younggun, et al.
Publicado: (2025)
Continual Learning on a Diet: Learning from Sparsely Labeled Streams Under Constrained Computation
por: Zhang, Wenxuan, et al.
Publicado: (2024)
por: Zhang, Wenxuan, et al.
Publicado: (2024)
Openstory++: A Large-scale Dataset and Benchmark for Instance-aware Open-domain Visual Storytelling
por: Ye, Zilyu, et al.
Publicado: (2024)
por: Ye, Zilyu, et al.
Publicado: (2024)
Step-by-step Layered Design Generation
por: Khan, Faizan Farooq, et al.
Publicado: (2025)
por: Khan, Faizan Farooq, et al.
Publicado: (2025)
IRG-MotionLLM: Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation
por: Li, Yuan-Ming, et al.
Publicado: (2025)
por: Li, Yuan-Ming, et al.
Publicado: (2025)
Motion Prompting: Controlling Video Generation with Motion Trajectories
por: Geng, Daniel, et al.
Publicado: (2024)
por: Geng, Daniel, et al.
Publicado: (2024)
RDM: Recurrent Diffusion Model for Human Motion Generation
por: Mohamed, Mirgahney, et al.
Publicado: (2024)
por: Mohamed, Mirgahney, et al.
Publicado: (2024)
MG-MotionLLM: A Unified Framework for Motion Comprehension and Generation across Multiple Granularities
por: Wu, Bizhu, et al.
Publicado: (2025)
por: Wu, Bizhu, et al.
Publicado: (2025)
Mojito: Motion Trajectory and Intensity Control for Video Generation
por: He, Xuehai, et al.
Publicado: (2024)
por: He, Xuehai, et al.
Publicado: (2024)
Ejemplares similares
-
MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
por: Ataallah, Kirolos, et al.
Publicado: (2024) -
StoryGPT-V: Large Language Models as Consistent Story Visualizers
por: Shen, Xiaoqian, et al.
Publicado: (2023) -
FishNet++: Analyzing the capabilities of Multimodal Large Language Models in marine biology
por: Khan, Faizan Farooq, et al.
Publicado: (2025) -
CoT3DRef: Chain-of-Thoughts Data-Efficient 3D Visual Grounding
por: Abdelrahman, Eslam, et al.
Publicado: (2023) -
Kestrel: 3D Multimodal LLM for Part-Aware Grounded Description
por: Ahmed, Mahmoud, et al.
Publicado: (2024)