Saved in:
| Main Authors: | Liu, Kun, Liu, Qi, Liu, Xinchen, Li, Jie, Zhang, Yongdong, Luo, Jiebo, He, Xiaodong, Liu, Wu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2503.23715 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Paradigm Shift: Fully End-to-End Training for Temporal Sentence Grounding in Videos
by: He, Allen, et al.
Published: (2026)
by: He, Allen, et al.
Published: (2026)
T-SVG: Text-Driven Stereoscopic Video Generation
by: Jin, Qiao, et al.
Published: (2024)
by: Jin, Qiao, et al.
Published: (2024)
SigFormer: Sparse Signal-Guided Transformer for Multi-Modal Human Action Segmentation
by: Liu, Qi, et al.
Published: (2023)
by: Liu, Qi, et al.
Published: (2023)
Motion Capture from Inertial and Vision Sensors
by: Chen, Xiaodong, et al.
Published: (2024)
by: Chen, Xiaodong, et al.
Published: (2024)
OmniPrism: Learning Disentangled Visual Concept for Image Generation
by: Li, Yangyang, et al.
Published: (2024)
by: Li, Yangyang, et al.
Published: (2024)
OmniHuman: A Large-scale Dataset and Benchmark for Human-Centric Video Generation
by: Zhu, Lei, et al.
Published: (2026)
by: Zhu, Lei, et al.
Published: (2026)
Scaling Down Text Encoders of Text-to-Image Diffusion Models
by: Wang, Lifu, et al.
Published: (2025)
by: Wang, Lifu, et al.
Published: (2025)
It Takes Two: Accurate Gait Recognition in the Wild via Cross-granularity Alignment
by: Zheng, Jinkai, et al.
Published: (2024)
by: Zheng, Jinkai, et al.
Published: (2024)
Latent-Reframe: Enabling Camera Control for Video Diffusion Model without Training
by: Zhou, Zhenghong, et al.
Published: (2024)
by: Zhou, Zhenghong, et al.
Published: (2024)
VideoSeek: Long-Horizon Video Agent with Tool-Guided Seeking
by: Lin, Jingyang, et al.
Published: (2026)
by: Lin, Jingyang, et al.
Published: (2026)
LOVO: Efficient Complex Object Query in Large-Scale Video Datasets
by: Liu, Yuxin, et al.
Published: (2025)
by: Liu, Yuxin, et al.
Published: (2025)
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
by: Wang, Yi, et al.
Published: (2023)
by: Wang, Yi, et al.
Published: (2023)
OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation
by: Zhou, Donghao, et al.
Published: (2026)
by: Zhou, Donghao, et al.
Published: (2026)
Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human Interactions
by: Xu, Liang, et al.
Published: (2025)
by: Xu, Liang, et al.
Published: (2025)
VAGNet: Grounding 3D Affordance from Human-Object Interactions in Videos
by: Mao, Aihua, et al.
Published: (2026)
by: Mao, Aihua, et al.
Published: (2026)
Unleashing Hour-Scale Video Training for Long Video-Language Understanding
by: Lin, Jingyang, et al.
Published: (2025)
by: Lin, Jingyang, et al.
Published: (2025)
MaMi-HOI: Harmonizing Global Kinematics and Local Geometry for Human-Object Interaction Generation
by: Wang, Hao, et al.
Published: (2026)
by: Wang, Hao, et al.
Published: (2026)
HumanNeRF-SE: A Simple yet Effective Approach to Animate HumanNeRF with Diverse Poses
by: Ma, Caoyuan, et al.
Published: (2023)
by: Ma, Caoyuan, et al.
Published: (2023)
SportsHHI: A Dataset for Human-Human Interaction Detection in Sports Videos
by: Wu, Tao, et al.
Published: (2024)
by: Wu, Tao, et al.
Published: (2024)
HOI4D: A 4D Egocentric Dataset for Category-Level Human-Object Interaction
by: Liu, Yunze, et al.
Published: (2022)
by: Liu, Yunze, et al.
Published: (2022)
CORE4D: A 4D Human-Object-Human Interaction Dataset for Collaborative Object REarrangement
by: Liu, Yun, et al.
Published: (2024)
by: Liu, Yun, et al.
Published: (2024)
LSVOS Challenge Report: Large-scale Complex and Long Video Object Segmentation
by: Ding, Henghui, et al.
Published: (2024)
by: Ding, Henghui, et al.
Published: (2024)
Diffusion Transformer-to-Mamba Distillation for High-Resolution Image Generation
by: Yao, Yuan, et al.
Published: (2025)
by: Yao, Yuan, et al.
Published: (2025)
MOVi: Training-free Text-conditioned Multi-Object Video Generation
by: Rahman, Aimon, et al.
Published: (2025)
by: Rahman, Aimon, et al.
Published: (2025)
VideoAfford: Grounding 3D Affordance from Human-Object-Interaction Videos via Multimodal Large Language Model
by: Wang, Hanqing, et al.
Published: (2026)
by: Wang, Hanqing, et al.
Published: (2026)
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding
by: Chen, Zhen, et al.
Published: (2025)
by: Chen, Zhen, et al.
Published: (2025)
Human vs. LMMs: Exploring the Discrepancy in Emoji Interpretation and Usage in Digital Communication
by: Lyu, Hanjia, et al.
Published: (2024)
by: Lyu, Hanjia, et al.
Published: (2024)
Controllable Human-Object Interaction Synthesis
by: Li, Jiaman, et al.
Published: (2023)
by: Li, Jiaman, et al.
Published: (2023)
Interact-Custom: Customized Human Object Interaction Image Generation
by: Xu, Zhu, et al.
Published: (2025)
by: Xu, Zhu, et al.
Published: (2025)
A Survey of Interactive Generative Video
by: Yu, Jiwen, et al.
Published: (2025)
by: Yu, Jiwen, et al.
Published: (2025)
Human-Object Interaction from Human-Level Instructions
by: Wu, Zhen, et al.
Published: (2024)
by: Wu, Zhen, et al.
Published: (2024)
Incremental Human-Object Interaction Detection with Invariant Relation Representation Learning
by: Wei, Yana, et al.
Published: (2025)
by: Wei, Yana, et al.
Published: (2025)
EvalCrafter: Benchmarking and Evaluating Large Video Generation Models
by: Liu, Yaofang, et al.
Published: (2023)
by: Liu, Yaofang, et al.
Published: (2023)
ATRNet-STAR: A Large Dataset and Benchmark Towards Remote Sensing Object Recognition in the Wild
by: Liu, Yongxiang, et al.
Published: (2025)
by: Liu, Yongxiang, et al.
Published: (2025)
Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting
by: Tang, Yunlong, et al.
Published: (2025)
by: Tang, Yunlong, et al.
Published: (2025)
Mask$^2$DiT: Dual Mask-based Diffusion Transformer for Multi-Scene Long Video Generation
by: Qi, Tianhao, et al.
Published: (2025)
by: Qi, Tianhao, et al.
Published: (2025)
JavisDiT++: Unified Modeling and Optimization for Joint Audio-Video Generation
by: Liu, Kai, et al.
Published: (2026)
by: Liu, Kai, et al.
Published: (2026)
Lumina-Video: Efficient and Flexible Video Generation with Multi-scale Next-DiT
by: Liu, Dongyang, et al.
Published: (2025)
by: Liu, Dongyang, et al.
Published: (2025)
XS-VID: An Extremely Small Video Object Detection Dataset
by: Guo, Jiahao, et al.
Published: (2024)
by: Guo, Jiahao, et al.
Published: (2024)
A Review of Human-Object Interaction Detection
by: Wang, Yuxiao, et al.
Published: (2024)
by: Wang, Yuxiao, et al.
Published: (2024)
Similar Items
-
A Paradigm Shift: Fully End-to-End Training for Temporal Sentence Grounding in Videos
by: He, Allen, et al.
Published: (2026) -
T-SVG: Text-Driven Stereoscopic Video Generation
by: Jin, Qiao, et al.
Published: (2024) -
SigFormer: Sparse Signal-Guided Transformer for Multi-Modal Human Action Segmentation
by: Liu, Qi, et al.
Published: (2023) -
Motion Capture from Inertial and Vision Sensors
by: Chen, Xiaodong, et al.
Published: (2024) -
OmniPrism: Learning Disentangled Visual Concept for Image Generation
by: Li, Yangyang, et al.
Published: (2024)