Towards Consistent Long-Term Pose Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Yayuan, Bellos, Filippos, Corso, Jason |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Effective Human-in-the-Loop Assistive AI Agents
by: Bellos, Filippos, et al.
Published: (2025)
by: Bellos, Filippos, et al.
Published: (2025)
Mistake Attribution: Fine-Grained Mistake Understanding in Egocentric Videos
by: Li, Yayuan, et al.
Published: (2025)
by: Li, Yayuan, et al.
Published: (2025)
EchoVQA: Enabling Conversational Assistance for Point-of-Care Cardiac Ultrasound
by: Bellos, Filippos, et al.
Published: (2026)
by: Bellos, Filippos, et al.
Published: (2026)
HANDI: Hand-Centric Text-and-Image Conditioned Video Generation
by: Li, Yayuan, et al.
Published: (2024)
by: Li, Yayuan, et al.
Published: (2024)
When to Think and When to Look: Uncertainty-Guided Lookback
by: Bi, Jing, et al.
Published: (2025)
by: Bi, Jing, et al.
Published: (2025)
Omni-Judge: Can Omni-LLMs Serve as Human-Aligned Judges for Text-Conditioned Audio-Video Generation?
by: Liang, Susan, et al.
Published: (2026)
by: Liang, Susan, et al.
Published: (2026)
Follow Your Heart: Landmark-Guided Transducer Pose Scoring for Point-of-Care Echocardiography
by: Guo, Zaiyang, et al.
Published: (2026)
by: Guo, Zaiyang, et al.
Published: (2026)
Memorize-and-Generate: Towards Long-Term Consistency in Real-Time Video Generation
by: Zhu, Tianrui, et al.
Published: (2025)
by: Zhu, Tianrui, et al.
Published: (2025)
Measuring Physical Plausibility of 3D Human Poses Using Physics Simulation
by: Louis, Nathan, et al.
Published: (2025)
by: Louis, Nathan, et al.
Published: (2025)
BiMotion: B-spline Motion for Text-guided Dynamic 3D Character Generation
by: Wang, Miaowei, et al.
Published: (2026)
by: Wang, Miaowei, et al.
Published: (2026)
StableWorld: Towards Stable and Consistent Long Interactive Video Generation
by: Yang, Ying, et al.
Published: (2026)
by: Yang, Ying, et al.
Published: (2026)
Identity-Consistent Multi-Pose Generation of Contactless Fingerprints
by: Pan, Zhiyu, et al.
Published: (2026)
by: Pan, Zhiyu, et al.
Published: (2026)
WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling
by: Sun, Wenqiang, et al.
Published: (2025)
by: Sun, Wenqiang, et al.
Published: (2025)
Substantial, Decomposable, and Invisible: Visual Context Misalignment in Instructional Videos for Physical Tasks
by: Li, Yayuan, et al.
Published: (2026)
by: Li, Yayuan, et al.
Published: (2026)
Multi Positive Contrastive Learning with Pose-Consistent Generated Images
by: Inayoshi, Sho, et al.
Published: (2024)
by: Inayoshi, Sho, et al.
Published: (2024)
Temporally Guided Articulated Hand Pose Tracking in Surgical Videos
by: Louis, Nathan, et al.
Published: (2021)
by: Louis, Nathan, et al.
Published: (2021)
DecMem: Towards Minute-Long Consistent World Generation with Decoupled Memory
by: Yang, Zhenhao, et al.
Published: (2026)
by: Yang, Zhenhao, et al.
Published: (2026)
CoDi: Subject-Consistent and Pose-Diverse Text-to-Image Generation
by: Gao, Zhanxin, et al.
Published: (2025)
by: Gao, Zhanxin, et al.
Published: (2025)
GRIP: Generating Interaction Poses Using Spatial Cues and Latent Consistency
by: Taheri, Omid, et al.
Published: (2023)
by: Taheri, Omid, et al.
Published: (2023)
Temporally Consistent Long-Term Memory for 3D Single Object Tracking
by: Yoo, Jaejoon, et al.
Published: (2026)
by: Yoo, Jaejoon, et al.
Published: (2026)
InfinityHuman: Towards Long-Term Audio-Driven Human
by: Li, Xiaodi, et al.
Published: (2025)
by: Li, Xiaodi, et al.
Published: (2025)
Learn Your Scales: Towards Scale-Consistent Generative Novel View Synthesis
by: Forghani, Fereshteh, et al.
Published: (2025)
by: Forghani, Fereshteh, et al.
Published: (2025)
PoseGen: In-Context LoRA Finetuning for Pose-Controllable Long Human Video Generation
by: He, Jingxuan, et al.
Published: (2025)
by: He, Jingxuan, et al.
Published: (2025)
Can Generative Video Models Help Pose Estimation?
by: Cai, Ruojin, et al.
Published: (2024)
by: Cai, Ruojin, et al.
Published: (2024)
Gen3DEval: Using vLLMs for Automatic Evaluation of Generated 3D Objects
by: Maiti, Shalini, et al.
Published: (2025)
by: Maiti, Shalini, et al.
Published: (2025)
OmniPose6D: Towards Short-Term Object Pose Tracking in Dynamic Scenes from Monocular RGB
by: Lin, Yunzhi, et al.
Published: (2024)
by: Lin, Yunzhi, et al.
Published: (2024)
Towards Better Robustness: Pose-Free 3D Gaussian Splatting for Arbitrarily Long Videos
by: Dong, Zhen-Hui, et al.
Published: (2025)
by: Dong, Zhen-Hui, et al.
Published: (2025)
Text and Image Are Mutually Beneficial: Enhancing Training-Free Few-Shot Classification with CLIP
by: Li, Yayuan, et al.
Published: (2024)
by: Li, Yayuan, et al.
Published: (2024)
Towards Long Term SLAM on Thermal Imagery
by: Keil, Colin, et al.
Published: (2024)
by: Keil, Colin, et al.
Published: (2024)
Collaborative Face Experts Fusion in Video Generation: Boosting Identity Consistency Across Large Face Poses
by: Wang, Yuji, et al.
Published: (2025)
by: Wang, Yuji, et al.
Published: (2025)
Zero-Shot Coreset Selection via Iterative Subspace Sampling
by: Griffin, Brent A., et al.
Published: (2024)
by: Griffin, Brent A., et al.
Published: (2024)
Geometry Depth Consistency in RGBD Relative Pose Estimation
by: Kumar, Sourav, et al.
Published: (2024)
by: Kumar, Sourav, et al.
Published: (2024)
Phantom-Data : Towards a General Subject-Consistent Video Generation Dataset
by: Chen, Zhuowei, et al.
Published: (2025)
by: Chen, Zhuowei, et al.
Published: (2025)
SCAIL: Towards Studio-Grade Character Animation via In-Context Learning of 3D-Consistent Pose Representations
by: Yan, Wenhao, et al.
Published: (2025)
by: Yan, Wenhao, et al.
Published: (2025)
Anticipating Object State Changes in Long Procedural Videos
by: Manousaki, Victoria, et al.
Published: (2024)
by: Manousaki, Victoria, et al.
Published: (2024)
StoryMaker: Towards Holistic Consistent Characters in Text-to-image Generation
by: Zhou, Zhengguang, et al.
Published: (2024)
by: Zhou, Zhengguang, et al.
Published: (2024)
Temporal Context Consistency Above All: Enhancing Long-Term Anticipation by Learning and Enforcing Temporal Constraints
by: Maté, Alberto, et al.
Published: (2024)
by: Maté, Alberto, et al.
Published: (2024)
CamDirector: Towards Long-Term Coherent Video Trajectory Editing
by: Shi, Zhihao, et al.
Published: (2026)
by: Shi, Zhihao, et al.
Published: (2026)
SecondPose: SE(3)-Consistent Dual-Stream Feature Fusion for Category-Level Pose Estimation
by: Chen, Yamei, et al.
Published: (2023)
by: Chen, Yamei, et al.
Published: (2023)
Pose Magic: Efficient and Temporally Consistent Human Pose Estimation with a Hybrid Mamba-GCN Network
by: Zhang, Xinyi, et al.
Published: (2024)
by: Zhang, Xinyi, et al.
Published: (2024)
Similar Items
-
Towards Effective Human-in-the-Loop Assistive AI Agents
by: Bellos, Filippos, et al.
Published: (2025) -
Mistake Attribution: Fine-Grained Mistake Understanding in Egocentric Videos
by: Li, Yayuan, et al.
Published: (2025) -
EchoVQA: Enabling Conversational Assistance for Point-of-Care Cardiac Ultrasound
by: Bellos, Filippos, et al.
Published: (2026) -
HANDI: Hand-Centric Text-and-Image Conditioned Video Generation
by: Li, Yayuan, et al.
Published: (2024) -
When to Think and When to Look: Uncertainty-Guided Lookback
by: Bi, Jing, et al.
Published: (2025)