From Frames to Sequences: Temporally Consistent Human-Centric Dense Prediction
Fuente:
arXiv
Salvato in:
| Autori principali: | Miao, Xingyu, Dong, Junting, Zhao, Qin, Yang, Yuhang, Chen, Junhao, Long, Yang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Rethinking Temporal Consistency in Video Object-Centric Learning: From Prediction to Correspondence
di: Li, Zhiyuan, et al.
Pubblicazione: (2026)
di: Li, Zhiyuan, et al.
Pubblicazione: (2026)
TrajVG: 3D Trajectory-Coupled Visual Geometry Learning
di: Miao, Xingyu, et al.
Pubblicazione: (2026)
di: Miao, Xingyu, et al.
Pubblicazione: (2026)
ShotVL: Human-Centric Highlight Frame Retrieval via Language Queries
di: Xue, Wangyu, et al.
Pubblicazione: (2024)
di: Xue, Wangyu, et al.
Pubblicazione: (2024)
From Frames to Events: Rethinking Evaluation in Human-Centric Video Anomaly Detection
di: Rashvand, Narges, et al.
Pubblicazione: (2026)
di: Rashvand, Narges, et al.
Pubblicazione: (2026)
SIGMAN:Scaling 3D Human Gaussian Generation with Millions of Assets
di: Yang, Yuhang, et al.
Pubblicazione: (2025)
di: Yang, Yuhang, et al.
Pubblicazione: (2025)
STCDiT: Spatio-Temporally Consistent Diffusion Transformer for High-Quality Video Super-Resolution
di: Chen, Junyang, et al.
Pubblicazione: (2025)
di: Chen, Junyang, et al.
Pubblicazione: (2025)
DanceTogether! Identity-Preserving Multi-Person Interactive Video Generation
di: Chen, Junhao, et al.
Pubblicazione: (2025)
di: Chen, Junhao, et al.
Pubblicazione: (2025)
Internalizing Temporal Consistency in Video Object-Centric Learning without Explicit Regularization
di: Zhao, Rongzhen, et al.
Pubblicazione: (2026)
di: Zhao, Rongzhen, et al.
Pubblicazione: (2026)
From Sparse to Dense: Spatio-Temporal Fusion for Multi-View 3D Human Pose Estimation with DenseWarper
di: Li, Ling, et al.
Pubblicazione: (2026)
di: Li, Ling, et al.
Pubblicazione: (2026)
Auteur: Language-Driven Cinematographic Framing for Human-Centric Video Generation
di: Kizil, Muhammed Burak, et al.
Pubblicazione: (2026)
di: Kizil, Muhammed Burak, et al.
Pubblicazione: (2026)
Motion Keyframe Interpolation for Any Human Skeleton via Temporally Consistent Point Cloud Sampling and Reconstruction
di: Mo, Clinton, et al.
Pubblicazione: (2024)
di: Mo, Clinton, et al.
Pubblicazione: (2024)
Beyond Static Frames: Temporal Aggregate-and-Restore Vision Transformer for Human Pose Estimation
di: Fang, Hongwei, et al.
Pubblicazione: (2026)
di: Fang, Hongwei, et al.
Pubblicazione: (2026)
Spatial-Temporal-Spectral Unified Modeling for Remote Sensing Dense Prediction
di: Zhao, Sijie, et al.
Pubblicazione: (2025)
di: Zhao, Sijie, et al.
Pubblicazione: (2025)
Rethinking Score Distilling Sampling for 3D Editing and Generation
di: Miao, Xingyu, et al.
Pubblicazione: (2025)
di: Miao, Xingyu, et al.
Pubblicazione: (2025)
BiDense: Binarization for Dense Prediction
di: Yin, Rui, et al.
Pubblicazione: (2024)
di: Yin, Rui, et al.
Pubblicazione: (2024)
V-CORE: Temporally Consistent Video Understanding for Video-LLM
di: Kang, Zhengjian, et al.
Pubblicazione: (2026)
di: Kang, Zhengjian, et al.
Pubblicazione: (2026)
Cycle Consistency in Video Object-Centric Learning
di: Zhao, Rongzhen, et al.
Pubblicazione: (2026)
di: Zhao, Rongzhen, et al.
Pubblicazione: (2026)
Bidirectional Temporal Diffusion Model for Temporally Consistent Human Animation
di: Adiya, Tserendorj, et al.
Pubblicazione: (2023)
di: Adiya, Tserendorj, et al.
Pubblicazione: (2023)
HERM: Benchmarking and Enhancing Multimodal LLMs for Human-Centric Understanding
di: Li, Keliang, et al.
Pubblicazione: (2024)
di: Li, Keliang, et al.
Pubblicazione: (2024)
Beyond Boundary Frames: Context-Centric Video Interpolation with Audio-Visual Semantics
di: Deng, Yuchen, et al.
Pubblicazione: (2025)
di: Deng, Yuchen, et al.
Pubblicazione: (2025)
From Spots to Pixels: Dense Spatial Gene Expression Prediction from Histology Images
di: Zhang, Ruikun, et al.
Pubblicazione: (2025)
di: Zhang, Ruikun, et al.
Pubblicazione: (2025)
MASRA: MLLM-Assisted Semantic-Relational Consistent Alignment for Video Temporal Grounding
di: Ran, Ran, et al.
Pubblicazione: (2026)
di: Ran, Ran, et al.
Pubblicazione: (2026)
Unified Dense Prediction of Video Diffusion
di: Yang, Lehan, et al.
Pubblicazione: (2025)
di: Yang, Lehan, et al.
Pubblicazione: (2025)
DreamFit: Garment-Centric Human Generation via a Lightweight Anything-Dressing Encoder
di: Lin, Ente, et al.
Pubblicazione: (2024)
di: Lin, Ente, et al.
Pubblicazione: (2024)
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference
di: Yang, Yuhang, et al.
Pubblicazione: (2024)
di: Yang, Yuhang, et al.
Pubblicazione: (2024)
IC-Mapper: Instance-Centric Spatio-Temporal Modeling for Online Vectorized Map Construction
di: Zhu, Jiangtong, et al.
Pubblicazione: (2025)
di: Zhu, Jiangtong, et al.
Pubblicazione: (2025)
ConsistentAvatar: Learning to Diffuse Fully Consistent Talking Head Avatar with Temporal Guidance
di: Yang, Haijie, et al.
Pubblicazione: (2024)
di: Yang, Haijie, et al.
Pubblicazione: (2024)
TRACE: Temporally Reliable Anatomically-Conditioned 3D CT Generation with Enhanced Efficiency
di: Shao, Minye, et al.
Pubblicazione: (2025)
di: Shao, Minye, et al.
Pubblicazione: (2025)
Storynizor: Consistent Story Generation via Inter-Frame Synchronized and Shuffled ID Injection
di: Ma, Yuhang, et al.
Pubblicazione: (2024)
di: Ma, Yuhang, et al.
Pubblicazione: (2024)
Vision Transformers: From Semantic Segmentation to Dense Prediction
di: Zhang, Li, et al.
Pubblicazione: (2022)
di: Zhang, Li, et al.
Pubblicazione: (2022)
Zero-Shot Video Translation and Editing with Frame Spatial-Temporal Correspondence
di: Yang, Shuai, et al.
Pubblicazione: (2025)
di: Yang, Shuai, et al.
Pubblicazione: (2025)
Decoding Visual Neural Representations by Multimodal with Dynamic Balancing
di: sun, Kaili, et al.
Pubblicazione: (2025)
di: sun, Kaili, et al.
Pubblicazione: (2025)
Deep Learning in Concealed Dense Prediction
di: Zhao, Pancheng, et al.
Pubblicazione: (2025)
di: Zhao, Pancheng, et al.
Pubblicazione: (2025)
VEOcc: Voxel-Centric Online Semantic Occupancy Prediction For Embodied Scene Understanding
di: Wang, Ruoyu, et al.
Pubblicazione: (2026)
di: Wang, Ruoyu, et al.
Pubblicazione: (2026)
From Recognition to Prediction: Leveraging Sequence Reasoning for Action Anticipation
di: Liu, Xin, et al.
Pubblicazione: (2024)
di: Liu, Xin, et al.
Pubblicazione: (2024)
TDViT: Temporal Dilated Video Transformer for Dense Video Tasks
di: Sun, Guanxiong, et al.
Pubblicazione: (2024)
di: Sun, Guanxiong, et al.
Pubblicazione: (2024)
Drawing2CAD: Sequence-to-Sequence Learning for CAD Generation from Vector Drawings
di: Qin, Feiwei, et al.
Pubblicazione: (2025)
di: Qin, Feiwei, et al.
Pubblicazione: (2025)
Memory Efficient Transformer Adapter for Dense Predictions
di: Zhang, Dong, et al.
Pubblicazione: (2025)
di: Zhang, Dong, et al.
Pubblicazione: (2025)
Frame Order Matters: A Temporal Sequence-Aware Model for Few-Shot Action Recognition
di: Li, Bozheng, et al.
Pubblicazione: (2024)
di: Li, Bozheng, et al.
Pubblicazione: (2024)
OmniHuman: A Large-scale Dataset and Benchmark for Human-Centric Video Generation
di: Zhu, Lei, et al.
Pubblicazione: (2026)
di: Zhu, Lei, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Rethinking Temporal Consistency in Video Object-Centric Learning: From Prediction to Correspondence
di: Li, Zhiyuan, et al.
Pubblicazione: (2026) -
TrajVG: 3D Trajectory-Coupled Visual Geometry Learning
di: Miao, Xingyu, et al.
Pubblicazione: (2026) -
ShotVL: Human-Centric Highlight Frame Retrieval via Language Queries
di: Xue, Wangyu, et al.
Pubblicazione: (2024) -
From Frames to Events: Rethinking Evaluation in Human-Centric Video Anomaly Detection
di: Rashvand, Narges, et al.
Pubblicazione: (2026) -
SIGMAN:Scaling 3D Human Gaussian Generation with Millions of Assets
di: Yang, Yuhang, et al.
Pubblicazione: (2025)