DiffVC: A Non-autoregressive Framework Based on Diffusion Model for Video Captioning
Fuente:
arXiv
Guardado en:
| Autores principales: | Wang, Junbo, Fu, Liangyu, Li, Yuke, Zhu, Yining, Jing, Ya, Wu, Xuecheng, Zheng, Jiangbin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
EPIR: An Efficient Patch Tokenization, Integration and Representation Framework for Micro-expression Recognition
por: Wang, Junbo, et al.
Publicado: (2026)
por: Wang, Junbo, et al.
Publicado: (2026)
DiffVC-OSD: One-Step Diffusion-based Perceptual Neural Video Compression Framework
por: Ma, Wenzhuo, et al.
Publicado: (2025)
por: Ma, Wenzhuo, et al.
Publicado: (2025)
DiffVC-RT: Towards Practical Real-Time Diffusion-based Perceptual Neural Video Compression
por: Ma, Wenzhuo, et al.
Publicado: (2026)
por: Ma, Wenzhuo, et al.
Publicado: (2026)
TACR-YOLO: A Real-time Detection Framework for Abnormal Human Behaviors Enhanced with Coordinate and Task-Aware Representations
por: Yin, Xinyi, et al.
Publicado: (2025)
por: Yin, Xinyi, et al.
Publicado: (2025)
FAMNet: Integrating 2D and 3D Features for Micro-expression Recognition via Multi-task Learning and Hierarchical Attention
por: Fu, Liangyu, et al.
Publicado: (2025)
por: Fu, Liangyu, et al.
Publicado: (2025)
VC4VG: Optimizing Video Captions for Text-to-Video Generation
por: Du, Yang, et al.
Publicado: (2025)
por: Du, Yang, et al.
Publicado: (2025)
Diff-3DCap: Shape Captioning with Diffusion Models
por: Shu, Zhenyu, et al.
Publicado: (2025)
por: Shu, Zhenyu, et al.
Publicado: (2025)
DAWN: Dynamic Frame Avatar with Non-autoregressive Diffusion Framework for Talking Head Video Generation
por: Cheng, Hanbo, et al.
Publicado: (2024)
por: Cheng, Hanbo, et al.
Publicado: (2024)
VC-Inspector: Advancing Reference-free Evaluation of Video Captions with Factual Analysis
por: Dipta, Shubhashis Roy, et al.
Publicado: (2025)
por: Dipta, Shubhashis Roy, et al.
Publicado: (2025)
UGC-VideoCaptioner: An Omni UGC Video Detail Caption Model and New Benchmarks
por: Wu, Peiran, et al.
Publicado: (2025)
por: Wu, Peiran, et al.
Publicado: (2025)
OneDiff: A Generalist Model for Image Difference Captioning
por: Hu, Erdong, et al.
Publicado: (2024)
por: Hu, Erdong, et al.
Publicado: (2024)
EventDiff: A Unified and Efficient Diffusion Model Framework for Event-based Video Frame Interpolation
por: Zheng, Hanle, et al.
Publicado: (2025)
por: Zheng, Hanle, et al.
Publicado: (2025)
Mitigate Replication and Copying in Diffusion Models with Generalized Caption and Dual Fusion Enhancement
por: Li, Chenghao, et al.
Publicado: (2023)
por: Li, Chenghao, et al.
Publicado: (2023)
MCF-VC: Mitigate Catastrophic Forgetting in Class-Incremental Learning for Multimodal Video Captioning
por: Xiong, Huiyu, et al.
Publicado: (2024)
por: Xiong, Huiyu, et al.
Publicado: (2024)
3A-YOLO: New Real-Time Object Detectors with Triple Discriminative Awareness and Coordinated Representations
por: Wu, Xuecheng, et al.
Publicado: (2024)
por: Wu, Xuecheng, et al.
Publicado: (2024)
Diff-BGM: A Diffusion Model for Video Background Music Generation
por: Li, Sizhe, et al.
Publicado: (2024)
por: Li, Sizhe, et al.
Publicado: (2024)
MaskINT: Video Editing via Interpolative Non-autoregressive Masked Transformers
por: Ma, Haoyu, et al.
Publicado: (2023)
por: Ma, Haoyu, et al.
Publicado: (2023)
End4: End-to-end Denoising Diffusion for Diffusion-Based Inpainting Detection
por: Wang, Fei, et al.
Publicado: (2025)
por: Wang, Fei, et al.
Publicado: (2025)
Scalable Audio-Visual Masked Autoencoders for Efficient Affective Video Facial Analysis
por: Wu, Xuecheng, et al.
Publicado: (2025)
por: Wu, Xuecheng, et al.
Publicado: (2025)
DaDiff: Domain-aware Diffusion Model for Nighttime UAV Tracking
por: Zuo, Haobo, et al.
Publicado: (2024)
por: Zuo, Haobo, et al.
Publicado: (2024)
Wolf: Dense Video Captioning with a World Summarization Framework
por: Li, Boyi, et al.
Publicado: (2024)
por: Li, Boyi, et al.
Publicado: (2024)
DiffClass: Diffusion-Based Class Incremental Learning
por: Meng, Zichong, et al.
Publicado: (2024)
por: Meng, Zichong, et al.
Publicado: (2024)
LaMamba-Diff: Linear-Time High-Fidelity Diffusion Models Based on Local Attention and Mamba
por: Fu, Yunxiang, et al.
Publicado: (2024)
por: Fu, Yunxiang, et al.
Publicado: (2024)
DiffMesh: A Motion-aware Diffusion Framework for Human Mesh Recovery from Videos
por: Zheng, Ce, et al.
Publicado: (2023)
por: Zheng, Ce, et al.
Publicado: (2023)
AdaQual-Diff: Diffusion-Based Image Restoration via Adaptive Quality Prompting
por: Su, Xin, et al.
Publicado: (2025)
por: Su, Xin, et al.
Publicado: (2025)
HaltingVT: Adaptive Token Halting Transformer for Efficient Video Recognition
por: Wu, Qian, et al.
Publicado: (2024)
por: Wu, Qian, et al.
Publicado: (2024)
DiffHDR: Re-Exposing LDR Videos with Video Diffusion Models
por: Yu, Zhengming, et al.
Publicado: (2026)
por: Yu, Zhengming, et al.
Publicado: (2026)
MagDiff: Multi-Alignment Diffusion for High-Fidelity Video Generation and Editing
por: Zhao, Haoyu, et al.
Publicado: (2023)
por: Zhao, Haoyu, et al.
Publicado: (2023)
ICANet: A Method of Short Video Emotion Recognition Driven by Multimodal Data
por: Wu, Xuecheng, et al.
Publicado: (2022)
por: Wu, Xuecheng, et al.
Publicado: (2022)
DiffSim: Taming Diffusion Models for Evaluating Visual Similarity
por: Song, Yiren, et al.
Publicado: (2024)
por: Song, Yiren, et al.
Publicado: (2024)
IS-Diff: Improving Diffusion-Based Inpainting with Better Initial Seed
por: Lyu, Yongzhe, et al.
Publicado: (2025)
por: Lyu, Yongzhe, et al.
Publicado: (2025)
Non-autoregressive Sequence-to-Sequence Vision-Language Models
por: Shi, Kunyu, et al.
Publicado: (2024)
por: Shi, Kunyu, et al.
Publicado: (2024)
MAMS: Model-Agnostic Module Selection Framework for Video Captioning
por: Lee, Sangho, et al.
Publicado: (2025)
por: Lee, Sangho, et al.
Publicado: (2025)
Exploiting Auxiliary Caption for Video Grounding
por: Li, Hongxiang, et al.
Publicado: (2023)
por: Li, Hongxiang, et al.
Publicado: (2023)
StereoDiff: Stereo-Diffusion Synergy for Video Depth Estimation
por: Li, Haodong, et al.
Publicado: (2025)
por: Li, Haodong, et al.
Publicado: (2025)
Attention Based Encoder Decoder Model for Video Captioning in Nepali (2023)
por: Parajuli, Kabita, et al.
Publicado: (2023)
por: Parajuli, Kabita, et al.
Publicado: (2023)
OBS-Diff: Accurate Pruning For Diffusion Models in One-Shot
por: Zhu, Junhan, et al.
Publicado: (2025)
por: Zhu, Junhan, et al.
Publicado: (2025)
DualDiff+: Dual-Branch Diffusion for High-Fidelity Video Generation with Reward Guidance
por: Yang, Zhao, et al.
Publicado: (2025)
por: Yang, Zhao, et al.
Publicado: (2025)
IF-VidCap: Can Video Caption Models Follow Instructions?
por: Li, Shihao, et al.
Publicado: (2025)
por: Li, Shihao, et al.
Publicado: (2025)
eMotions: A Large-Scale Dataset and Audio-Visual Fusion Network for Emotion Analysis in Short-form Videos
por: Wu, Xuecheng, et al.
Publicado: (2025)
por: Wu, Xuecheng, et al.
Publicado: (2025)
Ejemplares similares
-
EPIR: An Efficient Patch Tokenization, Integration and Representation Framework for Micro-expression Recognition
por: Wang, Junbo, et al.
Publicado: (2026) -
DiffVC-OSD: One-Step Diffusion-based Perceptual Neural Video Compression Framework
por: Ma, Wenzhuo, et al.
Publicado: (2025) -
DiffVC-RT: Towards Practical Real-Time Diffusion-based Perceptual Neural Video Compression
por: Ma, Wenzhuo, et al.
Publicado: (2026) -
TACR-YOLO: A Real-time Detection Framework for Abnormal Human Behaviors Enhanced with Coordinate and Task-Aware Representations
por: Yin, Xinyi, et al.
Publicado: (2025) -
FAMNet: Integrating 2D and 3D Features for Micro-expression Recognition via Multi-task Learning and Hierarchical Attention
por: Fu, Liangyu, et al.
Publicado: (2025)