ViDiC: Video Difference Captioning
Fuente:
arXiv
Guardado en:
| Autores principales: | Wu, Jiangtao, Li, Shihao, Bian, Zhaozhou, Chen, Jialu, Wen, Runzhe, Ping, An, He, Yiwen, Wang, Jiakai, Zhang, Yuanxing, Liu, Jiaheng |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
IF-VidCap: Can Video Caption Models Follow Instructions?
por: Li, Shihao, et al.
Publicado: (2025)
por: Li, Shihao, et al.
Publicado: (2025)
DiViD: Disentangled Video Diffusion for Static-Dynamic Factorization
por: Gheisari, Marzieh, et al.
Publicado: (2025)
por: Gheisari, Marzieh, et al.
Publicado: (2025)
InsViE-1M: Effective Instruction-based Video Editing with Elaborate Dataset Construction
por: Wu, Yuhui, et al.
Publicado: (2025)
por: Wu, Yuhui, et al.
Publicado: (2025)
TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions
por: Yao, Linli, et al.
Publicado: (2026)
por: Yao, Linli, et al.
Publicado: (2026)
AVoCaDO: An Audiovisual Video Captioner Driven by Temporal Orchestration
por: Chen, Xinlong, et al.
Publicado: (2025)
por: Chen, Xinlong, et al.
Publicado: (2025)
LongCaptioning: Unlocking the Power of Long Video Caption Generation in Large Multimodal Models
por: Wei, Hongchen, et al.
Publicado: (2025)
por: Wei, Hongchen, et al.
Publicado: (2025)
ViCaS: A Dataset for Combining Holistic and Pixel-level Video Understanding using Captions with Grounded Segmentation
por: Athar, Ali, et al.
Publicado: (2024)
por: Athar, Ali, et al.
Publicado: (2024)
T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation
por: Cao, Zhe, et al.
Publicado: (2025)
por: Cao, Zhe, et al.
Publicado: (2025)
ViDiT-Q: Efficient and Accurate Quantization of Diffusion Transformers for Image and Video Generation
por: Zhao, Tianchen, et al.
Publicado: (2024)
por: Zhao, Tianchen, et al.
Publicado: (2024)
ViTOC: Vision Transformer and Object-aware Captioner
por: Huang, Feiyang
Publicado: (2024)
por: Huang, Feiyang
Publicado: (2024)
MT-Video-Bench: A Holistic Video Understanding Benchmark for Evaluating Multimodal LLMs in Multi-Turn Dialogues
por: Pan, Yaning, et al.
Publicado: (2025)
por: Pan, Yaning, et al.
Publicado: (2025)
ViSIL: Unified Evaluation of Information Loss in Multimodal Video Captioning
por: Li, Po-han, et al.
Publicado: (2026)
por: Li, Po-han, et al.
Publicado: (2026)
UGC-VideoCaptioner: An Omni UGC Video Detail Caption Model and New Benchmarks
por: Wu, Peiran, et al.
Publicado: (2025)
por: Wu, Peiran, et al.
Publicado: (2025)
SGCap: Decoding Semantic Group for Zero-shot Video Captioning
por: Pan, Zeyu, et al.
Publicado: (2025)
por: Pan, Zeyu, et al.
Publicado: (2025)
Geo-Align: Video Generation Alignment via Metric Geometry Reward
por: Li, Zizun, et al.
Publicado: (2026)
por: Li, Zizun, et al.
Publicado: (2026)
ViSTec: Video Modeling for Sports Technique Recognition and Tactical Analysis
por: He, Yuchen, et al.
Publicado: (2024)
por: He, Yuchen, et al.
Publicado: (2024)
Pseudo-labeling with Keyword Refining for Few-Supervised Video Captioning
por: Li, Ping, et al.
Publicado: (2024)
por: Li, Ping, et al.
Publicado: (2024)
LoViC: Efficient Long Video Generation with Context Compression
por: Jiang, Jiaxiu, et al.
Publicado: (2025)
por: Jiang, Jiaxiu, et al.
Publicado: (2025)
ViBe: Ultra-High-Resolution Video Synthesis Born from Pure Images
por: Wu, Yunfeng, et al.
Publicado: (2026)
por: Wu, Yunfeng, et al.
Publicado: (2026)
TiViBench: Benchmarking Think-in-Video Reasoning for Video Generative Models
por: Chen, Harold Haodong, et al.
Publicado: (2025)
por: Chen, Harold Haodong, et al.
Publicado: (2025)
Efficient Bayer-Domain Video Computer Vision with Fast Motion Estimation and Learned Perception Residual
por: Wang, Haichao, et al.
Publicado: (2025)
por: Wang, Haichao, et al.
Publicado: (2025)
Live Video Captioning
por: Blanco-Fernández, Eduardo, et al.
Publicado: (2024)
por: Blanco-Fernández, Eduardo, et al.
Publicado: (2024)
UltraVideo: High-Quality UHD Video Dataset with Comprehensive Captions
por: Xue, Zhucun, et al.
Publicado: (2025)
por: Xue, Zhucun, et al.
Publicado: (2025)
ViViD: Video Virtual Try-on using Diffusion Models
por: Fang, Zixun, et al.
Publicado: (2024)
por: Fang, Zixun, et al.
Publicado: (2024)
Retrieval-Augmented Egocentric Video Captioning
por: Xu, Jilan, et al.
Publicado: (2024)
por: Xu, Jilan, et al.
Publicado: (2024)
Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting
por: Tang, Yunlong, et al.
Publicado: (2025)
por: Tang, Yunlong, et al.
Publicado: (2025)
VR-Thinker: Boosting Video Reward Models through Thinking-with-Image Reasoning
por: Wang, Qunzhong, et al.
Publicado: (2025)
por: Wang, Qunzhong, et al.
Publicado: (2025)
Grounded Video Caption Generation
por: Kazakos, Evangelos, et al.
Publicado: (2024)
por: Kazakos, Evangelos, et al.
Publicado: (2024)
Streaming Dense Video Captioning
por: Zhou, Xingyi, et al.
Publicado: (2024)
por: Zhou, Xingyi, et al.
Publicado: (2024)
Technical Report for Soccernet 2023 -- Dense Video Captioning
por: Ruan, Zheng, et al.
Publicado: (2024)
por: Ruan, Zheng, et al.
Publicado: (2024)
Accurate and Fast Compressed Video Captioning
por: Shen, Yaojie, et al.
Publicado: (2023)
por: Shen, Yaojie, et al.
Publicado: (2023)
DiTPainter: Efficient Video Inpainting with Diffusion Transformers
por: Wu, Xian, et al.
Publicado: (2025)
por: Wu, Xian, et al.
Publicado: (2025)
Global2Local: A Joint-Hierarchical Attention for Video Captioning
por: Dai, Chengpeng, et al.
Publicado: (2022)
por: Dai, Chengpeng, et al.
Publicado: (2022)
SoccerNet-Caption: Dense Video Captioning for Soccer Broadcasts Commentaries
por: Mkhallati, Hassan, et al.
Publicado: (2023)
por: Mkhallati, Hassan, et al.
Publicado: (2023)
Sali4Vid: Saliency-Aware Video Reweighting and Adaptive Caption Retrieval for Dense Video Captioning
por: Jeon, MinJu, et al.
Publicado: (2025)
por: Jeon, MinJu, et al.
Publicado: (2025)
Dense Video Object Captioning from Disjoint Supervision
por: Zhou, Xingyi, et al.
Publicado: (2023)
por: Zhou, Xingyi, et al.
Publicado: (2023)
ViMix-14M: A Curated Multi-Source Video-Text Dataset with Long-Form, High-Quality Captions and Crawl-Free Access
por: Yang, Timing, et al.
Publicado: (2025)
por: Yang, Timing, et al.
Publicado: (2025)
DiVE: DiT-based Video Generation with Enhanced Control
por: Jiang, Junpeng, et al.
Publicado: (2024)
por: Jiang, Junpeng, et al.
Publicado: (2024)
OwlCap: Harmonizing Motion-Detail for Video Captioning via HMD-270K and Caption Set Equivalence Reward
por: Zhong, Chunlin, et al.
Publicado: (2025)
por: Zhong, Chunlin, et al.
Publicado: (2025)
Video Summarization: Towards Entity-Aware Captions
por: Ayyubi, Hammad A., et al.
Publicado: (2023)
por: Ayyubi, Hammad A., et al.
Publicado: (2023)
Ejemplares similares
-
IF-VidCap: Can Video Caption Models Follow Instructions?
por: Li, Shihao, et al.
Publicado: (2025) -
DiViD: Disentangled Video Diffusion for Static-Dynamic Factorization
por: Gheisari, Marzieh, et al.
Publicado: (2025) -
InsViE-1M: Effective Instruction-based Video Editing with Elaborate Dataset Construction
por: Wu, Yuhui, et al.
Publicado: (2025) -
TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions
por: Yao, Linli, et al.
Publicado: (2026) -
AVoCaDO: An Audiovisual Video Captioner Driven by Temporal Orchestration
por: Chen, Xinlong, et al.
Publicado: (2025)