Towards Fine-Grained Human Motion Video Captioning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Song, Guorui, Wang, Guocun, Huang, Zhe, Lin, Jing, Zhe, Xuefei, Li, Jian, Wang, Haoqian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Taming Hallucinations: Boosting MLLMs' Video Understanding via Counterfactual Video Generation
von: Huang, Zhe, et al.
Veröffentlicht: (2025)
von: Huang, Zhe, et al.
Veröffentlicht: (2025)
F$^3$Set: Towards Analyzing Fast, Frequent, and Fine-grained Events from Videos
von: Liu, Zhaoyu, et al.
Veröffentlicht: (2025)
von: Liu, Zhaoyu, et al.
Veröffentlicht: (2025)
DreamRunner: Fine-Grained Compositional Story-to-Video Generation with Retrieval-Augmented Motion Adaptation
von: Wang, Zun, et al.
Veröffentlicht: (2024)
von: Wang, Zun, et al.
Veröffentlicht: (2024)
Sam-Guided Enhanced Fine-Grained Encoding with Mixed Semantic Learning for Medical Image Captioning
von: Zhang, Zhenyu, et al.
Veröffentlicht: (2023)
von: Zhang, Zhenyu, et al.
Veröffentlicht: (2023)
FAVOR-Bench: A Comprehensive Benchmark for Fine-Grained Video Motion Understanding
von: Tu, Chongjun, et al.
Veröffentlicht: (2025)
von: Tu, Chongjun, et al.
Veröffentlicht: (2025)
Top-Down Semantic Refinement for Image Captioning
von: Zhang, Jusheng, et al.
Veröffentlicht: (2025)
von: Zhang, Jusheng, et al.
Veröffentlicht: (2025)
PhyMotion: Structured 3D Motion Reward for Physics-Grounded Human Video Generation
von: Huang, Yidong, et al.
Veröffentlicht: (2026)
von: Huang, Yidong, et al.
Veröffentlicht: (2026)
Towards Fine-Grained Video Question Answering
von: Dai, Wei, et al.
Veröffentlicht: (2025)
von: Dai, Wei, et al.
Veröffentlicht: (2025)
VERIFIED: A Video Corpus Moment Retrieval Benchmark for Fine-Grained Video Understanding
von: Chen, Houlun, et al.
Veröffentlicht: (2024)
von: Chen, Houlun, et al.
Veröffentlicht: (2024)
ReasonMap: Towards Fine-Grained Visual Reasoning from Transit Maps
von: Feng, Sicheng, et al.
Veröffentlicht: (2025)
von: Feng, Sicheng, et al.
Veröffentlicht: (2025)
VC4VG: Optimizing Video Captions for Text-to-Video Generation
von: Du, Yang, et al.
Veröffentlicht: (2025)
von: Du, Yang, et al.
Veröffentlicht: (2025)
FastV-RAG: Towards Fast and Fine-Grained Video QA with Retrieval-Augmented Generation
von: Li, Gen, et al.
Veröffentlicht: (2026)
von: Li, Gen, et al.
Veröffentlicht: (2026)
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans
von: Qiu, Yansheng, et al.
Veröffentlicht: (2025)
von: Qiu, Yansheng, et al.
Veröffentlicht: (2025)
Any2Caption:Interpreting Any Condition to Caption for Controllable Video Generation
von: Wu, Shengqiong, et al.
Veröffentlicht: (2025)
von: Wu, Shengqiong, et al.
Veröffentlicht: (2025)
Towards Multimodal Video Paragraph Captioning Models Robust to Missing Modality
von: Chen, Sishuo, et al.
Veröffentlicht: (2024)
von: Chen, Sishuo, et al.
Veröffentlicht: (2024)
SAM Guided Semantic and Motion Changed Region Mining for Remote Sensing Change Captioning
von: Wang, Futian, et al.
Veröffentlicht: (2025)
von: Wang, Futian, et al.
Veröffentlicht: (2025)
Video Summarization: Towards Entity-Aware Captions
von: Ayyubi, Hammad A., et al.
Veröffentlicht: (2023)
von: Ayyubi, Hammad A., et al.
Veröffentlicht: (2023)
LD-RPS: Zero-Shot Unified Image Restoration via Latent Diffusion Recurrent Posterior Sampling
von: Li, Huaqiu, et al.
Veröffentlicht: (2025)
von: Li, Huaqiu, et al.
Veröffentlicht: (2025)
Mining Fine-Grained Image-Text Alignment for Zero-Shot Captioning via Text-Only Training
von: Qiu, Longtian, et al.
Veröffentlicht: (2024)
von: Qiu, Longtian, et al.
Veröffentlicht: (2024)
Enhancing Self-Supervised Fine-Grained Video Object Tracking with Dynamic Memory Prediction
von: Zhou, Zihan, et al.
Veröffentlicht: (2025)
von: Zhou, Zihan, et al.
Veröffentlicht: (2025)
CoDiff: Conditional Diffusion Model for Collaborative 3D Object Detection
von: Huang, Zhe, et al.
Veröffentlicht: (2025)
von: Huang, Zhe, et al.
Veröffentlicht: (2025)
FILA: Fine-Grained Vision Language Models
von: Zhu, Shiding, et al.
Veröffentlicht: (2024)
von: Zhu, Shiding, et al.
Veröffentlicht: (2024)
Fake-in-Facext: Towards Fine-Grained Explainable DeepFake Analysis
von: Qin, Lixiong, et al.
Veröffentlicht: (2025)
von: Qin, Lixiong, et al.
Veröffentlicht: (2025)
Unlocking Vision-Language Models for Video Anomaly Detection via Fine-Grained Prompting
von: Zou, Shu, et al.
Veröffentlicht: (2025)
von: Zou, Shu, et al.
Veröffentlicht: (2025)
MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance
von: Zhang, Yuang, et al.
Veröffentlicht: (2024)
von: Zhang, Yuang, et al.
Veröffentlicht: (2024)
Image Synthesis under Limited Data: A Survey and Taxonomy
von: Yang, Mengping, et al.
Veröffentlicht: (2023)
von: Yang, Mengping, et al.
Veröffentlicht: (2023)
TexVocab: Texture Vocabulary-conditioned Human Avatars
von: Liu, Yuxiao, et al.
Veröffentlicht: (2024)
von: Liu, Yuxiao, et al.
Veröffentlicht: (2024)
QAPruner: Quantization-Aware Vision Token Pruning for Multimodal Large Language Models
von: Wang, Xinhao, et al.
Veröffentlicht: (2026)
von: Wang, Xinhao, et al.
Veröffentlicht: (2026)
Retrieval-Augmented Long-Context Translation for Cultural Image Captioning: Gators submission for AmericasNLP 2026 shared task
von: Dhawan, Aashish, et al.
Veröffentlicht: (2026)
von: Dhawan, Aashish, et al.
Veröffentlicht: (2026)
PixelSmile: Toward Fine-Grained Facial Expression Editing
von: Hua, Jiabin, et al.
Veröffentlicht: (2026)
von: Hua, Jiabin, et al.
Veröffentlicht: (2026)
Parrot Captions Teach CLIP to Spot Text
von: Lin, Yiqi, et al.
Veröffentlicht: (2023)
von: Lin, Yiqi, et al.
Veröffentlicht: (2023)
NOVA3D: Normal Aligned Video Diffusion Model for Single Image to 3D Generation
von: Yang, Yuxiao, et al.
Veröffentlicht: (2025)
von: Yang, Yuxiao, et al.
Veröffentlicht: (2025)
Accurate and Fast Compressed Video Captioning
von: Shen, Yaojie, et al.
Veröffentlicht: (2023)
von: Shen, Yaojie, et al.
Veröffentlicht: (2023)
FireSentry: A Multi-Modal Spatio-temporal Benchmark Dataset for Fine-Grained Wildfire Spread Forecasting
von: Zhou, Nan, et al.
Veröffentlicht: (2025)
von: Zhou, Nan, et al.
Veröffentlicht: (2025)
QCaption: Video Captioning and Q&A through Fusion of Large Multimodal Models
von: Wang, Jiale, et al.
Veröffentlicht: (2026)
von: Wang, Jiale, et al.
Veröffentlicht: (2026)
VideoScore: Building Automatic Metrics to Simulate Fine-grained Human Feedback for Video Generation
von: He, Xuan, et al.
Veröffentlicht: (2024)
von: He, Xuan, et al.
Veröffentlicht: (2024)
CoLLM-NAS: Collaborative Large Language Models for Efficient Knowledge-Guided Neural Architecture Search
von: Li, Zhe, et al.
Veröffentlicht: (2025)
von: Li, Zhe, et al.
Veröffentlicht: (2025)
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption
von: Fan, Tiehan, et al.
Veröffentlicht: (2024)
von: Fan, Tiehan, et al.
Veröffentlicht: (2024)
TokenMotion: Decoupled Motion Control via Token Disentanglement for Human-centric Video Generation
von: Li, Ruineng, et al.
Veröffentlicht: (2025)
von: Li, Ruineng, et al.
Veröffentlicht: (2025)
MiSCHiEF: A Benchmark in Minimal-Pairs of Safety and Culture for Holistic Evaluation of Fine-Grained Image-Caption Alignment
von: Banerjee, Sagarika, et al.
Veröffentlicht: (2026)
von: Banerjee, Sagarika, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Taming Hallucinations: Boosting MLLMs' Video Understanding via Counterfactual Video Generation
von: Huang, Zhe, et al.
Veröffentlicht: (2025) -
F$^3$Set: Towards Analyzing Fast, Frequent, and Fine-grained Events from Videos
von: Liu, Zhaoyu, et al.
Veröffentlicht: (2025) -
DreamRunner: Fine-Grained Compositional Story-to-Video Generation with Retrieval-Augmented Motion Adaptation
von: Wang, Zun, et al.
Veröffentlicht: (2024) -
Sam-Guided Enhanced Fine-Grained Encoding with Mixed Semantic Learning for Medical Image Captioning
von: Zhang, Zhenyu, et al.
Veröffentlicht: (2023) -
FAVOR-Bench: A Comprehensive Benchmark for Fine-Grained Video Motion Understanding
von: Tu, Chongjun, et al.
Veröffentlicht: (2025)