Fine-Grained Captioning of Long Videos through Scene Graph Consolidation
Fuente:
arXiv
Guardado en:
| Autores principales: | Chu, Sanghyeok, Seo, Seonguk, Han, Bohyung |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Learning to Translate Noise for Robust Image Denoising
por: Ha, Inju, et al.
Publicado: (2024)
por: Ha, Inju, et al.
Publicado: (2024)
Beyond the Ground Truth: Enhanced Supervision for Image Restoration
por: Ryou, Donghun, et al.
Publicado: (2025)
por: Ryou, Donghun, et al.
Publicado: (2025)
Scene Graph-guided SegCaptioning Transformer with Fine-grained Alignment for Controllable Video Segmentation and Captioning
por: Zhang, Xu, et al.
Publicado: (2026)
por: Zhang, Xu, et al.
Publicado: (2026)
Metric Compatible Training for Online Backfilling in Large-Scale Retrieval
por: Seo, Seonguk, et al.
Publicado: (2023)
por: Seo, Seonguk, et al.
Publicado: (2023)
Enhancing Mixture-of-Experts Specialization via Cluster-Aware Upcycling
por: Chu, Sanghyeok, et al.
Publicado: (2026)
por: Chu, Sanghyeok, et al.
Publicado: (2026)
Frame-Level Captions for Long Video Generation with Complex Multi Scenes
por: Zheng, Guangcong, et al.
Publicado: (2025)
por: Zheng, Guangcong, et al.
Publicado: (2025)
Towards Fine-Grained Human Motion Video Captioning
por: Song, Guorui, et al.
Publicado: (2025)
por: Song, Guorui, et al.
Publicado: (2025)
SG2VID: Scene Graphs Enable Fine-Grained Control for Video Synthesis
por: Sivakumar, Ssharvien Kumar, et al.
Publicado: (2025)
por: Sivakumar, Ssharvien Kumar, et al.
Publicado: (2025)
Map the Flow: Revealing Hidden Pathways of Information in VideoLLMs
por: Kim, Minji, et al.
Publicado: (2025)
por: Kim, Minji, et al.
Publicado: (2025)
Leveraging Temporal Contextualization for Video Action Recognition
por: Kim, Minji, et al.
Publicado: (2024)
por: Kim, Minji, et al.
Publicado: (2024)
Diffusion-Based Conditional Image Editing through Optimized Inference with Guidance
por: Lee, Hyunsoo, et al.
Publicado: (2024)
por: Lee, Hyunsoo, et al.
Publicado: (2024)
GP-4DGS: Probabilistic 4D Gaussian Splatting from Monocular Video via Variational Gaussian Processes
por: Kim, Mijeong, et al.
Publicado: (2026)
por: Kim, Mijeong, et al.
Publicado: (2026)
Addressing the ID-Matching Challenge in Long Video Captioning
por: Yang, Zhantao, et al.
Publicado: (2025)
por: Yang, Zhantao, et al.
Publicado: (2025)
LongCaptioning: Unlocking the Power of Long Video Caption Generation in Large Multimodal Models
por: Wei, Hongchen, et al.
Publicado: (2025)
por: Wei, Hongchen, et al.
Publicado: (2025)
Memory Consolidation Enables Long-Context Video Understanding
por: Balažević, Ivana, et al.
Publicado: (2024)
por: Balažević, Ivana, et al.
Publicado: (2024)
VITRIX-CLIPIN: Enhancing Fine-Grained Visual Understanding in CLIP via Instruction Editing Data and Long Captions
por: Wang, Ziteng, et al.
Publicado: (2025)
por: Wang, Ziteng, et al.
Publicado: (2025)
STR-Match: Matching SpatioTemporal Relevance Score for Training-Free Video Editing
por: Lee, Junsung, et al.
Publicado: (2025)
por: Lee, Junsung, et al.
Publicado: (2025)
TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions
por: Yao, Linli, et al.
Publicado: (2026)
por: Yao, Linli, et al.
Publicado: (2026)
Video ReCap: Recursive Captioning of Hour-Long Videos
por: Islam, Md Mohaiminul, et al.
Publicado: (2024)
por: Islam, Md Mohaiminul, et al.
Publicado: (2024)
Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting
por: Tang, Yunlong, et al.
Publicado: (2025)
por: Tang, Yunlong, et al.
Publicado: (2025)
vid-TLDR: Training Free Token merging for Light-weight Video Transformer
por: Choi, Joonmyung, et al.
Publicado: (2024)
por: Choi, Joonmyung, et al.
Publicado: (2024)
FiLA-Video: Spatio-Temporal Compression for Fine-Grained Long Video Understanding
por: Guo, Yanan, et al.
Publicado: (2025)
por: Guo, Yanan, et al.
Publicado: (2025)
GEM-VPC: A dual Graph-Enhanced Multimodal integration for Video Paragraph Captioning
por: Wang, Eileen, et al.
Publicado: (2024)
por: Wang, Eileen, et al.
Publicado: (2024)
No Detail Left Behind: Revisiting Self-Retrieval for Fine-Grained Image Captioning
por: Gaur, Manu, et al.
Publicado: (2024)
por: Gaur, Manu, et al.
Publicado: (2024)
Diffusion-Based Image-to-Image Translation by Noise Correction via Prompt Interpolation
por: Lee, Junsung, et al.
Publicado: (2024)
por: Lee, Junsung, et al.
Publicado: (2024)
SGDiff: Scene Graph Guided Diffusion Model for Image Collaborative SegCaptioning
por: Zhang, Xu, et al.
Publicado: (2025)
por: Zhang, Xu, et al.
Publicado: (2025)
FIFO-Diffusion: Generating Infinite Videos from Text without Training
por: Kim, Jihwan, et al.
Publicado: (2024)
por: Kim, Jihwan, et al.
Publicado: (2024)
CaReBench: A Fine-Grained Benchmark for Video Captioning and Retrieval
por: Xu, Yifan, et al.
Publicado: (2024)
por: Xu, Yifan, et al.
Publicado: (2024)
The Devil is in the Distributions: Explicit Modeling of Scene Content is Key in Zero-Shot Video Captioning
por: Tian, Mingkai, et al.
Publicado: (2025)
por: Tian, Mingkai, et al.
Publicado: (2025)
OmniSparse: Training-Aware Fine-Grained Sparse Attention for Long-Video MLLMs
por: Chen, Feng, et al.
Publicado: (2025)
por: Chen, Feng, et al.
Publicado: (2025)
3D Scene Prompting for Scene-Consistent Camera-Controllable Video Generation
por: Lee, JoungBin, et al.
Publicado: (2025)
por: Lee, JoungBin, et al.
Publicado: (2025)
Fine-Grained Scene Image Classification with Modality-Agnostic Adapter
por: Wang, Yiqun, et al.
Publicado: (2024)
por: Wang, Yiqun, et al.
Publicado: (2024)
COCONut-PanCap: Joint Panoptic Segmentation and Grounded Captions for Fine-Grained Understanding and Generation
por: Deng, Xueqing, et al.
Publicado: (2025)
por: Deng, Xueqing, et al.
Publicado: (2025)
Fine-Grained Scene Graph Generation via Sample-Level Bias Prediction
por: Li, Yansheng, et al.
Publicado: (2024)
por: Li, Yansheng, et al.
Publicado: (2024)
Live Video Captioning
por: Blanco-Fernández, Eduardo, et al.
Publicado: (2024)
por: Blanco-Fernández, Eduardo, et al.
Publicado: (2024)
Implicit and Explicit Commonsense for Multi-sentence Video Captioning
por: Chou, Shih-Han, et al.
Publicado: (2023)
por: Chou, Shih-Han, et al.
Publicado: (2023)
MJ-VIDEO: Fine-Grained Benchmarking and Rewarding Video Preferences in Video Generation
por: Tong, Haibo, et al.
Publicado: (2025)
por: Tong, Haibo, et al.
Publicado: (2025)
Benchmarking Large Vision-Language Models via Directed Scene Graph for Comprehensive Image Captioning
por: Lu, Fan, et al.
Publicado: (2024)
por: Lu, Fan, et al.
Publicado: (2024)
UGC-VideoCaptioner: An Omni UGC Video Detail Caption Model and New Benchmarks
por: Wu, Peiran, et al.
Publicado: (2025)
por: Wu, Peiran, et al.
Publicado: (2025)
DreamScene4D: Dynamic Multi-Object Scene Generation from Monocular Videos
por: Chu, Wen-Hsuan, et al.
Publicado: (2024)
por: Chu, Wen-Hsuan, et al.
Publicado: (2024)
Ejemplares similares
-
Learning to Translate Noise for Robust Image Denoising
por: Ha, Inju, et al.
Publicado: (2024) -
Beyond the Ground Truth: Enhanced Supervision for Image Restoration
por: Ryou, Donghun, et al.
Publicado: (2025) -
Scene Graph-guided SegCaptioning Transformer with Fine-grained Alignment for Controllable Video Segmentation and Captioning
por: Zhang, Xu, et al.
Publicado: (2026) -
Metric Compatible Training for Online Backfilling in Large-Scale Retrieval
por: Seo, Seonguk, et al.
Publicado: (2023) -
Enhancing Mixture-of-Experts Specialization via Cluster-Aware Upcycling
por: Chu, Sanghyeok, et al.
Publicado: (2026)