Chapter-Llama: Efficient Chaptering in Hour-Long Videos with LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Ventura, Lucas, Yang, Antoine, Schmid, Cordelia, Varol, Gül |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CoVR-2: Automatic Data Construction for Composed Video Retrieval
by: Ventura, Lucas, et al.
Published: (2023)
by: Ventura, Lucas, et al.
Published: (2023)
Learning text-to-video retrieval from image captioning
by: Ventura, Lucas, et al.
Published: (2024)
by: Ventura, Lucas, et al.
Published: (2024)
Grounded Video Caption Generation
by: Kazakos, Evangelos, et al.
Published: (2024)
by: Kazakos, Evangelos, et al.
Published: (2024)
Large-scale Pre-training for Grounded Video Caption Generation
by: Kazakos, Evangelos, et al.
Published: (2025)
by: Kazakos, Evangelos, et al.
Published: (2025)
ARC-Chapter: Structuring Hour-Long Videos into Navigable Chapters and Hierarchical Summaries
by: Pu, Junfu, et al.
Published: (2025)
by: Pu, Junfu, et al.
Published: (2025)
VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation
by: Bousselham, Walid, et al.
Published: (2025)
by: Bousselham, Walid, et al.
Published: (2025)
CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects
by: Fiastre, Gabriel, et al.
Published: (2025)
by: Fiastre, Gabriel, et al.
Published: (2025)
RAVEN: Rethinking Adversarial Video Generation with Efficient Tri-plane Networks
by: Ghosh, Partha, et al.
Published: (2024)
by: Ghosh, Partha, et al.
Published: (2024)
InterPose: Learning to Generate Human-Object Interactions from Large-Scale Web Videos
by: Zhang, Yangsong, et al.
Published: (2025)
by: Zhang, Yangsong, et al.
Published: (2025)
A Cross-Dataset Study for Text-based 3D Human Motion Retrieval
by: Bensabath, Léore, et al.
Published: (2024)
by: Bensabath, Léore, et al.
Published: (2024)
Dense Video Object Captioning from Disjoint Supervision
by: Zhou, Xingyi, et al.
Published: (2023)
by: Zhou, Xingyi, et al.
Published: (2023)
OVFact: Measuring and Improving Open-Vocabulary Factuality for Long Caption Models
by: Wysoczańska, Monika, et al.
Published: (2025)
by: Wysoczańska, Monika, et al.
Published: (2025)
BrickNet: Graph-Backed Generative Brick Assembly
by: Kulits, Peter, et al.
Published: (2026)
by: Kulits, Peter, et al.
Published: (2026)
ComposeAnything: Composite Object Priors for Text-to-Image Generation
by: Khan, Zeeshan, et al.
Published: (2025)
by: Khan, Zeeshan, et al.
Published: (2025)
MINERVA-Cultural: A Benchmark for Cultural and Multilingual Long Video Reasoning
by: Singh, Darshan, et al.
Published: (2026)
by: Singh, Darshan, et al.
Published: (2026)
What Are You Doing? A Closer Look at Controllable Human Video Generation
by: Bugliarello, Emanuele, et al.
Published: (2025)
by: Bugliarello, Emanuele, et al.
Published: (2025)
SINC: Spatial Composition of 3D Human Motions for Simultaneous Action Generation
by: Athanasiou, Nikos, et al.
Published: (2023)
by: Athanasiou, Nikos, et al.
Published: (2023)
Dense Optical Tracking: Connecting the Dots
by: Moing, Guillaume Le, et al.
Published: (2023)
by: Moing, Guillaume Le, et al.
Published: (2023)
PointACT: Vision-Language-Action Models with Multi-Scale Point-Action Interaction
by: Chen, Shizhe, et al.
Published: (2026)
by: Chen, Shizhe, et al.
Published: (2026)
Towards Generalizable Vision-Language Robotic Manipulation: A Benchmark and LLM-guided 3D Policy
by: Garcia, Ricardo, et al.
Published: (2024)
by: Garcia, Ricardo, et al.
Published: (2024)
Time-, Memory- and Parameter-Efficient Visual Adaptation
by: Mercea, Otniel-Bogdan, et al.
Published: (2024)
by: Mercea, Otniel-Bogdan, et al.
Published: (2024)
Lost in Translation, Found in Context: Sign Language Translation with Contextual Cues
by: Jang, Youngjoon, et al.
Published: (2025)
by: Jang, Youngjoon, et al.
Published: (2025)
Web-Scale Visual Entity Recognition: An LLM-Driven Data Approach
by: Caron, Mathilde, et al.
Published: (2024)
by: Caron, Mathilde, et al.
Published: (2024)
SUGAR: Pre-training 3D Visual Representations for Robotics
by: Chen, Shizhe, et al.
Published: (2024)
by: Chen, Shizhe, et al.
Published: (2024)
Retrieval-Enhanced Contrastive Vision-Text Models
by: Iscen, Ahmet, et al.
Published: (2023)
by: Iscen, Ahmet, et al.
Published: (2023)
A Generative Approach for Wikipedia-Scale Visual Entity Recognition
by: Caron, Mathilde, et al.
Published: (2024)
by: Caron, Mathilde, et al.
Published: (2024)
Streaming Dense Video Captioning
by: Zhou, Xingyi, et al.
Published: (2024)
by: Zhou, Xingyi, et al.
Published: (2024)
HORT: Monocular Hand-held Objects Reconstruction with Transformers
by: Chen, Zerui, et al.
Published: (2025)
by: Chen, Zerui, et al.
Published: (2025)
LoFT: LoRA-fused Training Dataset Generation with Few-shot Guidance
by: Kim, Jae Myung, et al.
Published: (2025)
by: Kim, Jae Myung, et al.
Published: (2025)
Learning Correlation Structures for Vision Transformers
by: Kim, Manjin, et al.
Published: (2024)
by: Kim, Manjin, et al.
Published: (2024)
CAViAR: Critic-Augmented Video Agentic Reasoning
by: Menon, Sachit, et al.
Published: (2025)
by: Menon, Sachit, et al.
Published: (2025)
Scaling Cross-Environment Failure Reasoning Data for Vision-Language Robotic Manipulation
by: Pacaud, Paul, et al.
Published: (2025)
by: Pacaud, Paul, et al.
Published: (2025)
Online 3D Scene Reconstruction Using Neural Object Priors
by: Chabal, Thomas, et al.
Published: (2025)
by: Chabal, Thomas, et al.
Published: (2025)
FOM-Nav: Frontier-Object Maps for Object Goal Navigation
by: Chabal, Thomas, et al.
Published: (2025)
by: Chabal, Thomas, et al.
Published: (2025)
VoCap: Video Object Captioning and Segmentation from Any Prompt
by: Uijlings, Jasper, et al.
Published: (2025)
by: Uijlings, Jasper, et al.
Published: (2025)
STORM: Token-Efficient Long Video Understanding for Multimodal LLMs
by: Jiang, Jindong, et al.
Published: (2025)
by: Jiang, Jindong, et al.
Published: (2025)
AutoAD III: The Prequel -- Back to the Pixels
by: Han, Tengda, et al.
Published: (2024)
by: Han, Tengda, et al.
Published: (2024)
ViViDex: Learning Vision-based Dexterous Manipulation from Human Videos
by: Chen, Zerui, et al.
Published: (2024)
by: Chen, Zerui, et al.
Published: (2024)
MotionFix: Text-Driven 3D Human Motion Editing
by: Athanasiou, Nikos, et al.
Published: (2024)
by: Athanasiou, Nikos, et al.
Published: (2024)
Lost in Translation, Found in Embeddings: Sign Language Translation and Alignment
by: Jang, Youngjoon, et al.
Published: (2025)
by: Jang, Youngjoon, et al.
Published: (2025)
Similar Items
-
CoVR-2: Automatic Data Construction for Composed Video Retrieval
by: Ventura, Lucas, et al.
Published: (2023) -
Learning text-to-video retrieval from image captioning
by: Ventura, Lucas, et al.
Published: (2024) -
Grounded Video Caption Generation
by: Kazakos, Evangelos, et al.
Published: (2024) -
Large-scale Pre-training for Grounded Video Caption Generation
by: Kazakos, Evangelos, et al.
Published: (2025) -
ARC-Chapter: Structuring Hour-Long Videos into Navigable Chapters and Hierarchical Summaries
by: Pu, Junfu, et al.
Published: (2025)