DenseStep2M: A Scalable, Training-Free Pipeline for Dense Instructional Video Annotation
Fuente:
arXiv
Saved in:
| Main Authors: | Ge, Mingji, Chen, Qirui, Li, Zeqian, Xie, Weidi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multi-Sentence Grounding for Long-term Instructional Video
by: Li, Zeqian, et al.
Published: (2023)
by: Li, Zeqian, et al.
Published: (2023)
Learning Streaming Video Representation via Multitask Training
by: Yan, Yibin, et al.
Published: (2025)
by: Yan, Yibin, et al.
Published: (2025)
Grounded Multi-Hop VideoQA in Long-Form Egocentric Videos
by: Chen, Qirui, et al.
Published: (2024)
by: Chen, Qirui, et al.
Published: (2024)
Enhancing Video-LLM Reasoning via Agent-of-Thoughts Distillation
by: Shi, Yudi, et al.
Published: (2024)
by: Shi, Yudi, et al.
Published: (2024)
Universal Video Temporal Grounding with Generative Multi-modal Large Language Models
by: Li, Zeqian, et al.
Published: (2025)
by: Li, Zeqian, et al.
Published: (2025)
Weaver: End-to-End Agentic System Training for Video Interleaved Reasoning
by: Shi, Yudi, et al.
Published: (2026)
by: Shi, Yudi, et al.
Published: (2026)
DenseAnnotate: Enabling Scalable Dense Caption Collection for Images and 3D Scenes via Spoken Descriptions
by: Lin, Xiaoyu, et al.
Published: (2025)
by: Lin, Xiaoyu, et al.
Published: (2025)
DeVAn: Dense Video Annotation for Video-Language Models
by: Liu, Tingkai, et al.
Published: (2023)
by: Liu, Tingkai, et al.
Published: (2023)
Dense-Face: Personalized Face Generation Model via Dense Annotation Prediction
by: Guo, Xiao, et al.
Published: (2024)
by: Guo, Xiao, et al.
Published: (2024)
NOVA: Sparse Control, Dense Synthesis for Pair-Free Video Editing
by: Pan, Tianlin, et al.
Published: (2026)
by: Pan, Tianlin, et al.
Published: (2026)
Unified Dense Prediction of Video Diffusion
by: Yang, Lehan, et al.
Published: (2025)
by: Yang, Lehan, et al.
Published: (2025)
DenseScan: Advancing 3D Scene Understanding with 2D Dense Annotation
by: Wang, Zirui, et al.
Published: (2025)
by: Wang, Zirui, et al.
Published: (2025)
Streaming Dense Video Captioning
by: Zhou, Xingyi, et al.
Published: (2024)
by: Zhou, Xingyi, et al.
Published: (2024)
Technical Report for Soccernet 2023 -- Dense Video Captioning
by: Ruan, Zheng, et al.
Published: (2024)
by: Ruan, Zheng, et al.
Published: (2024)
PerSense: Training-Free Personalized Instance Segmentation in Dense Images
by: Siddiqui, Muhammad Ibraheem, et al.
Published: (2024)
by: Siddiqui, Muhammad Ibraheem, et al.
Published: (2024)
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World
by: Li, Xiangtai, et al.
Published: (2025)
by: Li, Xiangtai, et al.
Published: (2025)
RetinaGS: Scalable Training for Dense Scene Rendering with Billion-Scale 3D Gaussians
by: Li, Bingling, et al.
Published: (2024)
by: Li, Bingling, et al.
Published: (2024)
Grounded Question-Answering in Long Egocentric Videos
by: Di, Shangzhe, et al.
Published: (2023)
by: Di, Shangzhe, et al.
Published: (2023)
Colon-Bench: An Agentic Workflow for Scalable Dense Lesion Annotation in Full-Procedure Colonoscopy Videos
by: Hamdi, Abdullah, et al.
Published: (2026)
by: Hamdi, Abdullah, et al.
Published: (2026)
BiDense: Binarization for Dense Prediction
by: Yin, Rui, et al.
Published: (2024)
by: Yin, Rui, et al.
Published: (2024)
Dense-SfM: Structure from Motion with Dense Consistent Matching
by: Lee, JongMin, et al.
Published: (2025)
by: Lee, JongMin, et al.
Published: (2025)
Dense Matchers for Dense Tracking
by: Jelínek, Tomáš, et al.
Published: (2024)
by: Jelínek, Tomáš, et al.
Published: (2024)
Object-centric Video Question Answering with Visual Grounding and Referring
by: Wang, Haochen, et al.
Published: (2025)
by: Wang, Haochen, et al.
Published: (2025)
Exo2EgoDVC: Dense Video Captioning of Egocentric Procedural Activities Using Web Instructional Videos
by: Ohkawa, Takehiko, et al.
Published: (2023)
by: Ohkawa, Takehiko, et al.
Published: (2023)
Dense Video Object Captioning from Disjoint Supervision
by: Zhou, Xingyi, et al.
Published: (2023)
by: Zhou, Xingyi, et al.
Published: (2023)
Dense360: Dense Understanding from Omnidirectional Panoramas
by: Zhou, Yikang, et al.
Published: (2025)
by: Zhou, Yikang, et al.
Published: (2025)
Training-Free Dense Hand Contact Estimation with Multi-Modal Large Language Models
by: Jung, Daniel Sungho, et al.
Published: (2026)
by: Jung, Daniel Sungho, et al.
Published: (2026)
Towards PerSense++: Advancing Training-Free Personalized Instance Segmentation in Dense Images
by: Siddiqui, Muhammad Ibraheem, et al.
Published: (2025)
by: Siddiqui, Muhammad Ibraheem, et al.
Published: (2025)
Sparse-to-Dense: A Free Lunch for Lossless Acceleration of Video Understanding in LLMs
by: Zhang, Xuan, et al.
Published: (2025)
by: Zhang, Xuan, et al.
Published: (2025)
TEMSET-24K: Densely Annotated Dataset for Indexing Multipart Endoscopic Videos using Surgical Timeline Segmentation
by: Bilal, Muhammad, et al.
Published: (2025)
by: Bilal, Muhammad, et al.
Published: (2025)
A Unified Image-Dense Annotation Generation Model for Underwater Scenes
by: Lin, Hongkai, et al.
Published: (2025)
by: Lin, Hongkai, et al.
Published: (2025)
TDViT: Temporal Dilated Video Transformer for Dense Video Tasks
by: Sun, Guanxiong, et al.
Published: (2024)
by: Sun, Guanxiong, et al.
Published: (2024)
DiffKillR: Killing and Recreating Diffeomorphisms for Cell Annotation in Dense Microscopy Images
by: Liu, Chen, et al.
Published: (2024)
by: Liu, Chen, et al.
Published: (2024)
Exploring Temporal Event Cues for Dense Video Captioning in Cyclic Co-learning
by: Xie, Zhuyang, et al.
Published: (2024)
by: Xie, Zhuyang, et al.
Published: (2024)
TrajLoom: Dense Future Trajectory Generation from Video
by: Zhang, Zewei, et al.
Published: (2026)
by: Zhang, Zewei, et al.
Published: (2026)
DenseSR: Image Shadow Removal as Dense Prediction
by: Lin, Yu-Fan, et al.
Published: (2025)
by: Lin, Yu-Fan, et al.
Published: (2025)
Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions
by: Shen, Yijun, et al.
Published: (2025)
by: Shen, Yijun, et al.
Published: (2025)
What Matters When Repurposing Diffusion Models for General Dense Perception Tasks?
by: Xu, Guangkai, et al.
Published: (2024)
by: Xu, Guangkai, et al.
Published: (2024)
Dense Video Captioning Using Unsupervised Semantic Information
by: Estevam, Valter, et al.
Published: (2021)
by: Estevam, Valter, et al.
Published: (2021)
Step Differences in Instructional Video
by: Nagarajan, Tushar, et al.
Published: (2024)
by: Nagarajan, Tushar, et al.
Published: (2024)
Similar Items
-
Multi-Sentence Grounding for Long-term Instructional Video
by: Li, Zeqian, et al.
Published: (2023) -
Learning Streaming Video Representation via Multitask Training
by: Yan, Yibin, et al.
Published: (2025) -
Grounded Multi-Hop VideoQA in Long-Form Egocentric Videos
by: Chen, Qirui, et al.
Published: (2024) -
Enhancing Video-LLM Reasoning via Agent-of-Thoughts Distillation
by: Shi, Yudi, et al.
Published: (2024) -
Universal Video Temporal Grounding with Generative Multi-modal Large Language Models
by: Li, Zeqian, et al.
Published: (2025)