SceneScribe-1M: A Large-Scale Video Dataset with Comprehensive Geometric and Semantic Annotations
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Yunnan, Zheng, Kecheng, Wang, Jianyuan, Chen, Minghao, Novotny, David, Rupprecht, Christian, Xu, Yinghao, Zhu, Xing, Zeng, Wenjun, Jin, Xin, Shen, Yujun |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PoseDiffusion: Solving Pose Estimation via Diffusion-aided Bundle Adjustment
by: Wang, Jianyuan, et al.
Published: (2023)
by: Wang, Jianyuan, et al.
Published: (2023)
VGGT: Visual Geometry Grounded Transformer
by: Wang, Jianyuan, et al.
Published: (2025)
by: Wang, Jianyuan, et al.
Published: (2025)
Vision-Centric Activation and Coordination for Multimodal Large Language Models
by: Wang, Yunnan, et al.
Published: (2025)
by: Wang, Yunnan, et al.
Published: (2025)
Bridging Stereo Geometry and BEV Representation with Reliable Mutual Interaction for Semantic Scene Completion
by: Li, Bohan, et al.
Published: (2023)
by: Li, Bohan, et al.
Published: (2023)
FLARE: Feed-forward Geometry, Appearance and Camera Estimation from Uncalibrated Sparse Views
by: Zhang, Shangzhan, et al.
Published: (2025)
by: Zhang, Shangzhan, et al.
Published: (2025)
Scene Graph Disentanglement and Composition for Generalizable Complex Image Generation
by: Wang, Yunnan, et al.
Published: (2024)
by: Wang, Yunnan, et al.
Published: (2024)
Graph-based Unsupervised Disentangled Representation Learning via Multimodal Large Language Models
by: Xie, Baao, et al.
Published: (2024)
by: Xie, Baao, et al.
Published: (2024)
VideoCoT: A Video Chain-of-Thought Dataset with Active Annotation Tool
by: Wang, Yan, et al.
Published: (2024)
by: Wang, Yan, et al.
Published: (2024)
Scaling Instruction-Based Video Editing with a High-Quality Synthetic Dataset
by: Bai, Qingyan, et al.
Published: (2025)
by: Bai, Qingyan, et al.
Published: (2025)
Hierarchical Context Alignment with Disentangled Geometric and Temporal Modeling for Semantic Occupancy Prediction
by: Li, Bohan, et al.
Published: (2024)
by: Li, Bohan, et al.
Published: (2024)
VGGT-$Ω$
by: Wang, Jianyuan, et al.
Published: (2026)
by: Wang, Jianyuan, et al.
Published: (2026)
Spatial Steerability of GANs via Self-Supervision from Discriminator
by: Wang, Jianyuan, et al.
Published: (2023)
by: Wang, Jianyuan, et al.
Published: (2023)
SynMotion: Semantic-Visual Adaptation for Motion Customized Video Generation
by: Tan, Shuai, et al.
Published: (2025)
by: Tan, Shuai, et al.
Published: (2025)
HowToCaption: Prompting LLMs to Transform Video Annotations at Scale
by: Shvetsova, Nina, et al.
Published: (2023)
by: Shvetsova, Nina, et al.
Published: (2023)
CoTracker3: Simpler and Better Point Tracking by Pseudo-Labelling Real Videos
by: Karaev, Nikita, et al.
Published: (2024)
by: Karaev, Nikita, et al.
Published: (2024)
Hierarchical Temporal Context Learning for Camera-based Semantic Scene Completion
by: Li, Bohan, et al.
Published: (2024)
by: Li, Bohan, et al.
Published: (2024)
LISNeRF Mapping: LiDAR-based Implicit Mapping via Semantic Neural Fields for Large-Scale 3D Scenes
by: Zhang, Jianyuan, et al.
Published: (2023)
by: Zhang, Jianyuan, et al.
Published: (2023)
OccScene: Semantic Occupancy-based Cross-task Mutual Learning for 3D Scene Generation
by: Li, Bohan, et al.
Published: (2024)
by: Li, Bohan, et al.
Published: (2024)
PartGen: Part-level 3D Generation and Reconstruction with Multi-View Diffusion Models
by: Chen, Minghao, et al.
Published: (2024)
by: Chen, Minghao, et al.
Published: (2024)
Contextual AD Narration with Interleaved Multimodal Sequence
by: Wang, Hanlin, et al.
Published: (2024)
by: Wang, Hanlin, et al.
Published: (2024)
Benchmarking Large Vision-Language Models via Directed Scene Graph for Comprehensive Image Captioning
by: Lu, Fan, et al.
Published: (2024)
by: Lu, Fan, et al.
Published: (2024)
Learning Naturally Aggregated Appearance for Efficient 3D Editing
by: Cheng, Ka Leong, et al.
Published: (2023)
by: Cheng, Ka Leong, et al.
Published: (2023)
SpatialVID: A Large-Scale Video Dataset with Spatial Annotations
by: Wang, Jiahao, et al.
Published: (2025)
by: Wang, Jiahao, et al.
Published: (2025)
Feed-Forward SceneDINO for Unsupervised Semantic Scene Completion
by: Jevtić, Aleksandar, et al.
Published: (2025)
by: Jevtić, Aleksandar, et al.
Published: (2025)
Disentangled World Models: Learning to Transfer Semantic Knowledge from Distracting Videos for Reinforcement Learning
by: Wang, Qi, et al.
Published: (2025)
by: Wang, Qi, et al.
Published: (2025)
DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge
by: Zhang, Wenyao, et al.
Published: (2025)
by: Zhang, Wenyao, et al.
Published: (2025)
On the sub-adjacent Hopf algebra of the universal enveloping algebra of a post-Lie algebra
by: Li, Yunnan
Published: (2024)
by: Li, Yunnan
Published: (2024)
Matched pairs and Yang-Baxter operators
by: Li, Yunnan
Published: (2025)
by: Li, Yunnan
Published: (2025)
Label-efficient Semantic Scene Completion with Scribble Annotations
by: Wang, Song, et al.
Published: (2024)
by: Wang, Song, et al.
Published: (2024)
Diffusion Models for Image Restoration and Enhancement: A Comprehensive Survey
by: Li, Xin, et al.
Published: (2023)
by: Li, Xin, et al.
Published: (2023)
VGGHeads: 3D Multi Head Alignment with a Large-Scale Synthetic Dataset
by: Kupyn, Orest, et al.
Published: (2024)
by: Kupyn, Orest, et al.
Published: (2024)
B-Fredholm theory in Banach algebras
by: Zhang, Yunnan, et al.
Published: (2024)
by: Zhang, Yunnan, et al.
Published: (2024)
MultiEgo: A Multi-View Egocentric Video Dataset for 4D Scene Reconstruction
by: Li, Bate, et al.
Published: (2025)
by: Li, Bate, et al.
Published: (2025)
Hybrid-grained Feature Aggregation with Coarse-to-fine Language Guidance for Self-supervised Monocular Depth Estimation
by: Zhang, Wenyao, et al.
Published: (2025)
by: Zhang, Wenyao, et al.
Published: (2025)
MotionBank: A Large-scale Video Motion Benchmark with Disentangled Rule-based Annotations
by: Xu, Liang, et al.
Published: (2024)
by: Xu, Liang, et al.
Published: (2024)
MSVCOD:A Large-Scale Multi-Scene Dataset for Video Camouflage Object Detection
by: Gao, Shuyong, et al.
Published: (2025)
by: Gao, Shuyong, et al.
Published: (2025)
GA-VLN: Geometry-Aware BEV Representation for Efficient Vision-Language Navigation
by: Yang, Jiahao, et al.
Published: (2026)
by: Yang, Jiahao, et al.
Published: (2026)
Dataset Enhancement with Instance-Level Augmentations
by: Kupyn, Orest, et al.
Published: (2024)
by: Kupyn, Orest, et al.
Published: (2024)
ScribeAgent: Towards Specialized Web Agents Using Production-Scale Workflow Data
by: Shen, Junhong, et al.
Published: (2024)
by: Shen, Junhong, et al.
Published: (2024)
CoDeF: Content Deformation Fields for Temporally Consistent Video Processing
by: Ouyang, Hao, et al.
Published: (2023)
by: Ouyang, Hao, et al.
Published: (2023)
Similar Items
-
PoseDiffusion: Solving Pose Estimation via Diffusion-aided Bundle Adjustment
by: Wang, Jianyuan, et al.
Published: (2023) -
VGGT: Visual Geometry Grounded Transformer
by: Wang, Jianyuan, et al.
Published: (2025) -
Vision-Centric Activation and Coordination for Multimodal Large Language Models
by: Wang, Yunnan, et al.
Published: (2025) -
Bridging Stereo Geometry and BEV Representation with Reliable Mutual Interaction for Semantic Scene Completion
by: Li, Bohan, et al.
Published: (2023) -
FLARE: Feed-forward Geometry, Appearance and Camera Estimation from Uncalibrated Sparse Views
by: Zhang, Shangzhan, et al.
Published: (2025)