Scaling Dense Event-Stream Pretraining from Visual Foundation Models
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Zhiwen, Hou, Junhui, Zhu, Zhiyu, Wu, Jinjian, Shi, Guangming |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
E-Motion: Future Motion Simulation via Event Sequence Diffusion
by: Wu, Song, et al.
Published: (2024)
by: Wu, Song, et al.
Published: (2024)
Optimizing Multi-Modality Trackers via Significance-Regularized Tuning
by: Chen, Zhiwen, et al.
Published: (2025)
by: Chen, Zhiwen, et al.
Published: (2025)
Modeling State Shifting via Local-Global Distillation for Event-Frame Gaze Tracking
by: Li, Jiading, et al.
Published: (2024)
by: Li, Jiading, et al.
Published: (2024)
UniE2F: A Unified Diffusion Framework for Event-to-Frame Reconstruction with Video Foundation Models
by: Xu, Gang, et al.
Published: (2026)
by: Xu, Gang, et al.
Published: (2026)
Energy-oriented Diffusion Bridge for Image Restoration with Foundational Diffusion Models
by: Hou, Jinhui, et al.
Published: (2026)
by: Hou, Jinhui, et al.
Published: (2026)
Fast Window-Based Event Denoising with Spatiotemporal Correlation Enhancement
by: Fang, Huachen, et al.
Published: (2024)
by: Fang, Huachen, et al.
Published: (2024)
Self-supervised Learning of LiDAR 3D Point Clouds via 2D-3D Neural Calibration
by: Zhang, Yifan, et al.
Published: (2024)
by: Zhang, Yifan, et al.
Published: (2024)
Fine-grained Image-to-LiDAR Contrastive Distillation with Visual Foundation Models
by: Zhang, Yifan, et al.
Published: (2024)
by: Zhang, Yifan, et al.
Published: (2024)
Diffusion Image Generation with Explicit Modeling of Data Manifold Geometry
by: Xue, Duoduo, et al.
Published: (2026)
by: Xue, Duoduo, et al.
Published: (2026)
Spatial-Temporal Graph Enhanced DETR Towards Multi-Frame 3D Object Detection
by: Zhang, Yifan, et al.
Published: (2023)
by: Zhang, Yifan, et al.
Published: (2023)
NVS-Solver: Video Diffusion Model as Zero-Shot Novel View Synthesizer
by: You, Meng, et al.
Published: (2024)
by: You, Meng, et al.
Published: (2024)
Acc3D: Accelerating Single Image to 3D Diffusion Models via Edge Consistency Guided Score Distillation
by: Liu, Kendong, et al.
Published: (2025)
by: Liu, Kendong, et al.
Published: (2025)
ResFlow: Fine-tuning Residual Optical Flow for Event-based High Temporal Resolution Motion Estimation
by: Zhou, Qianang, et al.
Published: (2024)
by: Zhou, Qianang, et al.
Published: (2024)
Scaling Video Pretraining for Surgical Foundation Models
by: Lu, Sicheng, et al.
Published: (2026)
by: Lu, Sicheng, et al.
Published: (2026)
Learning Efficient and Effective Trajectories for Differential Equation-based Image Restoration
by: Zhu, Zhiyu, et al.
Published: (2024)
by: Zhu, Zhiyu, et al.
Published: (2024)
Generative Event Pretraining with Foundation Model Alignment
by: Cao, Jianwen, et al.
Published: (2026)
by: Cao, Jianwen, et al.
Published: (2026)
From Contrast to Consistency: Rethinking Event-based Continuous-Time Optical Flow Estimation
by: Hu, Rui, et al.
Published: (2026)
by: Hu, Rui, et al.
Published: (2026)
Visual Instruction Pretraining for Domain-Specific Foundation Models
by: Li, Yuxuan, et al.
Published: (2025)
by: Li, Yuxuan, et al.
Published: (2025)
PrefPaint: Aligning Image Inpainting Diffusion Model with Human Preference
by: Liu, Kendong, et al.
Published: (2024)
by: Liu, Kendong, et al.
Published: (2024)
GLENet: Boosting 3D Object Detectors with Generative Label Uncertainty Estimation
by: Zhang, Yifan, et al.
Published: (2022)
by: Zhang, Yifan, et al.
Published: (2022)
FoundationSLAM: Unleashing the Power of Depth Foundation Models for End-to-End Dense Visual SLAM
by: Wu, Yuchen, et al.
Published: (2025)
by: Wu, Yuchen, et al.
Published: (2025)
NYC-Event-VPR: A Large-Scale High-Resolution Event-Based Visual Place Recognition Dataset in Dense Urban Environments
by: Pan, Taiyi, et al.
Published: (2024)
by: Pan, Taiyi, et al.
Published: (2024)
Beyond ViT Tokens: Masked-Diffusion Pretrained Convolutional Pathology Foundation Model for Cell-Level Dense Prediction
by: Chen, Weiming, et al.
Published: (2026)
by: Chen, Weiming, et al.
Published: (2026)
Lotus: Diffusion-based Visual Foundation Model for High-quality Dense Prediction
by: He, Jing, et al.
Published: (2024)
by: He, Jing, et al.
Published: (2024)
CogVLM: Visual Expert for Pretrained Language Models
by: Wang, Weihan, et al.
Published: (2023)
by: Wang, Weihan, et al.
Published: (2023)
Language-Guided Visual Perception Disentanglement for Image Quality Assessment and Conditional Image Generation
by: Yang, Zhichao, et al.
Published: (2025)
by: Yang, Zhichao, et al.
Published: (2025)
Can Test-Time Scaling Improve World Foundation Model?
by: Cong, Wenyan, et al.
Published: (2025)
by: Cong, Wenyan, et al.
Published: (2025)
Temporal-Guided Visual Foundation Models for Event-Based Vision
by: Xia, Ruihao, et al.
Published: (2025)
by: Xia, Ruihao, et al.
Published: (2025)
Flow-Based Visual Stream Compression for Event Cameras
by: Stumpp, Daniel C., et al.
Published: (2024)
by: Stumpp, Daniel C., et al.
Published: (2024)
Multimodal Collaboration Networks for Geospatial Vehicle Detection in Dense, Occluded, and Large-Scale Events
by: Wu, Xin, et al.
Published: (2024)
by: Wu, Xin, et al.
Published: (2024)
Streaming Dense Video Captioning
by: Zhou, Xingyi, et al.
Published: (2024)
by: Zhou, Xingyi, et al.
Published: (2024)
Parse Graph-Based Visual-Language Interaction for Human Pose Estimation
by: Liu, Shibang, et al.
Published: (2025)
by: Liu, Shibang, et al.
Published: (2025)
InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
by: Chen, Zhe, et al.
Published: (2023)
by: Chen, Zhe, et al.
Published: (2023)
Enhancing Representation in Medical Vision-Language Foundation Models via Multi-Scale Information Extraction Techniques
by: Huang, Weijian, et al.
Published: (2024)
by: Huang, Weijian, et al.
Published: (2024)
DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding
by: Ahmadian, Mona, et al.
Published: (2025)
by: Ahmadian, Mona, et al.
Published: (2025)
GSStream: 3D Gaussian Splatting based Volumetric Scene Streaming System
by: Tang, Zhiye, et al.
Published: (2026)
by: Tang, Zhiye, et al.
Published: (2026)
Applications of Large Scale Foundation Models for Autonomous Driving
by: Huang, Yu, et al.
Published: (2023)
by: Huang, Yu, et al.
Published: (2023)
Deep Diversity-Enhanced Feature Representation of Hyperspectral Images
by: Hou, Jinhui, et al.
Published: (2023)
by: Hou, Jinhui, et al.
Published: (2023)
M2P: Improving Visual Foundation Models with Mask-to-Point Weakly-Supervised Learning for Dense Point Tracking
by: Wu, Qiangqiang, et al.
Published: (2026)
by: Wu, Qiangqiang, et al.
Published: (2026)
SemGauss-SLAM: Dense Semantic Gaussian Splatting SLAM
by: Zhu, Siting, et al.
Published: (2024)
by: Zhu, Siting, et al.
Published: (2024)
Similar Items
-
E-Motion: Future Motion Simulation via Event Sequence Diffusion
by: Wu, Song, et al.
Published: (2024) -
Optimizing Multi-Modality Trackers via Significance-Regularized Tuning
by: Chen, Zhiwen, et al.
Published: (2025) -
Modeling State Shifting via Local-Global Distillation for Event-Frame Gaze Tracking
by: Li, Jiading, et al.
Published: (2024) -
UniE2F: A Unified Diffusion Framework for Event-to-Frame Reconstruction with Video Foundation Models
by: Xu, Gang, et al.
Published: (2026) -
Energy-oriented Diffusion Bridge for Image Restoration with Foundational Diffusion Models
by: Hou, Jinhui, et al.
Published: (2026)