End-to-End Dense Video Grounding via Parallel Regression
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shi, Fengyuan, Huang, Weilin, Wang, Limin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2021
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Enhancing Traffic Safety with Parallel Dense Video Captioning for End-to-End Event Analysis
von: Shoman, Maged, et al.
Veröffentlicht: (2024)
von: Shoman, Maged, et al.
Veröffentlicht: (2024)
End-to-End Training for Autoregressive Video Diffusion via Self-Resampling
von: Guo, Yuwei, et al.
Veröffentlicht: (2025)
von: Guo, Yuwei, et al.
Veröffentlicht: (2025)
End-to-End Facial Expression Detection in Long Videos
von: Fang, Yini, et al.
Veröffentlicht: (2025)
von: Fang, Yini, et al.
Veröffentlicht: (2025)
RT-DETRv3: Real-time End-to-End Object Detection with Hierarchical Dense Positive Supervision
von: Wang, Shuo, et al.
Veröffentlicht: (2024)
von: Wang, Shuo, et al.
Veröffentlicht: (2024)
Weaver: End-to-End Agentic System Training for Video Interleaved Reasoning
von: Shi, Yudi, et al.
Veröffentlicht: (2026)
von: Shi, Yudi, et al.
Veröffentlicht: (2026)
MoGA: Mixture-of-Groups Attention for End-to-End Long Video Generation
von: Jia, Weinan, et al.
Veröffentlicht: (2025)
von: Jia, Weinan, et al.
Veröffentlicht: (2025)
ExploreVLA: Dense World Modeling and Exploration for End-to-End Autonomous Driving
von: Sheng, Zihao, et al.
Veröffentlicht: (2026)
von: Sheng, Zihao, et al.
Veröffentlicht: (2026)
ImmersePro: End-to-End Stereo Video Synthesis Via Implicit Disparity Learning
von: Shi, Jian, et al.
Veröffentlicht: (2024)
von: Shi, Jian, et al.
Veröffentlicht: (2024)
BIVDiff: A Training-Free Framework for General-Purpose Video Synthesis via Bridging Image and Video Diffusion Models
von: Shi, Fengyuan, et al.
Veröffentlicht: (2023)
von: Shi, Fengyuan, et al.
Veröffentlicht: (2023)
LMVC: An End-to-End Learned Multiview Video Coding Framework
von: Sheng, Xihua, et al.
Veröffentlicht: (2025)
von: Sheng, Xihua, et al.
Veröffentlicht: (2025)
A Paradigm Shift: Fully End-to-End Training for Temporal Sentence Grounding in Videos
von: He, Allen, et al.
Veröffentlicht: (2026)
von: He, Allen, et al.
Veröffentlicht: (2026)
Action Images: End-to-End Policy Learning via Multiview Video Generation
von: Zhen, Haoyu, et al.
Veröffentlicht: (2026)
von: Zhen, Haoyu, et al.
Veröffentlicht: (2026)
Leveraging Image Matching Toward End-to-End Relative Camera Pose Regression
von: Khatib, Fadi, et al.
Veröffentlicht: (2022)
von: Khatib, Fadi, et al.
Veröffentlicht: (2022)
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs
von: Cheng, Dabing, et al.
Veröffentlicht: (2025)
von: Cheng, Dabing, et al.
Veröffentlicht: (2025)
FoundationSLAM: Unleashing the Power of Depth Foundation Models for End-to-End Dense Visual SLAM
von: Wu, Yuchen, et al.
Veröffentlicht: (2025)
von: Wu, Yuchen, et al.
Veröffentlicht: (2025)
Drive&Gen: Co-Evaluating End-to-End Driving and Video Generation Models
von: Wang, Jiahao, et al.
Veröffentlicht: (2025)
von: Wang, Jiahao, et al.
Veröffentlicht: (2025)
End-to-End Multi-Person Pose Estimation with Pose-Aware Video Transformer
von: Yu, Yonghui, et al.
Veröffentlicht: (2025)
von: Yu, Yonghui, et al.
Veröffentlicht: (2025)
An End-to-End Framework for Video Multi-Person Pose Estimation
von: Wei, Zhihong
Veröffentlicht: (2025)
von: Wei, Zhihong
Veröffentlicht: (2025)
MoCha:End-to-End Video Character Replacement without Structural Guidance
von: Xu, Zhengbo, et al.
Veröffentlicht: (2026)
von: Xu, Zhengbo, et al.
Veröffentlicht: (2026)
SparseDrive: End-to-End Autonomous Driving via Sparse Scene Representation
von: Sun, Wenchao, et al.
Veröffentlicht: (2024)
von: Sun, Wenchao, et al.
Veröffentlicht: (2024)
EVE: Towards End-to-End Video Subtitle Extraction with Vision-Language Models
von: Yu, Haiyang, et al.
Veröffentlicht: (2025)
von: Yu, Haiyang, et al.
Veröffentlicht: (2025)
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding
von: Zheng, Henry, et al.
Veröffentlicht: (2025)
von: Zheng, Henry, et al.
Veröffentlicht: (2025)
SOLVE: Synergy of Language-Vision and End-to-End Networks for Autonomous Driving
von: Chen, Xuesong, et al.
Veröffentlicht: (2025)
von: Chen, Xuesong, et al.
Veröffentlicht: (2025)
End-to-End Streaming Video Temporal Action Segmentation with Reinforce Learning
von: Zhang, Jinrong, et al.
Veröffentlicht: (2023)
von: Zhang, Jinrong, et al.
Veröffentlicht: (2023)
End-to-End Probabilistic Geometry-Guided Regression for 6DoF Object Pose Estimation
von: Pöllabauer, Thomas, et al.
Veröffentlicht: (2024)
von: Pöllabauer, Thomas, et al.
Veröffentlicht: (2024)
Puzzles: Unbounded Video-Depth Augmentation for Scalable End-to-End 3D Reconstruction
von: Ma, Jiahao, et al.
Veröffentlicht: (2025)
von: Ma, Jiahao, et al.
Veröffentlicht: (2025)
Scanning Only Once: An End-to-end Framework for Fast Temporal Grounding in Long Videos
von: Pan, Yulin, et al.
Veröffentlicht: (2023)
von: Pan, Yulin, et al.
Veröffentlicht: (2023)
Drive-JEPA: Video JEPA Meets Multimodal Trajectory Distillation for End-to-End Driving
von: Wang, Linhan, et al.
Veröffentlicht: (2026)
von: Wang, Linhan, et al.
Veröffentlicht: (2026)
End4: End-to-end Denoising Diffusion for Diffusion-Based Inpainting Detection
von: Wang, Fei, et al.
Veröffentlicht: (2025)
von: Wang, Fei, et al.
Veröffentlicht: (2025)
End-to-End Visual Autonomous Parking via Control-Aided Attention
von: Chen, Chao, et al.
Veröffentlicht: (2025)
von: Chen, Chao, et al.
Veröffentlicht: (2025)
Differentiable NMS via Sinkhorn Matching for End-to-End Fabric Defect Detection
von: Lu, Zhengyang, et al.
Veröffentlicht: (2025)
von: Lu, Zhengyang, et al.
Veröffentlicht: (2025)
Unleashing Generalization of End-to-End Autonomous Driving with Controllable Long Video Generation
von: Ma, Enhui, et al.
Veröffentlicht: (2024)
von: Ma, Enhui, et al.
Veröffentlicht: (2024)
WorldJen: An End-to-End Multi-Dimensional Benchmark for Generative Video Models
von: Inbasekar, Karthik, et al.
Veröffentlicht: (2026)
von: Inbasekar, Karthik, et al.
Veröffentlicht: (2026)
SMTrack: End-to-End Trained Spiking Neural Networks for Multi-Object Tracking in RGB Videos
von: Zhong, Pengzhi, et al.
Veröffentlicht: (2025)
von: Zhong, Pengzhi, et al.
Veröffentlicht: (2025)
End-to-End Video Question Answering with Frame Scoring Mechanisms and Adaptive Sampling
von: Liang, Jianxin, et al.
Veröffentlicht: (2024)
von: Liang, Jianxin, et al.
Veröffentlicht: (2024)
SparseWorld: Enhancing End-to-End Autonomous Driving via World Models with Sparse Scene Representation
von: Wang, Ruoyu, et al.
Veröffentlicht: (2026)
von: Wang, Ruoyu, et al.
Veröffentlicht: (2026)
Deep-JGAC: End-to-End Deep Joint Geometry and Attribute Compression for Dense Colored Point Clouds
von: Zhang, Yun, et al.
Veröffentlicht: (2025)
von: Zhang, Yun, et al.
Veröffentlicht: (2025)
STORM: End-to-End Referring Multi-Object Tracking in Videos
von: Lu, Zijia, et al.
Veröffentlicht: (2026)
von: Lu, Zijia, et al.
Veröffentlicht: (2026)
End-to-End Vision Tokenizer Tuning
von: Wang, Wenxuan, et al.
Veröffentlicht: (2025)
von: Wang, Wenxuan, et al.
Veröffentlicht: (2025)
Universal Video Temporal Grounding with Generative Multi-modal Large Language Models
von: Li, Zeqian, et al.
Veröffentlicht: (2025)
von: Li, Zeqian, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Enhancing Traffic Safety with Parallel Dense Video Captioning for End-to-End Event Analysis
von: Shoman, Maged, et al.
Veröffentlicht: (2024) -
End-to-End Training for Autoregressive Video Diffusion via Self-Resampling
von: Guo, Yuwei, et al.
Veröffentlicht: (2025) -
End-to-End Facial Expression Detection in Long Videos
von: Fang, Yini, et al.
Veröffentlicht: (2025) -
RT-DETRv3: Real-time End-to-End Object Detection with Hierarchical Dense Positive Supervision
von: Wang, Shuo, et al.
Veröffentlicht: (2024) -
Weaver: End-to-End Agentic System Training for Video Interleaved Reasoning
von: Shi, Yudi, et al.
Veröffentlicht: (2026)