SnAG: Scalable and Accurate Video Grounding
Fuente:
arXiv
Saved in:
| Main Authors: | Mu, Fangzhou, Mo, Sicheng, Li, Yin |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Ctrl-X: Controlling Structure and Appearance for Text-To-Image Generation Without Guidance
by: Lin, Kuan Heng, et al.
Published: (2024)
by: Lin, Kuan Heng, et al.
Published: (2024)
HERO: Hierarchical Embedding-Refinement for Open-Vocabulary Temporal Sentence Grounding in Videos
by: Han, Tingting, et al.
Published: (2026)
by: Han, Tingting, et al.
Published: (2026)
AG-VPReID 2025: Aerial-Ground Video-based Person Re-identification Challenge Results
by: Nguyen, Kien, et al.
Published: (2025)
by: Nguyen, Kien, et al.
Published: (2025)
AG-VPReID.VIR: Bridging Aerial and Ground Platforms for Video-based Visible-Infrared Person Re-ID
by: Nguyen, Huy, et al.
Published: (2025)
by: Nguyen, Huy, et al.
Published: (2025)
AG-VPReID: A Challenging Large-Scale Benchmark for Aerial-Ground Video-based Person Re-Identification
by: Nguyen, Huy, et al.
Published: (2025)
by: Nguyen, Huy, et al.
Published: (2025)
Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding
by: Zheng, Minghang, et al.
Published: (2025)
by: Zheng, Minghang, et al.
Published: (2025)
Recovering Parametric Scenes from Very Few Time-of-Flight Pixels
by: Sifferman, Carter, et al.
Published: (2025)
by: Sifferman, Carter, et al.
Published: (2025)
Synthetic Object Compositions for Scalable and Accurate Learning in Detection, Segmentation, and Grounding
by: Huang, Weikai, et al.
Published: (2025)
by: Huang, Weikai, et al.
Published: (2025)
AG-ReID.v2: Bridging Aerial and Ground Views for Person Re-identification
by: Nguyen, Huy, et al.
Published: (2024)
by: Nguyen, Huy, et al.
Published: (2024)
TimeLoc: A Unified End-to-End Framework for Precise Timestamp Localization in Long Videos
by: Zhang, Chen-Lin, et al.
Published: (2025)
by: Zhang, Chen-Lin, et al.
Published: (2025)
Grounded Forcing: Bridging Time-Independent Semantics and Proximal Dynamics in Autoregressive Video Synthesis
by: Chen, Jintao, et al.
Published: (2026)
by: Chen, Jintao, et al.
Published: (2026)
Exploiting Auxiliary Caption for Video Grounding
by: Li, Hongxiang, et al.
Published: (2023)
by: Li, Hongxiang, et al.
Published: (2023)
Customizing Visual Emotion Evaluation for MLLMs: An Open-vocabulary, Multifaceted, and Scalable Approach
by: Wu, Daiqing, et al.
Published: (2025)
by: Wu, Daiqing, et al.
Published: (2025)
EvoGround: Self-Evolving Video Agents for Video Temporal Grounding
by: Jung, Minjoon, et al.
Published: (2026)
by: Jung, Minjoon, et al.
Published: (2026)
DiTVR: Zero-Shot Diffusion Transformer for Video Restoration
by: Gao, Sicheng, et al.
Published: (2025)
by: Gao, Sicheng, et al.
Published: (2025)
GIFStream: 4D Gaussian-based Immersive Video with Feature Stream
by: Li, Hao, et al.
Published: (2025)
by: Li, Hao, et al.
Published: (2025)
Chronologically Accurate Retrieval for Temporal Grounding of Motion-Language Models
by: Fujiwara, Kent, et al.
Published: (2024)
by: Fujiwara, Kent, et al.
Published: (2024)
LINEA: Fast and Accurate Line Detection Using Scalable Transformers
by: Janampa, Sebastian, et al.
Published: (2025)
by: Janampa, Sebastian, et al.
Published: (2025)
Video-GroundingDINO: Towards Open-Vocabulary Spatio-Temporal Video Grounding
by: Wasim, Syed Talal, et al.
Published: (2023)
by: Wasim, Syed Talal, et al.
Published: (2023)
VideoTIR: Accurate Understanding for Long Videos with Efficient Tool-Integrated Reasoning
by: Gao, Zhe, et al.
Published: (2026)
by: Gao, Zhe, et al.
Published: (2026)
Weakly-Supervised Affordance Grounding Guided by Part-Level Semantic Priors
by: Xu, Peiran, et al.
Published: (2025)
by: Xu, Peiran, et al.
Published: (2025)
Martian World Model: Controllable Video Synthesis with Physically Accurate 3D Reconstructions
by: Li, Longfei, et al.
Published: (2025)
by: Li, Longfei, et al.
Published: (2025)
Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation
by: Li, Jiaze, et al.
Published: (2026)
by: Li, Jiaze, et al.
Published: (2026)
GVDIFF: Grounded Text-to-Video Generation with Diffusion Models
by: Dou, Huanzhang, et al.
Published: (2024)
by: Dou, Huanzhang, et al.
Published: (2024)
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection
by: Zhuang, Weijun, et al.
Published: (2025)
by: Zhuang, Weijun, et al.
Published: (2025)
Clinically-Grounded Counterfactual Reasoning for Medical Video Diagnosis
by: Gao, Jianzhe, et al.
Published: (2026)
by: Gao, Jianzhe, et al.
Published: (2026)
MoGe: Unlocking Accurate Monocular Geometry Estimation for Open-Domain Images with Optimal Training Supervision
by: Wang, Ruicheng, et al.
Published: (2024)
by: Wang, Ruicheng, et al.
Published: (2024)
Multi-Object Tracking Retrieval with LLaVA-Video: A Training-Free Solution to MOT25-StAG Challenge
by: Yang, Yi, et al.
Published: (2025)
by: Yang, Yi, et al.
Published: (2025)
VideoMAP: Toward Scalable Mamba-based Video Autoregressive Pretraining
by: Liu, Yunze, et al.
Published: (2025)
by: Liu, Yunze, et al.
Published: (2025)
Mobile-VideoGPT: Fast and Accurate Model for Mobile Video Understanding
by: Shaker, Abdelrahman, et al.
Published: (2025)
by: Shaker, Abdelrahman, et al.
Published: (2025)
UrbanGS: A Scalable and Efficient Architecture for Geometrically Accurate Large-Scene Reconstruction
by: Li, Changbai, et al.
Published: (2026)
by: Li, Changbai, et al.
Published: (2026)
Grounded Video Caption Generation
by: Kazakos, Evangelos, et al.
Published: (2024)
by: Kazakos, Evangelos, et al.
Published: (2024)
Manifold-Aware Local Feature Modeling for Semi-Supervised Medical Image Segmentation
by: Shen, Sicheng, et al.
Published: (2024)
by: Shen, Sicheng, et al.
Published: (2024)
Foresee-to-Ground: From Predictive Temporal Perception to Evidence-Driven Reasoning for Video Temporal Grounding
by: Zheng, Zelin, et al.
Published: (2026)
by: Zheng, Zelin, et al.
Published: (2026)
Physics-Aware Video Instance Removal Benchmark
by: Li, Zirui, et al.
Published: (2026)
by: Li, Zirui, et al.
Published: (2026)
MoGe-2: Accurate Monocular Geometry with Metric Scale and Sharp Details
by: Wang, Ruicheng, et al.
Published: (2025)
by: Wang, Ruicheng, et al.
Published: (2025)
Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark
by: Yang, Shuyu, et al.
Published: (2025)
by: Yang, Shuyu, et al.
Published: (2025)
Gaussian Sequences with Multi-Scale Dynamics for 4D Reconstruction from Monocular Casual Videos
by: Li, Can, et al.
Published: (2026)
by: Li, Can, et al.
Published: (2026)
Multi-sentence Video Grounding for Long Video Generation
by: Feng, Wei, et al.
Published: (2024)
by: Feng, Wei, et al.
Published: (2024)
Long-Horizon Streaming Video Generation via Hybrid Attention with Decoupled Distillation
by: Li, Ruibin, et al.
Published: (2026)
by: Li, Ruibin, et al.
Published: (2026)
Similar Items
-
Ctrl-X: Controlling Structure and Appearance for Text-To-Image Generation Without Guidance
by: Lin, Kuan Heng, et al.
Published: (2024) -
HERO: Hierarchical Embedding-Refinement for Open-Vocabulary Temporal Sentence Grounding in Videos
by: Han, Tingting, et al.
Published: (2026) -
AG-VPReID 2025: Aerial-Ground Video-based Person Re-identification Challenge Results
by: Nguyen, Kien, et al.
Published: (2025) -
AG-VPReID.VIR: Bridging Aerial and Ground Platforms for Video-based Visible-Infrared Person Re-ID
by: Nguyen, Huy, et al.
Published: (2025) -
AG-VPReID: A Challenging Large-Scale Benchmark for Aerial-Ground Video-based Person Re-Identification
by: Nguyen, Huy, et al.
Published: (2025)