SnAG: Scalable and Accurate Video Grounding
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Mu, Fangzhou, Mo, Sicheng, Li, Yin |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Ctrl-X: Controlling Structure and Appearance for Text-To-Image Generation Without Guidance
par: Lin, Kuan Heng, et autres
Publié: (2024)
par: Lin, Kuan Heng, et autres
Publié: (2024)
HERO: Hierarchical Embedding-Refinement for Open-Vocabulary Temporal Sentence Grounding in Videos
par: Han, Tingting, et autres
Publié: (2026)
par: Han, Tingting, et autres
Publié: (2026)
AG-VPReID 2025: Aerial-Ground Video-based Person Re-identification Challenge Results
par: Nguyen, Kien, et autres
Publié: (2025)
par: Nguyen, Kien, et autres
Publié: (2025)
AG-VPReID.VIR: Bridging Aerial and Ground Platforms for Video-based Visible-Infrared Person Re-ID
par: Nguyen, Huy, et autres
Publié: (2025)
par: Nguyen, Huy, et autres
Publié: (2025)
AG-VPReID: A Challenging Large-Scale Benchmark for Aerial-Ground Video-based Person Re-Identification
par: Nguyen, Huy, et autres
Publié: (2025)
par: Nguyen, Huy, et autres
Publié: (2025)
Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding
par: Zheng, Minghang, et autres
Publié: (2025)
par: Zheng, Minghang, et autres
Publié: (2025)
Recovering Parametric Scenes from Very Few Time-of-Flight Pixels
par: Sifferman, Carter, et autres
Publié: (2025)
par: Sifferman, Carter, et autres
Publié: (2025)
Synthetic Object Compositions for Scalable and Accurate Learning in Detection, Segmentation, and Grounding
par: Huang, Weikai, et autres
Publié: (2025)
par: Huang, Weikai, et autres
Publié: (2025)
AG-ReID.v2: Bridging Aerial and Ground Views for Person Re-identification
par: Nguyen, Huy, et autres
Publié: (2024)
par: Nguyen, Huy, et autres
Publié: (2024)
TimeLoc: A Unified End-to-End Framework for Precise Timestamp Localization in Long Videos
par: Zhang, Chen-Lin, et autres
Publié: (2025)
par: Zhang, Chen-Lin, et autres
Publié: (2025)
Grounded Forcing: Bridging Time-Independent Semantics and Proximal Dynamics in Autoregressive Video Synthesis
par: Chen, Jintao, et autres
Publié: (2026)
par: Chen, Jintao, et autres
Publié: (2026)
Exploiting Auxiliary Caption for Video Grounding
par: Li, Hongxiang, et autres
Publié: (2023)
par: Li, Hongxiang, et autres
Publié: (2023)
Customizing Visual Emotion Evaluation for MLLMs: An Open-vocabulary, Multifaceted, and Scalable Approach
par: Wu, Daiqing, et autres
Publié: (2025)
par: Wu, Daiqing, et autres
Publié: (2025)
EvoGround: Self-Evolving Video Agents for Video Temporal Grounding
par: Jung, Minjoon, et autres
Publié: (2026)
par: Jung, Minjoon, et autres
Publié: (2026)
DiTVR: Zero-Shot Diffusion Transformer for Video Restoration
par: Gao, Sicheng, et autres
Publié: (2025)
par: Gao, Sicheng, et autres
Publié: (2025)
GIFStream: 4D Gaussian-based Immersive Video with Feature Stream
par: Li, Hao, et autres
Publié: (2025)
par: Li, Hao, et autres
Publié: (2025)
Chronologically Accurate Retrieval for Temporal Grounding of Motion-Language Models
par: Fujiwara, Kent, et autres
Publié: (2024)
par: Fujiwara, Kent, et autres
Publié: (2024)
LINEA: Fast and Accurate Line Detection Using Scalable Transformers
par: Janampa, Sebastian, et autres
Publié: (2025)
par: Janampa, Sebastian, et autres
Publié: (2025)
Video-GroundingDINO: Towards Open-Vocabulary Spatio-Temporal Video Grounding
par: Wasim, Syed Talal, et autres
Publié: (2023)
par: Wasim, Syed Talal, et autres
Publié: (2023)
VideoTIR: Accurate Understanding for Long Videos with Efficient Tool-Integrated Reasoning
par: Gao, Zhe, et autres
Publié: (2026)
par: Gao, Zhe, et autres
Publié: (2026)
Weakly-Supervised Affordance Grounding Guided by Part-Level Semantic Priors
par: Xu, Peiran, et autres
Publié: (2025)
par: Xu, Peiran, et autres
Publié: (2025)
Martian World Model: Controllable Video Synthesis with Physically Accurate 3D Reconstructions
par: Li, Longfei, et autres
Publié: (2025)
par: Li, Longfei, et autres
Publié: (2025)
Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation
par: Li, Jiaze, et autres
Publié: (2026)
par: Li, Jiaze, et autres
Publié: (2026)
GVDIFF: Grounded Text-to-Video Generation with Diffusion Models
par: Dou, Huanzhang, et autres
Publié: (2024)
par: Dou, Huanzhang, et autres
Publié: (2024)
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection
par: Zhuang, Weijun, et autres
Publié: (2025)
par: Zhuang, Weijun, et autres
Publié: (2025)
Clinically-Grounded Counterfactual Reasoning for Medical Video Diagnosis
par: Gao, Jianzhe, et autres
Publié: (2026)
par: Gao, Jianzhe, et autres
Publié: (2026)
MoGe: Unlocking Accurate Monocular Geometry Estimation for Open-Domain Images with Optimal Training Supervision
par: Wang, Ruicheng, et autres
Publié: (2024)
par: Wang, Ruicheng, et autres
Publié: (2024)
Multi-Object Tracking Retrieval with LLaVA-Video: A Training-Free Solution to MOT25-StAG Challenge
par: Yang, Yi, et autres
Publié: (2025)
par: Yang, Yi, et autres
Publié: (2025)
VideoMAP: Toward Scalable Mamba-based Video Autoregressive Pretraining
par: Liu, Yunze, et autres
Publié: (2025)
par: Liu, Yunze, et autres
Publié: (2025)
Mobile-VideoGPT: Fast and Accurate Model for Mobile Video Understanding
par: Shaker, Abdelrahman, et autres
Publié: (2025)
par: Shaker, Abdelrahman, et autres
Publié: (2025)
UrbanGS: A Scalable and Efficient Architecture for Geometrically Accurate Large-Scene Reconstruction
par: Li, Changbai, et autres
Publié: (2026)
par: Li, Changbai, et autres
Publié: (2026)
Grounded Video Caption Generation
par: Kazakos, Evangelos, et autres
Publié: (2024)
par: Kazakos, Evangelos, et autres
Publié: (2024)
Manifold-Aware Local Feature Modeling for Semi-Supervised Medical Image Segmentation
par: Shen, Sicheng, et autres
Publié: (2024)
par: Shen, Sicheng, et autres
Publié: (2024)
Foresee-to-Ground: From Predictive Temporal Perception to Evidence-Driven Reasoning for Video Temporal Grounding
par: Zheng, Zelin, et autres
Publié: (2026)
par: Zheng, Zelin, et autres
Publié: (2026)
Physics-Aware Video Instance Removal Benchmark
par: Li, Zirui, et autres
Publié: (2026)
par: Li, Zirui, et autres
Publié: (2026)
MoGe-2: Accurate Monocular Geometry with Metric Scale and Sharp Details
par: Wang, Ruicheng, et autres
Publié: (2025)
par: Wang, Ruicheng, et autres
Publié: (2025)
Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark
par: Yang, Shuyu, et autres
Publié: (2025)
par: Yang, Shuyu, et autres
Publié: (2025)
Gaussian Sequences with Multi-Scale Dynamics for 4D Reconstruction from Monocular Casual Videos
par: Li, Can, et autres
Publié: (2026)
par: Li, Can, et autres
Publié: (2026)
Multi-sentence Video Grounding for Long Video Generation
par: Feng, Wei, et autres
Publié: (2024)
par: Feng, Wei, et autres
Publié: (2024)
Long-Horizon Streaming Video Generation via Hybrid Attention with Decoupled Distillation
par: Li, Ruibin, et autres
Publié: (2026)
par: Li, Ruibin, et autres
Publié: (2026)
Documents similaires
-
Ctrl-X: Controlling Structure and Appearance for Text-To-Image Generation Without Guidance
par: Lin, Kuan Heng, et autres
Publié: (2024) -
HERO: Hierarchical Embedding-Refinement for Open-Vocabulary Temporal Sentence Grounding in Videos
par: Han, Tingting, et autres
Publié: (2026) -
AG-VPReID 2025: Aerial-Ground Video-based Person Re-identification Challenge Results
par: Nguyen, Kien, et autres
Publié: (2025) -
AG-VPReID.VIR: Bridging Aerial and Ground Platforms for Video-based Visible-Infrared Person Re-ID
par: Nguyen, Huy, et autres
Publié: (2025) -
AG-VPReID: A Challenging Large-Scale Benchmark for Aerial-Ground Video-based Person Re-Identification
par: Nguyen, Huy, et autres
Publié: (2025)