TimeLoc: A Unified End-to-End Framework for Precise Timestamp Localization in Long Videos
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Chen-Lin, Sui, Lin, Liu, Shuming, Mu, Fangzhou, Wang, Zhangcheng, Ghanem, Bernard |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
End-to-End Temporal Action Detection with 1B Parameters Across 1000 Frames
von: Liu, Shuming, et al.
Veröffentlicht: (2023)
von: Liu, Shuming, et al.
Veröffentlicht: (2023)
End-to-End Optimized Image Compression with the Frequency-Oriented Transform
von: Zhang, Yuefeng, et al.
Veröffentlicht: (2024)
von: Zhang, Yuefeng, et al.
Veröffentlicht: (2024)
Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation
von: Chen, Yuheng, et al.
Veröffentlicht: (2026)
von: Chen, Yuheng, et al.
Veröffentlicht: (2026)
Harnessing Temporal Causality for Advanced Temporal Action Detection
von: Liu, Shuming, et al.
Veröffentlicht: (2024)
von: Liu, Shuming, et al.
Veröffentlicht: (2024)
Multiscale Feature Importance-based Bit Allocation for End-to-End Feature Coding for Machines
von: Liu, Junle, et al.
Veröffentlicht: (2025)
von: Liu, Junle, et al.
Veröffentlicht: (2025)
DualComp: End-to-End Learning of a Unified Dual-Modality Lossless Compressor
von: Zhao, Yan, et al.
Veröffentlicht: (2025)
von: Zhao, Yan, et al.
Veröffentlicht: (2025)
Recent Advances of End-to-End Video Coding Technologies for AVS Standard Development
von: Sheng, Xihua, et al.
Veröffentlicht: (2026)
von: Sheng, Xihua, et al.
Veröffentlicht: (2026)
Deep-JGAC: End-to-End Deep Joint Geometry and Attribute Compression for Dense Colored Point Clouds
von: Zhang, Yun, et al.
Veröffentlicht: (2025)
von: Zhang, Yun, et al.
Veröffentlicht: (2025)
LongInsightBench: A Comprehensive Benchmark for Evaluating Omni-Modal Models on Human-Centric Long-Video Understanding
von: Han, ZhaoYang, et al.
Veröffentlicht: (2025)
von: Han, ZhaoYang, et al.
Veröffentlicht: (2025)
End-to-end Semantic-centric Video-based Multimodal Affective Computing
von: Lin, Ronghao, et al.
Veröffentlicht: (2024)
von: Lin, Ronghao, et al.
Veröffentlicht: (2024)
MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
von: Fang, Xinyu, et al.
Veröffentlicht: (2024)
von: Fang, Xinyu, et al.
Veröffentlicht: (2024)
One Framework to Rule Them All: Unifying Multimodal Tasks with LLM Neural-Tuning
von: Sun, Hao, et al.
Veröffentlicht: (2024)
von: Sun, Hao, et al.
Veröffentlicht: (2024)
Bridging Your Imagination with Audio-Video Generation via a Unified Director
von: Zhang, Jiaxu, et al.
Veröffentlicht: (2025)
von: Zhang, Jiaxu, et al.
Veröffentlicht: (2025)
VIA: Unified Spatiotemporal Video Adaptation Framework for Global and Local Video Editing
von: Gu, Jing, et al.
Veröffentlicht: (2024)
von: Gu, Jing, et al.
Veröffentlicht: (2024)
Error Analyses of Auto-Regressive Video Diffusion Models: A Unified Framework
von: Wang, Jing, et al.
Veröffentlicht: (2025)
von: Wang, Jing, et al.
Veröffentlicht: (2025)
Loc4Plan: Locating Before Planning for Outdoor Vision and Language Navigation
von: Tian, Huilin, et al.
Veröffentlicht: (2024)
von: Tian, Huilin, et al.
Veröffentlicht: (2024)
MotionPro: A Precise Motion Controller for Image-to-Video Generation
von: Zhang, Zhongwei, et al.
Veröffentlicht: (2025)
von: Zhang, Zhongwei, et al.
Veröffentlicht: (2025)
Perceptual Learned Image Compression via End-to-End JND-Based Optimization
von: Pakdaman, Farhad, et al.
Veröffentlicht: (2024)
von: Pakdaman, Farhad, et al.
Veröffentlicht: (2024)
Generative Frame Sampler for Long Video Understanding
von: Yao, Linli, et al.
Veröffentlicht: (2025)
von: Yao, Linli, et al.
Veröffentlicht: (2025)
End-to-End RGB-IR Joint Image Compression With Channel-wise Cross-modality Entropy Model
von: Wang, Haofeng, et al.
Veröffentlicht: (2025)
von: Wang, Haofeng, et al.
Veröffentlicht: (2025)
When Video Coding Meets Multimodal Large Language Models: A Unified Paradigm for Video Coding
von: Zhang, Pingping, et al.
Veröffentlicht: (2024)
von: Zhang, Pingping, et al.
Veröffentlicht: (2024)
Memory-enhanced Retrieval Augmentation for Long Video Understanding
von: Yuan, Huaying, et al.
Veröffentlicht: (2025)
von: Yuan, Huaying, et al.
Veröffentlicht: (2025)
UniTalking: A Unified Audio-Video Framework for Talking Portrait Generation
von: Li, Hebeizi, et al.
Veröffentlicht: (2026)
von: Li, Hebeizi, et al.
Veröffentlicht: (2026)
Joint End-to-End Image Compression and Denoising: Leveraging Contrastive Learning and Multi-Scale Self-ONNs
von: Xie, Yuxin, et al.
Veröffentlicht: (2024)
von: Xie, Yuxin, et al.
Veröffentlicht: (2024)
Where to Focus: Query-Modulated Multimodal Keyframe Selection for Long Video Understanding
von: Wang, Shaoguang, et al.
Veröffentlicht: (2026)
von: Wang, Shaoguang, et al.
Veröffentlicht: (2026)
Unify-Agent: A Unified Multimodal Agent for World-Grounded Image Synthesis
von: Chen, Shuang, et al.
Veröffentlicht: (2026)
von: Chen, Shuang, et al.
Veröffentlicht: (2026)
OmniScript: Towards Audio-Visual Script Generation for Long-Form Cinematic Video
von: Pu, Junfu, et al.
Veröffentlicht: (2026)
von: Pu, Junfu, et al.
Veröffentlicht: (2026)
SkyLink: Unifying Street-Satellite Geo-Localization via UAV-Mediated 3D Scene Alignment
von: Zhang, Hongyang, et al.
Veröffentlicht: (2025)
von: Zhang, Hongyang, et al.
Veröffentlicht: (2025)
VideoMem: Constructing, Analyzing, Predicting Short-term and Long-term Video Memorability
von: Cohendet, Romain, et al.
Veröffentlicht: (2018)
von: Cohendet, Romain, et al.
Veröffentlicht: (2018)
VAGU & GtS: LLM-Based Benchmark and Framework for Joint Video Anomaly Grounding and Understanding
von: Gao, Shibo, et al.
Veröffentlicht: (2025)
von: Gao, Shibo, et al.
Veröffentlicht: (2025)
Delving Deeper: Hierarchical Visual Perception for Robust Video-Text Retrieval
von: Xie, Zequn, et al.
Veröffentlicht: (2026)
von: Xie, Zequn, et al.
Veröffentlicht: (2026)
PG-Attack: A Precision-Guided Adversarial Attack Framework Against Vision Foundation Models for Autonomous Driving
von: Fu, Jiyuan, et al.
Veröffentlicht: (2024)
von: Fu, Jiyuan, et al.
Veröffentlicht: (2024)
UBiSS: A Unified Framework for Bimodal Semantic Summarization of Videos
von: Mei, Yuting, et al.
Veröffentlicht: (2024)
von: Mei, Yuting, et al.
Veröffentlicht: (2024)
JavisDiT++: Unified Modeling and Optimization for Joint Audio-Video Generation
von: Liu, Kai, et al.
Veröffentlicht: (2026)
von: Liu, Kai, et al.
Veröffentlicht: (2026)
Emotion-Qwen: A Unified Framework for Emotion and Vision Understanding
von: Huang, Dawei, et al.
Veröffentlicht: (2025)
von: Huang, Dawei, et al.
Veröffentlicht: (2025)
VidCompress: Memory-Enhanced Temporal Compression for Video Understanding in Large Language Models
von: Lan, Xiaohan, et al.
Veröffentlicht: (2024)
von: Lan, Xiaohan, et al.
Veröffentlicht: (2024)
Learning Video Context as Interleaved Multimodal Sequences
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2024)
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2024)
StreamingEval: A Unified Evaluation Protocol towards Realistic Streaming Video Understanding
von: Tang, Guowei, et al.
Veröffentlicht: (2026)
von: Tang, Guowei, et al.
Veröffentlicht: (2026)
Hybrid Local-Global Context Learning for Neural Video Compression
von: Zhai, Yongqi, et al.
Veröffentlicht: (2024)
von: Zhai, Yongqi, et al.
Veröffentlicht: (2024)
Omni2Sound: Towards Unified Video-Text-to-Audio Generation
von: Dai, Yusheng, et al.
Veröffentlicht: (2026)
von: Dai, Yusheng, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
End-to-End Temporal Action Detection with 1B Parameters Across 1000 Frames
von: Liu, Shuming, et al.
Veröffentlicht: (2023) -
End-to-End Optimized Image Compression with the Frequency-Oriented Transform
von: Zhang, Yuefeng, et al.
Veröffentlicht: (2024) -
Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation
von: Chen, Yuheng, et al.
Veröffentlicht: (2026) -
Harnessing Temporal Causality for Advanced Temporal Action Detection
von: Liu, Shuming, et al.
Veröffentlicht: (2024) -
Multiscale Feature Importance-based Bit Allocation for End-to-End Feature Coding for Machines
von: Liu, Junle, et al.
Veröffentlicht: (2025)