Length Matters: Length-Aware Transformer for Temporal Sentence Grounding
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Yifan, Liu, Ziyi, Sun, Xiaolong, Wang, Jiawei, Liu, Hongmin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Diversifying Query: Region-Guided Transformer for Temporal Sentence Grounding
by: Sun, Xiaolong, et al.
Published: (2024)
by: Sun, Xiaolong, et al.
Published: (2024)
Hierarchical Local-Global Transformer for Temporal Sentence Grounding
by: Fang, Xiang, et al.
Published: (2022)
by: Fang, Xiang, et al.
Published: (2022)
Contrast-Unity for Partially-Supervised Temporal Sentence Grounding
by: Wang, Haicheng, et al.
Published: (2025)
by: Wang, Haicheng, et al.
Published: (2025)
Moment Quantization for Video Temporal Grounding
by: Sun, Xiaolong, et al.
Published: (2025)
by: Sun, Xiaolong, et al.
Published: (2025)
Weakly Supervised Temporal Sentence Grounding via Positive Sample Mining
by: Dong, Lu, et al.
Published: (2025)
by: Dong, Lu, et al.
Published: (2025)
Multi-Pair Temporal Sentence Grounding via Multi-Thread Knowledge Transfer Network
by: Fang, Xiang, et al.
Published: (2024)
by: Fang, Xiang, et al.
Published: (2024)
Sim-DETR: Unlock DETR for Temporal Sentence Grounding
by: Tang, Jiajin, et al.
Published: (2025)
by: Tang, Jiajin, et al.
Published: (2025)
A Paradigm Shift: Fully End-to-End Training for Temporal Sentence Grounding in Videos
by: He, Allen, et al.
Published: (2026)
by: He, Allen, et al.
Published: (2026)
Boosting Temporal Sentence Grounding via Causal Inference
by: Tang, Kefan, et al.
Published: (2025)
by: Tang, Kefan, et al.
Published: (2025)
Sequence Length Scaling in Vision Transformers for Scientific Images on Frontier
by: Tsaris, Aristeidis, et al.
Published: (2024)
by: Tsaris, Aristeidis, et al.
Published: (2024)
MomentMix Augmentation with Length-Aware DETR for Temporally Robust Moment Retrieval
by: Park, Seojeong, et al.
Published: (2024)
by: Park, Seojeong, et al.
Published: (2024)
Length-Aware Motion Synthesis via Latent Diffusion
by: Sampieri, Alessio, et al.
Published: (2024)
by: Sampieri, Alessio, et al.
Published: (2024)
Efficient Temporal Sentence Grounding in Videos with Multi-Teacher Knowledge Distillation
by: Liang, Renjie, et al.
Published: (2023)
by: Liang, Renjie, et al.
Published: (2023)
Bias-Conflict Sample Synthesis and Adversarial Removal Debias Strategy for Temporal Sentence Grounding in Video
by: Qi, Zhaobo, et al.
Published: (2024)
by: Qi, Zhaobo, et al.
Published: (2024)
Two-Stage Decoupling Framework for Variable-Length Glaucoma Prognosis
by: Song, Yiran, et al.
Published: (2025)
by: Song, Yiran, et al.
Published: (2025)
BAM-DETR: Boundary-Aligned Moment Detection Transformer for Temporal Sentence Grounding in Videos
by: Lee, Pilhyeon, et al.
Published: (2023)
by: Lee, Pilhyeon, et al.
Published: (2023)
Adapting to Length Shift: FlexiLength Network for Trajectory Prediction
by: Xu, Yi, et al.
Published: (2024)
by: Xu, Yi, et al.
Published: (2024)
HERO: Hierarchical Embedding-Refinement for Open-Vocabulary Temporal Sentence Grounding in Videos
by: Han, Tingting, et al.
Published: (2026)
by: Han, Tingting, et al.
Published: (2026)
Annotations Are Not All You Need: A Cross-modal Knowledge Transfer Network for Unsupervised Temporal Sentence Grounding
by: Fang, Xiang, et al.
Published: (2026)
by: Fang, Xiang, et al.
Published: (2026)
Empower Words: DualGround for Structured Phrase and Sentence-Level Temporal Grounding
by: Kang, Minseok, et al.
Published: (2025)
by: Kang, Minseok, et al.
Published: (2025)
UniLayDiff: A Unified Diffusion Transformer for Content-Aware Layout Generation
by: Liu, Zeyang, et al.
Published: (2025)
by: Liu, Zeyang, et al.
Published: (2025)
Let Me Finish My Sentence: Video Temporal Grounding with Holistic Text Understanding
by: Woo, Jongbhin, et al.
Published: (2024)
by: Woo, Jongbhin, et al.
Published: (2024)
Frame Order Matters: A Temporal Sequence-Aware Model for Few-Shot Action Recognition
by: Li, Bozheng, et al.
Published: (2024)
by: Li, Bozheng, et al.
Published: (2024)
Frozen LLMs as Map-Aware Spatio-Temporal Reasoners for Vehicle Trajectory Prediction
by: Liu, Yanjiao, et al.
Published: (2026)
by: Liu, Yanjiao, et al.
Published: (2026)
Multi-Sentence Grounding for Long-term Instructional Video
by: Li, Zeqian, et al.
Published: (2023)
by: Li, Zeqian, et al.
Published: (2023)
Scaling up Multi-domain Semantic Segmentation with Sentence Embeddings
by: Yin, Wei, et al.
Published: (2022)
by: Yin, Wei, et al.
Published: (2022)
LEDiT: Your Length-Extrapolatable Diffusion Transformer without Positional Encoding
by: Zhang, Shen, et al.
Published: (2025)
by: Zhang, Shen, et al.
Published: (2025)
RIFLEx: A Free Lunch for Length Extrapolation in Video Diffusion Transformers
by: Zhao, Min, et al.
Published: (2025)
by: Zhao, Min, et al.
Published: (2025)
Progressive Token Length Scaling in Transformer Encoders for Efficient Universal Segmentation
by: Aich, Abhishek, et al.
Published: (2024)
by: Aich, Abhishek, et al.
Published: (2024)
Don't Look Twice: Faster Video Transformers with Run-Length Tokenization
by: Choudhury, Rohan, et al.
Published: (2024)
by: Choudhury, Rohan, et al.
Published: (2024)
ALTo: Adaptive-Length Tokenizer for Autoregressive Mask Generation
by: Wang, Lingfeng, et al.
Published: (2025)
by: Wang, Lingfeng, et al.
Published: (2025)
Prototype-Aware Multimodal Alignment for Open-Vocabulary Visual Grounding
by: Xie, Jiangnan, et al.
Published: (2025)
by: Xie, Jiangnan, et al.
Published: (2025)
Temporal Adaptive RGBT Tracking with Modality Prompt
by: Wang, Hongyu, et al.
Published: (2024)
by: Wang, Hongyu, et al.
Published: (2024)
Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length
by: Huang, Yubo, et al.
Published: (2025)
by: Huang, Yubo, et al.
Published: (2025)
AVID: Any-Length Video Inpainting with Diffusion Model
by: Zhang, Zhixing, et al.
Published: (2023)
by: Zhang, Zhixing, et al.
Published: (2023)
RMT: Retentive Networks Meet Vision Transformers
by: Fan, Qihang, et al.
Published: (2023)
by: Fan, Qihang, et al.
Published: (2023)
Advancing Vision Transformer with Enhanced Spatial Priors
by: Fan, Qihang, et al.
Published: (2026)
by: Fan, Qihang, et al.
Published: (2026)
Images are Worth Variable Length of Representations
by: Mao, Lingjun, et al.
Published: (2025)
by: Mao, Lingjun, et al.
Published: (2025)
LoopExpose: An Unsupervised Framework for Arbitrary-Length Exposure Correction
by: Li, Ao, et al.
Published: (2025)
by: Li, Ao, et al.
Published: (2025)
MLP: Motion Label Prior for Temporal Sentence Localization in Untrimmed 3D Human Motions
by: Yan, Sheng, et al.
Published: (2024)
by: Yan, Sheng, et al.
Published: (2024)
Similar Items
-
Diversifying Query: Region-Guided Transformer for Temporal Sentence Grounding
by: Sun, Xiaolong, et al.
Published: (2024) -
Hierarchical Local-Global Transformer for Temporal Sentence Grounding
by: Fang, Xiang, et al.
Published: (2022) -
Contrast-Unity for Partially-Supervised Temporal Sentence Grounding
by: Wang, Haicheng, et al.
Published: (2025) -
Moment Quantization for Video Temporal Grounding
by: Sun, Xiaolong, et al.
Published: (2025) -
Weakly Supervised Temporal Sentence Grounding via Positive Sample Mining
by: Dong, Lu, et al.
Published: (2025)