Structured Context Learning for Generic Event Boundary Detection
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gu, Xin, Li, Congcong, Wang, Xinyao, Hong, Dexiang, Zhang, Libo, Luo, Tiejian, Wen, Longyin, Fan, Heng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Context-Guided Spatio-Temporal Video Grounding
von: Gu, Xin, et al.
Veröffentlicht: (2024)
von: Gu, Xin, et al.
Veröffentlicht: (2024)
Edit3K: Universal Representation Learning for Video Editing Components
von: Gu, Xin, et al.
Veröffentlicht: (2024)
von: Gu, Xin, et al.
Veröffentlicht: (2024)
Multi-Reward as Condition for Instruction-based Image Editing
von: Gu, Xin, et al.
Veröffentlicht: (2024)
von: Gu, Xin, et al.
Veröffentlicht: (2024)
Robust Domain Adaptive Object Detection with Unified Multi-Granularity Alignment
von: Zhang, Libo, et al.
Veröffentlicht: (2023)
von: Zhang, Libo, et al.
Veröffentlicht: (2023)
Accurate and Fast Compressed Video Captioning
von: Shen, Yaojie, et al.
Veröffentlicht: (2023)
von: Shen, Yaojie, et al.
Veröffentlicht: (2023)
Knowing Your Target: Target-Aware Transformer Makes Better Spatio-Temporal Video Grounding
von: Gu, Xin, et al.
Veröffentlicht: (2025)
von: Gu, Xin, et al.
Veröffentlicht: (2025)
Thinking With Bounding Boxes: Enhancing Spatio-Temporal Video Grounding via Reinforcement Fine-Tuning
von: Gu, Xin, et al.
Veröffentlicht: (2025)
von: Gu, Xin, et al.
Veröffentlicht: (2025)
EchoFoley: Event-Centric Hierarchical Control for Video Grounded Creative Sound Generation
von: Li, Bingxuan, et al.
Veröffentlicht: (2025)
von: Li, Bingxuan, et al.
Veröffentlicht: (2025)
Fine-grained Dynamic Network for Generic Event Boundary Detection
von: Zheng, Ziwei, et al.
Veröffentlicht: (2024)
von: Zheng, Ziwei, et al.
Veröffentlicht: (2024)
Flow-Guided Diffusion for Video Inpainting
von: Gu, Bohai, et al.
Veröffentlicht: (2023)
von: Gu, Bohai, et al.
Veröffentlicht: (2023)
CuMo: Scaling Multimodal LLM with Co-Upcycled Mixture-of-Experts
von: Li, Jiachen, et al.
Veröffentlicht: (2024)
von: Li, Jiachen, et al.
Veröffentlicht: (2024)
Robust Ego-Exo Correspondence with Long-Term Memory
von: Hu, Yijun, et al.
Veröffentlicht: (2025)
von: Hu, Yijun, et al.
Veröffentlicht: (2025)
Beyond Raw Videos: Understanding Edited Videos with Large Multimodal Model
von: Xu, Lu, et al.
Veröffentlicht: (2024)
von: Xu, Lu, et al.
Veröffentlicht: (2024)
CGTrack: Cascade Gating Network with Hierarchical Feature Aggregation for UAV Tracking
von: Li, Weihong, et al.
Veröffentlicht: (2025)
von: Li, Weihong, et al.
Veröffentlicht: (2025)
Towards Long-Form Spatio-Temporal Video Grounding
von: Gu, Xin, et al.
Veröffentlicht: (2026)
von: Gu, Xin, et al.
Veröffentlicht: (2026)
Rethinking the Architecture Design for Efficient Generic Event Boundary Detection
von: Zheng, Ziwei, et al.
Veröffentlicht: (2024)
von: Zheng, Ziwei, et al.
Veröffentlicht: (2024)
D-Attn: Decomposed Attention for Large Vision-and-Language Models
von: Kuo, Chia-Wen, et al.
Veröffentlicht: (2025)
von: Kuo, Chia-Wen, et al.
Veröffentlicht: (2025)
BPDO:Boundary Points Dynamic Optimization for Arbitrary Shape Scene Text Detection
von: Zheng, Jinzhi, et al.
Veröffentlicht: (2024)
von: Zheng, Jinzhi, et al.
Veröffentlicht: (2024)
Generalization-Enhanced Few-Shot Object Detection in Remote Sensing
von: Lin, Hui, et al.
Veröffentlicht: (2025)
von: Lin, Hui, et al.
Veröffentlicht: (2025)
SuperEdit: Rectifying and Facilitating Supervision for Instruction-Based Image Editing
von: Li, Ming, et al.
Veröffentlicht: (2025)
von: Li, Ming, et al.
Veröffentlicht: (2025)
High-Fidelity Image Inpainting with Multimodal Guided GAN Inversion
von: Zhang, Libo, et al.
Veröffentlicht: (2025)
von: Zhang, Libo, et al.
Veröffentlicht: (2025)
OmniSTVG: Toward Spatio-Temporal Omni-Object Video Grounding
von: Yao, Jiali, et al.
Veröffentlicht: (2025)
von: Yao, Jiali, et al.
Veröffentlicht: (2025)
Online Generic Event Boundary Detection
von: Jung, Hyungrok, et al.
Veröffentlicht: (2025)
von: Jung, Hyungrok, et al.
Veröffentlicht: (2025)
LaMOT: Language-Guided Multi-Object Tracking
von: Li, Yunhao, et al.
Veröffentlicht: (2024)
von: Li, Yunhao, et al.
Veröffentlicht: (2024)
G3CN: Gaussian Topology Refinement Gated Graph Convolutional Network for Skeleton-Based Action Recognition
von: Ren, Haiqing, et al.
Veröffentlicht: (2025)
von: Ren, Haiqing, et al.
Veröffentlicht: (2025)
CreatiParser: Generative Image Parsing of Raster Graphic Designs into Editable Layers
von: Chen, Weidong, et al.
Veröffentlicht: (2026)
von: Chen, Weidong, et al.
Veröffentlicht: (2026)
GSOT3D: Towards Generic 3D Single Object Tracking in the Wild
von: Jiao, Yifan, et al.
Veröffentlicht: (2024)
von: Jiao, Yifan, et al.
Veröffentlicht: (2024)
Generic Event Boundary Detection via Denoising Diffusion
von: Hwang, Jaejun, et al.
Veröffentlicht: (2025)
von: Hwang, Jaejun, et al.
Veröffentlicht: (2025)
DreamPoster: A Unified Framework for Image-Conditioned Generative Poster Design
von: Hu, Xiwei, et al.
Veröffentlicht: (2025)
von: Hu, Xiwei, et al.
Veröffentlicht: (2025)
CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation
von: Zhang, Hui, et al.
Veröffentlicht: (2024)
von: Zhang, Hui, et al.
Veröffentlicht: (2024)
What's in the Flow? Exploiting Temporal Motion Cues for Unsupervised Generic Event Boundary Detection
von: Gothe, Sourabh Vasant, et al.
Veröffentlicht: (2024)
von: Gothe, Sourabh Vasant, et al.
Veröffentlicht: (2024)
VisRefiner: Learning from Visual Differences for Screenshot-to-Code Generation
von: Deng, Jie, et al.
Veröffentlicht: (2026)
von: Deng, Jie, et al.
Veröffentlicht: (2026)
Real-time Stereo-based 3D Object Detection for Streaming Perception
von: Li, Changcai, et al.
Veröffentlicht: (2024)
von: Li, Changcai, et al.
Veröffentlicht: (2024)
A Dual-Branch Framework for Semantic Change Detection with Boundary and Temporal Awareness
von: Li, Yun-Cheng, et al.
Veröffentlicht: (2026)
von: Li, Yun-Cheng, et al.
Veröffentlicht: (2026)
Attention to Trajectory: Trajectory-Aware Open-Vocabulary Tracking
von: Li, Yunhao, et al.
Veröffentlicht: (2025)
von: Li, Yunhao, et al.
Veröffentlicht: (2025)
CreatiPoster: Towards Editable and Controllable Multi-Layer Graphic Design Generation
von: Zhang, Zhao, et al.
Veröffentlicht: (2025)
von: Zhang, Zhao, et al.
Veröffentlicht: (2025)
Leveraging Satellite Image Time Series for Accurate Extreme Event Detection
von: Fang, Heng, et al.
Veröffentlicht: (2025)
von: Fang, Heng, et al.
Veröffentlicht: (2025)
PlanarTrack: A high-quality and challenging benchmark for large-scale planar object tracking
von: Jiao, Yifan, et al.
Veröffentlicht: (2025)
von: Jiao, Yifan, et al.
Veröffentlicht: (2025)
FedRSClip: Federated Learning for Remote Sensing Scene Classification Using Vision-Language Models
von: Lin, Hui, et al.
Veröffentlicht: (2025)
von: Lin, Hui, et al.
Veröffentlicht: (2025)
Learning Dynamic Structural Specialization for Underwater Salient Object Detection
von: Hong, Lin, et al.
Veröffentlicht: (2026)
von: Hong, Lin, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Context-Guided Spatio-Temporal Video Grounding
von: Gu, Xin, et al.
Veröffentlicht: (2024) -
Edit3K: Universal Representation Learning for Video Editing Components
von: Gu, Xin, et al.
Veröffentlicht: (2024) -
Multi-Reward as Condition for Instruction-based Image Editing
von: Gu, Xin, et al.
Veröffentlicht: (2024) -
Robust Domain Adaptive Object Detection with Unified Multi-Granularity Alignment
von: Zhang, Libo, et al.
Veröffentlicht: (2023) -
Accurate and Fast Compressed Video Captioning
von: Shen, Yaojie, et al.
Veröffentlicht: (2023)