A Multimodal Transformer for Live Streaming Highlight Prediction
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Deng, Jiaxin, Wang, Shiyao, Shen, Dong, Zhao, Liqin, Yang, Fan, Zhou, Guorui, Meng, Gaofeng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TR-DETR: Task-Reciprocal Transformer for Joint Moment Retrieval and Highlight Detection
von: Sun, Hao, et al.
Veröffentlicht: (2024)
von: Sun, Hao, et al.
Veröffentlicht: (2024)
High-Quality Live Video Streaming via Transcoding Time Prediction and Preset Selection
von: Shahre-Babak, Zahra Nabizadeh, et al.
Veröffentlicht: (2023)
von: Shahre-Babak, Zahra Nabizadeh, et al.
Veröffentlicht: (2023)
Embedded Heterogeneous Attention Transformer for Cross-lingual Image Captioning
von: Song, Zijie, et al.
Veröffentlicht: (2023)
von: Song, Zijie, et al.
Veröffentlicht: (2023)
Viewport Prediction for Volumetric Video Streaming by Exploring Video Saliency and Trajectory Information
von: Li, Jie, et al.
Veröffentlicht: (2023)
von: Li, Jie, et al.
Veröffentlicht: (2023)
Multimodal Class-aware Semantic Enhancement Network for Audio-Visual Video Parsing
von: Zhao, Pengcheng, et al.
Veröffentlicht: (2024)
von: Zhao, Pengcheng, et al.
Veröffentlicht: (2024)
Hallo-Live: Real-Time Streaming Joint Audio-Video Avatar Generation with Asynchronous Dual-Stream and Human-Centric Preference Distillation
von: Li, Chunyu, et al.
Veröffentlicht: (2026)
von: Li, Chunyu, et al.
Veröffentlicht: (2026)
Learning Long-Range Action Representation by Two-Stream Mamba Pyramid Network for Figure Skating Assessment
von: Wang, Fengshun, et al.
Veröffentlicht: (2025)
von: Wang, Fengshun, et al.
Veröffentlicht: (2025)
SMSP: A Plug-and-Play Strategy of Multi-Scale Perception for MLLMs to Perceive Visual Illusions
von: Tu, Jinzhe, et al.
Veröffentlicht: (2026)
von: Tu, Jinzhe, et al.
Veröffentlicht: (2026)
StreamingEval: A Unified Evaluation Protocol towards Realistic Streaming Video Understanding
von: Tang, Guowei, et al.
Veröffentlicht: (2026)
von: Tang, Guowei, et al.
Veröffentlicht: (2026)
Detached and Interactive Multimodal Learning
von: Fan, Yunfeng, et al.
Veröffentlicht: (2024)
von: Fan, Yunfeng, et al.
Veröffentlicht: (2024)
Adaptive 3D Gaussian Splatting Video Streaming
von: Gong, Han, et al.
Veröffentlicht: (2025)
von: Gong, Han, et al.
Veröffentlicht: (2025)
Advancing Unsupervised Low-light Image Enhancement: Noise Estimation, Illumination Interpolation, and Self-Regulation
von: Liu, Xiaofeng, et al.
Veröffentlicht: (2023)
von: Liu, Xiaofeng, et al.
Veröffentlicht: (2023)
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval
von: Kong, Fanheng, et al.
Veröffentlicht: (2025)
von: Kong, Fanheng, et al.
Veröffentlicht: (2025)
PLayerTV: Advanced Player Tracking and Identification for Automatic Soccer Highlight Clips
von: Solberg, Håkon Maric, et al.
Veröffentlicht: (2024)
von: Solberg, Håkon Maric, et al.
Veröffentlicht: (2024)
ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
von: Zhang, Zhenxing, et al.
Veröffentlicht: (2024)
von: Zhang, Zhenxing, et al.
Veröffentlicht: (2024)
OralGPT-Omni: A Versatile Dental Multimodal Large Language Model
von: Hao, Jing, et al.
Veröffentlicht: (2025)
von: Hao, Jing, et al.
Veröffentlicht: (2025)
Scaling up Multimodal Pre-training for Sign Language Understanding
von: Zhou, Wengang, et al.
Veröffentlicht: (2024)
von: Zhou, Wengang, et al.
Veröffentlicht: (2024)
Towards Robust and Realible Multimodal Misinformation Recognition with Incomplete Modality
von: Zhou, Hengyang, et al.
Veröffentlicht: (2025)
von: Zhou, Hengyang, et al.
Veröffentlicht: (2025)
Efficient Token Compression for Vision Transformer with Spatial Information Preserved
von: Mao, Junzhu, et al.
Veröffentlicht: (2025)
von: Mao, Junzhu, et al.
Veröffentlicht: (2025)
Multi-scale Bottleneck Transformer for Weakly Supervised Multimodal Violence Detection
von: Sun, Shengyang, et al.
Veröffentlicht: (2024)
von: Sun, Shengyang, et al.
Veröffentlicht: (2024)
Tile Classification Based Viewport Prediction with Multi-modal Fusion Transformer
von: Zhang, Zhihao, et al.
Veröffentlicht: (2023)
von: Zhang, Zhihao, et al.
Veröffentlicht: (2023)
Other Tokens Matter: Exploring Global and Local Features of Vision Transformers for Object Re-Identification
von: Wang, Yingquan, et al.
Veröffentlicht: (2024)
von: Wang, Yingquan, et al.
Veröffentlicht: (2024)
MHAD: Multimodal Home Activity Dataset with Multi-Angle Videos and Synchronized Physiological Signals
von: Yu, Lei, et al.
Veröffentlicht: (2024)
von: Yu, Lei, et al.
Veröffentlicht: (2024)
SSNVC: Single Stream Neural Video Compression with Implicit Temporal Information
von: Wang, Feng, et al.
Veröffentlicht: (2024)
von: Wang, Feng, et al.
Veröffentlicht: (2024)
SpecFLASH: A Latent-Guided Semi-autoregressive Speculative Decoding Framework for Efficient Multimodal Generation
von: Wang, Zihua, et al.
Veröffentlicht: (2025)
von: Wang, Zihua, et al.
Veröffentlicht: (2025)
DRFormer: A Dual-Regularized Bidirectional Transformer for Person Re-identification
von: Shu, Ying, et al.
Veröffentlicht: (2026)
von: Shu, Ying, et al.
Veröffentlicht: (2026)
TraveLLaMA: A Multimodal Travel Assistant with Large-Scale Dataset and Structured Reasoning
von: Chu, Meng, et al.
Veröffentlicht: (2025)
von: Chu, Meng, et al.
Veröffentlicht: (2025)
Grounded Chain-of-Thought for Multimodal Large Language Models
von: Wu, Qiong, et al.
Veröffentlicht: (2025)
von: Wu, Qiong, et al.
Veröffentlicht: (2025)
CL2CM: Improving Cross-Lingual Cross-Modal Retrieval via Cross-Lingual Knowledge Transfer
von: Wang, Yabing, et al.
Veröffentlicht: (2023)
von: Wang, Yabing, et al.
Veröffentlicht: (2023)
PRVR: Partially Relevant Video Retrieval
von: Chen, Xianke, et al.
Veröffentlicht: (2022)
von: Chen, Xianke, et al.
Veröffentlicht: (2022)
DDNet: A Dual-Stream Graph Learning and Disentanglement Framework for Temporal Forgery Localization
von: Zhao, Boyang, et al.
Veröffentlicht: (2026)
von: Zhao, Boyang, et al.
Veröffentlicht: (2026)
Unify-Agent: A Unified Multimodal Agent for World-Grounded Image Synthesis
von: Chen, Shuang, et al.
Veröffentlicht: (2026)
von: Chen, Shuang, et al.
Veröffentlicht: (2026)
MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces
von: E, Shaojun, et al.
Veröffentlicht: (2025)
von: E, Shaojun, et al.
Veröffentlicht: (2025)
Context Guided Transformer Entropy Modeling for Video Compression
von: Tong, Junlong, et al.
Veröffentlicht: (2025)
von: Tong, Junlong, et al.
Veröffentlicht: (2025)
DPDETR: Decoupled Position Detection Transformer for Infrared-Visible Object Detection
von: Guo, Junjie, et al.
Veröffentlicht: (2024)
von: Guo, Junjie, et al.
Veröffentlicht: (2024)
Transformer-based Video Saliency Prediction with High Temporal Dimension Decoding
von: Moradi, Morteza, et al.
Veröffentlicht: (2024)
von: Moradi, Morteza, et al.
Veröffentlicht: (2024)
HOP: Heterogeneous Topology-based Multimodal Entanglement for Co-Speech Gesture Generation
von: Cheng, Hongye, et al.
Veröffentlicht: (2025)
von: Cheng, Hongye, et al.
Veröffentlicht: (2025)
M2ORT: Many-To-One Regression Transformer for Spatial Transcriptomics Prediction from Histopathology Images
von: Wang, Hongyi, et al.
Veröffentlicht: (2024)
von: Wang, Hongyi, et al.
Veröffentlicht: (2024)
DeepSPG: Exploring Deep Semantic Prior Guidance for Low-light Image Enhancement with Multimodal Learning
von: Lu, Jialang, et al.
Veröffentlicht: (2025)
von: Lu, Jialang, et al.
Veröffentlicht: (2025)
Interpretable Embedding for Ad-hoc Video Search
von: Wu, Jiaxin, et al.
Veröffentlicht: (2024)
von: Wu, Jiaxin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
TR-DETR: Task-Reciprocal Transformer for Joint Moment Retrieval and Highlight Detection
von: Sun, Hao, et al.
Veröffentlicht: (2024) -
High-Quality Live Video Streaming via Transcoding Time Prediction and Preset Selection
von: Shahre-Babak, Zahra Nabizadeh, et al.
Veröffentlicht: (2023) -
Embedded Heterogeneous Attention Transformer for Cross-lingual Image Captioning
von: Song, Zijie, et al.
Veröffentlicht: (2023) -
Viewport Prediction for Volumetric Video Streaming by Exploring Video Saliency and Trajectory Information
von: Li, Jie, et al.
Veröffentlicht: (2023) -
Multimodal Class-aware Semantic Enhancement Network for Audio-Visual Video Parsing
von: Zhao, Pengcheng, et al.
Veröffentlicht: (2024)