UBiSS: A Unified Framework for Bimodal Semantic Summarization of Videos
Fuente:
arXiv
Guardado en:
| Autores principales: | Mei, Yuting, Yao, Linli, Jin, Qin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Unsupervised Transcript-assisted Video Summarization and Highlight Detection
por: Barbakos, Spyros, et al.
Publicado: (2025)
por: Barbakos, Spyros, et al.
Publicado: (2025)
VIA: Unified Spatiotemporal Video Adaptation Framework for Global and Local Video Editing
por: Gu, Jing, et al.
Publicado: (2024)
por: Gu, Jing, et al.
Publicado: (2024)
SoccerHigh: A Benchmark Dataset for Automatic Soccer Video Summarization
por: Díaz-Juan, Artur, et al.
Publicado: (2025)
por: Díaz-Juan, Artur, et al.
Publicado: (2025)
SD-VSum: A Method and Dataset for Script-Driven Video Summarization
por: Mylonas, Manolis, et al.
Publicado: (2025)
por: Mylonas, Manolis, et al.
Publicado: (2025)
Question-Answering Dense Video Events
por: Qin, Hangyu, et al.
Publicado: (2024)
por: Qin, Hangyu, et al.
Publicado: (2024)
Video Summarization: Towards Entity-Aware Captions
por: Ayyubi, Hammad A., et al.
Publicado: (2023)
por: Ayyubi, Hammad A., et al.
Publicado: (2023)
Bernini: Latent Semantic Planning for Video Diffusion
por: Bernini Team, et al.
Publicado: (2026)
por: Bernini Team, et al.
Publicado: (2026)
EgoEsportsQA: An Egocentric Video Benchmark for Perception and Reasoning in Esports
por: Ma, Jianzhe, et al.
Publicado: (2026)
por: Ma, Jianzhe, et al.
Publicado: (2026)
AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks
por: Ku, Max, et al.
Publicado: (2024)
por: Ku, Max, et al.
Publicado: (2024)
Reinforcement Learning for Unsupervised Video Summarization with Reward Generator Training
por: Abbasi, Mehryar, et al.
Publicado: (2024)
por: Abbasi, Mehryar, et al.
Publicado: (2024)
LayerT2V: A Unified Multi-Layer Video Generation Framework
por: Li, Guangzhao, et al.
Publicado: (2025)
por: Li, Guangzhao, et al.
Publicado: (2025)
Beyond Audio and Pose: A General-Purpose Framework for Video Synchronization
por: Shin, Yosub, et al.
Publicado: (2025)
por: Shin, Yosub, et al.
Publicado: (2025)
Lumos-1: On Autoregressive Video Generation with Discrete Diffusion from a Unified Model Perspective
por: Yuan, Hangjie, et al.
Publicado: (2025)
por: Yuan, Hangjie, et al.
Publicado: (2025)
SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment
por: Mao, Xinyu, et al.
Publicado: (2025)
por: Mao, Xinyu, et al.
Publicado: (2025)
Q-Ponder: A Unified Training Pipeline for Reasoning-based Visual Quality Assessment
por: Cai, Zhuoxuan, et al.
Publicado: (2025)
por: Cai, Zhuoxuan, et al.
Publicado: (2025)
Apollo: Unified Multi-Task Audio-Video Joint Generation
por: Wang, Jun, et al.
Publicado: (2026)
por: Wang, Jun, et al.
Publicado: (2026)
InstructVid2Vid: Controllable Video Editing with Natural Language Instructions
por: Qin, Bosheng, et al.
Publicado: (2023)
por: Qin, Bosheng, et al.
Publicado: (2023)
Can I Trust Your Answer? Visually Grounded Video Question Answering
por: Xiao, Junbin, et al.
Publicado: (2023)
por: Xiao, Junbin, et al.
Publicado: (2023)
AToken: A Unified Tokenizer for Vision
por: Lu, Jiasen, et al.
Publicado: (2025)
por: Lu, Jiasen, et al.
Publicado: (2025)
Edit As You Wish: Video Caption Editing with Multi-grained User Control
por: Yao, Linli, et al.
Publicado: (2023)
por: Yao, Linli, et al.
Publicado: (2023)
Using AI to Summarize US Presidential Campaign TV Advertisement Videos, 1952-2012
por: Breuer, Adam, et al.
Publicado: (2025)
por: Breuer, Adam, et al.
Publicado: (2025)
Multi-modal Segment Assemblage Network for Ad Video Editing with Importance-Coherence Reward
por: Tang, Yolo Yunlong, et al.
Publicado: (2022)
por: Tang, Yolo Yunlong, et al.
Publicado: (2022)
Grounding is All You Need? Dual Temporal Grounding for Video Dialog
por: Qin, You, et al.
Publicado: (2024)
por: Qin, You, et al.
Publicado: (2024)
From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents
por: Lian, Niu, et al.
Publicado: (2026)
por: Lian, Niu, et al.
Publicado: (2026)
MuKV: Multi-Grained KV Cache Compression for Long Streaming Video Question-Answering
por: Xiao, Junbin, et al.
Publicado: (2026)
por: Xiao, Junbin, et al.
Publicado: (2026)
Moiré Video Authentication: A Physical Signature Against AI Video Generation
por: Qing, Yuan, et al.
Publicado: (2026)
por: Qing, Yuan, et al.
Publicado: (2026)
Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
por: Liu, Jiajun, et al.
Publicado: (2024)
por: Liu, Jiajun, et al.
Publicado: (2024)
Video Seal: Open and Efficient Video Watermarking
por: Fernandez, Pierre, et al.
Publicado: (2024)
por: Fernandez, Pierre, et al.
Publicado: (2024)
FedVideoMAE: Efficient Privacy-Preserving Federated Video Moderation
por: Tao, Ziyuan, et al.
Publicado: (2025)
por: Tao, Ziyuan, et al.
Publicado: (2025)
MotionCtrl: A Unified and Flexible Motion Controller for Video Generation
por: Wang, Zhouxia, et al.
Publicado: (2023)
por: Wang, Zhouxia, et al.
Publicado: (2023)
Crab$^{+}$: A Scalable and Unified Audio-Visual Scene Understanding Model with Explicit Cooperation
por: Cai, Dongnuan, et al.
Publicado: (2026)
por: Cai, Dongnuan, et al.
Publicado: (2026)
RacketVision: A Multiple Racket Sports Benchmark for Unified Ball and Racket Analysis
por: Dong, Linfeng, et al.
Publicado: (2025)
por: Dong, Linfeng, et al.
Publicado: (2025)
Video-EM: Event-Centric Episodic Memory for Long-Form Video Understanding
por: Wang, Yun, et al.
Publicado: (2025)
por: Wang, Yun, et al.
Publicado: (2025)
Enhancing Self-Supervised Talking Head Forgery Detection via a Training-Free Dual-System Framework
por: Liu, Ke, et al.
Publicado: (2026)
por: Liu, Ke, et al.
Publicado: (2026)
Open-o3-Video: Grounded Video Reasoning with Explicit Spatio-Temporal Evidence
por: Meng, Jiahao, et al.
Publicado: (2025)
por: Meng, Jiahao, et al.
Publicado: (2025)
VideoPainter: Any-length Video Inpainting and Editing with Plug-and-Play Context Control
por: Bian, Yuxuan, et al.
Publicado: (2025)
por: Bian, Yuxuan, et al.
Publicado: (2025)
XEmoGPT: An Explainable Multimodal Emotion Recognition Framework with Cue-Level Perception and Reasoning
por: Zhang, Hanwen, et al.
Publicado: (2026)
por: Zhang, Hanwen, et al.
Publicado: (2026)
Long Video Diffusion Generation with Segmented Cross-Attention and Content-Rich Video Data Curation
por: Yan, Xin, et al.
Publicado: (2024)
por: Yan, Xin, et al.
Publicado: (2024)
DIBS: Enhancing Dense Video Captioning with Unlabeled Videos via Pseudo Boundary Enrichment and Online Refinement
por: Wu, Hao, et al.
Publicado: (2024)
por: Wu, Hao, et al.
Publicado: (2024)
LoViF 2026 The First Challenge on Weather Removal in Videos
por: Qian, Chenghao, et al.
Publicado: (2026)
por: Qian, Chenghao, et al.
Publicado: (2026)
Ejemplares similares
-
Unsupervised Transcript-assisted Video Summarization and Highlight Detection
por: Barbakos, Spyros, et al.
Publicado: (2025) -
VIA: Unified Spatiotemporal Video Adaptation Framework for Global and Local Video Editing
por: Gu, Jing, et al.
Publicado: (2024) -
SoccerHigh: A Benchmark Dataset for Automatic Soccer Video Summarization
por: Díaz-Juan, Artur, et al.
Publicado: (2025) -
SD-VSum: A Method and Dataset for Script-Driven Video Summarization
por: Mylonas, Manolis, et al.
Publicado: (2025) -
Question-Answering Dense Video Events
por: Qin, Hangyu, et al.
Publicado: (2024)