Language-Guided Graph Representation Learning for Video Summarization
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Wenrui, Han, Wei, Man, Hengyu, Zuo, Wangmeng, Fan, Xiaopeng, Tian, Yonghong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Spiking Variational Graph Representation Inference for Video Summarization
by: Li, Wenrui, et al.
Published: (2025)
by: Li, Wenrui, et al.
Published: (2025)
Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning
by: Li, Wenrui, et al.
Published: (2025)
by: Li, Wenrui, et al.
Published: (2025)
All-in-One Video Restoration under Smoothly Evolving Unknown Weather Degradations
by: Li, Wenrui, et al.
Published: (2026)
by: Li, Wenrui, et al.
Published: (2026)
T-GVC: Trajectory-Guided Generative Video Coding at Ultra-Low Bitrates
by: Wang, Zhitao, et al.
Published: (2025)
by: Wang, Zhitao, et al.
Published: (2025)
Hyperbolic-constraint Point Cloud Reconstruction from Single RGB-D Images
by: Li, Wenrui, et al.
Published: (2024)
by: Li, Wenrui, et al.
Published: (2024)
MRT: Learning Compact Representations with Mixed RWKV-Transformer for Extreme Image Compression
by: Liu, Han, et al.
Published: (2025)
by: Liu, Han, et al.
Published: (2025)
RoamScene3D: Immersive Text-to-3D Scene Generation via Adaptive Object-aware Roaming
by: Chu, Jisheng, et al.
Published: (2026)
by: Chu, Jisheng, et al.
Published: (2026)
Hyperbolic Hierarchical Alignment Reasoning Network for Text-3D Retrieval
by: Li, Wenrui, et al.
Published: (2025)
by: Li, Wenrui, et al.
Published: (2025)
SceneDreamer360: Text-Driven 3D-Consistent Scene Generation with Panoramic Gaussian Splatting
by: Li, Wenrui, et al.
Published: (2024)
by: Li, Wenrui, et al.
Published: (2024)
SpikeMba: Multi-Modal Spiking Saliency Mamba for Temporal Video Grounding
by: Li, Wenrui, et al.
Published: (2024)
by: Li, Wenrui, et al.
Published: (2024)
MTLSI-Net: A Linear Semantic Interaction Network for Parameter-Efficient Multi-Task Dense Prediction
by: Liu, Chen, et al.
Published: (2026)
by: Liu, Chen, et al.
Published: (2026)
Cluster-Wise Spatio-Temporal Masking for Efficient Video-Language Pretraining
by: Zhuang, Weijun, et al.
Published: (2026)
by: Zhuang, Weijun, et al.
Published: (2026)
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection
by: Zhuang, Weijun, et al.
Published: (2025)
by: Zhuang, Weijun, et al.
Published: (2025)
VideoSAGE: Video Summarization with Graph Representation Learning
by: Chaves, Jose M. Rojas, et al.
Published: (2024)
by: Chaves, Jose M. Rojas, et al.
Published: (2024)
Lie Flow: Video Dynamic Fields Modeling and Predicting with Lie Algebra as Geometric Physics Principle
by: Qiao, Weidong, et al.
Published: (2026)
by: Qiao, Weidong, et al.
Published: (2026)
SplatWeaver: Learning to Allocate Gaussian Primitives for Generalizable Novel View Synthesis
by: Wan, Yecong, et al.
Published: (2026)
by: Wan, Yecong, et al.
Published: (2026)
SelfHVD: Self-Supervised Handheld Video Deblurring
by: Xu, Honglei, et al.
Published: (2025)
by: Xu, Honglei, et al.
Published: (2025)
Multi-modal Crowd Counting via a Broker Modality
by: Meng, Haoliang, et al.
Published: (2024)
by: Meng, Haoliang, et al.
Published: (2024)
Reblurring-Guided Single Image Defocus Deblurring: A Learning Framework with Misaligned Training Pairs
by: Ren, Dongwei, et al.
Published: (2024)
by: Ren, Dongwei, et al.
Published: (2024)
Learning Spatially Decoupled Color Representations for Facial Image Colorization
by: Zhu, Hangyan, et al.
Published: (2024)
by: Zhu, Hangyan, et al.
Published: (2024)
HDI-Former: Hybrid Dynamic Interaction ANN-SNN Transformer for Object Detection Using Frames and Events
by: Li, Dianze, et al.
Published: (2024)
by: Li, Dianze, et al.
Published: (2024)
Dynamic Graph Representation with Knowledge-aware Attention for Histopathology Whole Slide Image Analysis
by: Li, Jiawen, et al.
Published: (2024)
by: Li, Jiawen, et al.
Published: (2024)
Prompts to Summaries: Zero-Shot Language-Guided Video Summarization with Large Language and Video Models
by: Barbara, Mario, et al.
Published: (2025)
by: Barbara, Mario, et al.
Published: (2025)
Improving Image Restoration through Removing Degradations in Textual Representations
by: Lin, Jingbo, et al.
Published: (2023)
by: Lin, Jingbo, et al.
Published: (2023)
Perceptual Quality Assessment of 3D Gaussian Splatting: A Subjective Dataset and Prediction Metric
by: Wan, Zhaolin, et al.
Published: (2025)
by: Wan, Zhaolin, et al.
Published: (2025)
Visual-O1: Understanding Ambiguous Instructions via Multi-modal Multi-turn Chain-of-thoughts Reasoning
by: Ni, Minheng, et al.
Published: (2024)
by: Ni, Minheng, et al.
Published: (2024)
Hyperbolic Distillation: Geometry-Guided Cross-Modal Transfer for Robust 3D Object Detection
by: Ning, Kanglin, et al.
Published: (2026)
by: Ning, Kanglin, et al.
Published: (2026)
Aggregating Nearest Sharp Features via Hybrid Transformers for Video Deblurring
by: Shang, Wei, et al.
Published: (2023)
by: Shang, Wei, et al.
Published: (2023)
Deformable Attention Graph Representation Learning for Histopathology Whole Slide Image Analysis
by: Fu, Mingxi, et al.
Published: (2025)
by: Fu, Mingxi, et al.
Published: (2025)
ACE: Anti-Editing Concept Erasure in Text-to-Image Models
by: Wang, Zihao, et al.
Published: (2025)
by: Wang, Zihao, et al.
Published: (2025)
Generative Inbetweening through Frame-wise Conditions-Driven Video Generation
by: Zhu, Tianyi, et al.
Published: (2024)
by: Zhu, Tianyi, et al.
Published: (2024)
ScrollScape: Unlocking 32K Image Generation With Video Diffusion Priors
by: Yu, Haodong, et al.
Published: (2026)
by: Yu, Haodong, et al.
Published: (2026)
FramePainter: Endowing Interactive Image Editing with Video Diffusion Priors
by: Zhang, Yabo, et al.
Published: (2025)
by: Zhang, Yabo, et al.
Published: (2025)
Enhanced Generative Structure Prior for Chinese Text Image Super-resolution
by: Li, Xiaoming, et al.
Published: (2025)
by: Li, Xiaoming, et al.
Published: (2025)
Deblur4DGS: 4D Gaussian Splatting from Blurry Monocular Video
by: Wu, Renlong, et al.
Published: (2024)
by: Wu, Renlong, et al.
Published: (2024)
Thin-Plate Spline-based Interpolation for Animation Line Inbetweening
by: Zhu, Tianyi, et al.
Published: (2024)
by: Zhu, Tianyi, et al.
Published: (2024)
Language-Inspired Relation Transfer for Few-shot Class-Incremental Learning
by: Zhao, Yifan, et al.
Published: (2025)
by: Zhao, Yifan, et al.
Published: (2025)
MetricDepth: Enhancing Monocular Depth Estimation with Deep Metric Learning
by: Liu, Chunpu, et al.
Published: (2024)
by: Liu, Chunpu, et al.
Published: (2024)
Riemann-based Multi-scale Attention Reasoning Network for Text-3D Retrieval
by: Li, Wenrui, et al.
Published: (2024)
by: Li, Wenrui, et al.
Published: (2024)
LoViC: Efficient Long Video Generation with Context Compression
by: Jiang, Jiaxiu, et al.
Published: (2025)
by: Jiang, Jiaxiu, et al.
Published: (2025)
Similar Items
-
Spiking Variational Graph Representation Inference for Video Summarization
by: Li, Wenrui, et al.
Published: (2025) -
Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning
by: Li, Wenrui, et al.
Published: (2025) -
All-in-One Video Restoration under Smoothly Evolving Unknown Weather Degradations
by: Li, Wenrui, et al.
Published: (2026) -
T-GVC: Trajectory-Guided Generative Video Coding at Ultra-Low Bitrates
by: Wang, Zhitao, et al.
Published: (2025) -
Hyperbolic-constraint Point Cloud Reconstruction from Single RGB-D Images
by: Li, Wenrui, et al.
Published: (2024)