SEAL: Semantic Attention Learning for Long Video Representation
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Lan, Chen, Yujia, Tran, Du, Boddeti, Vishnu Naresh, Chu, Wen-Sheng |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FairerCLIP: Debiasing CLIP's Zero-Shot Predictions using Functions in RKHSs
by: Dehdashtian, Sepehr, et al.
Published: (2024)
by: Dehdashtian, Sepehr, et al.
Published: (2024)
Action Reimagined: Text-to-Pose Video Editing for Dynamic Human Actions
by: Wang, Lan, et al.
Published: (2024)
by: Wang, Lan, et al.
Published: (2024)
CryptoFace: End-to-End Encrypted Face Recognition
by: Ao, Wei, et al.
Published: (2025)
by: Ao, Wei, et al.
Published: (2025)
Utility-Fairness Trade-Offs and How to Find Them
by: Dehdashtian, Sepehr, et al.
Published: (2024)
by: Dehdashtian, Sepehr, et al.
Published: (2024)
OASIS Uncovers: High-Quality T2I Models, Same Old Stereotypes
by: Dehdashtian, Sepehr, et al.
Published: (2025)
by: Dehdashtian, Sepehr, et al.
Published: (2025)
Show Me: Unifying Instructional Image and Video Generation with Diffusion Models
by: Pu, Yujiang, et al.
Published: (2025)
by: Pu, Yujiang, et al.
Published: (2025)
Layer-Aware Video Composition via Split-then-Merge
by: Kara, Ozgur, et al.
Published: (2025)
by: Kara, Ozgur, et al.
Published: (2025)
Fairness and Bias Mitigation in Computer Vision: A Survey
by: Dehdashtian, Sepehr, et al.
Published: (2024)
by: Dehdashtian, Sepehr, et al.
Published: (2024)
SEAL: Semantic-Aware Hierarchical Learning for Generalized Category Discovery
by: He, Zhenqi, et al.
Published: (2025)
by: He, Zhenqi, et al.
Published: (2025)
MM-SEAL: A Large-scale Video Dataset of Multi-person Multi-grained Spatio-temporally Action Localization
by: Chen, Shimin, et al.
Published: (2022)
by: Chen, Shimin, et al.
Published: (2022)
VideoSEAL: Mitigating Evidence Misalignment in Agentic Long Video Understanding by Decoupling Answer Authority
by: Qiu, Chenhao, et al.
Published: (2026)
by: Qiu, Chenhao, et al.
Published: (2026)
Semantic Image Inversion and Editing using Rectified Stochastic Differential Equations
by: Rout, Litu, et al.
Published: (2024)
by: Rout, Litu, et al.
Published: (2024)
A Deep Learning Framework for Three Dimensional Shape Reconstruction from Phaseless Acoustic Scattering Far-field Data
by: Dikbayir, Doga, et al.
Published: (2024)
by: Dikbayir, Doga, et al.
Published: (2024)
SEAL: Semantic Aware Image Watermarking
by: Arabi, Kasra, et al.
Published: (2025)
by: Arabi, Kasra, et al.
Published: (2025)
SeViCES: Unifying Semantic-Visual Evidence Consensus for Long Video Understanding
by: Sheng, Yuan, et al.
Published: (2025)
by: Sheng, Yuan, et al.
Published: (2025)
Seeing Eye to AI: Comparing Human Gaze and Model Attention in Video Memorability
by: Kumar, Prajneya, et al.
Published: (2023)
by: Kumar, Prajneya, et al.
Published: (2023)
Enhancing Privacy in Face Analytics Using Fully Homomorphic Encryption
by: Yalavarthi, Bharat, et al.
Published: (2024)
by: Yalavarthi, Bharat, et al.
Published: (2024)
Semantic Flow: Learning Semantic Field of Dynamic Scenes from Monocular Videos
by: Tian, Fengrui, et al.
Published: (2024)
by: Tian, Fengrui, et al.
Published: (2024)
Reframing Long-Tailed Learning via Loss Landscape Geometry
by: Chen, Shenghan, et al.
Published: (2026)
by: Chen, Shenghan, et al.
Published: (2026)
Learning Audio-guided Video Representation with Gated Attention for Video-Text Retrieval
by: Jeong, Boseung, et al.
Published: (2025)
by: Jeong, Boseung, et al.
Published: (2025)
SEAL: Semantic-aware Single-image Sticker Personalization with a Large-scale Sticker-tag Dataset
by: Roh, Changhyun, et al.
Published: (2026)
by: Roh, Changhyun, et al.
Published: (2026)
Learning Compact Video Representations for Efficient Long-form Video Understanding in Large Multimodal Models
by: Chen, Yuxiao, et al.
Published: (2026)
by: Chen, Yuxiao, et al.
Published: (2026)
SAVE: Speech-Aware Video Representation Learning for Video-Text Retrieval
by: Zhao, Ruixiang, et al.
Published: (2026)
by: Zhao, Ruixiang, et al.
Published: (2026)
Time2General: Learning Spatiotemporal Invariant Representations for Domain-Generalization Video Semantic Segmentation
by: Chen, Siyu, et al.
Published: (2026)
by: Chen, Siyu, et al.
Published: (2026)
Long Video Understanding with Learnable Retrieval in Video-Language Models
by: Xu, Jiaqi, et al.
Published: (2023)
by: Xu, Jiaqi, et al.
Published: (2023)
SEAL: A Framework for Systematic Evaluation of Real-World Super-Resolution
by: Zhang, Wenlong, et al.
Published: (2023)
by: Zhang, Wenlong, et al.
Published: (2023)
Improving Human Image Animation via Semantic Representation Alignment
by: Liu, Chang, et al.
Published: (2026)
by: Liu, Chang, et al.
Published: (2026)
Multi-entity Video Transformers for Fine-Grained Video Representation Learning
by: Walmer, Matthew, et al.
Published: (2023)
by: Walmer, Matthew, et al.
Published: (2023)
Edit3K: Universal Representation Learning for Video Editing Components
by: Gu, Xin, et al.
Published: (2024)
by: Gu, Xin, et al.
Published: (2024)
Long-Horizon Streaming Video Generation via Hybrid Attention with Decoupled Distillation
by: Li, Ruibin, et al.
Published: (2026)
by: Li, Ruibin, et al.
Published: (2026)
Learning Video Representations without Natural Videos
by: Yu, Xueyang, et al.
Published: (2024)
by: Yu, Xueyang, et al.
Published: (2024)
Latent Diffusion Unlearning: Protecting Against Unauthorized Personalization Through Trajectory Shifted Perturbations
by: Devulapally, Naresh Kumar, et al.
Published: (2025)
by: Devulapally, Naresh Kumar, et al.
Published: (2025)
Learnable Instance Attention Filtering for Adaptive Detector Distillation
by: Liu, Chen, et al.
Published: (2026)
by: Liu, Chen, et al.
Published: (2026)
Unlocking Exocentric Video-Language Data for Egocentric Video Representation Learning
by: Dou, Zi-Yi, et al.
Published: (2024)
by: Dou, Zi-Yi, et al.
Published: (2024)
Beyond Prototypes: Semantic Anchor Regularization for Better Representation Learning
by: Ge, Yanqi, et al.
Published: (2023)
by: Ge, Yanqi, et al.
Published: (2023)
Multi-Scale Representations by Varying Window Attention for Semantic Segmentation
by: Yan, Haotian, et al.
Published: (2024)
by: Yan, Haotian, et al.
Published: (2024)
LongLive: Real-time Interactive Long Video Generation
by: Yang, Shuai, et al.
Published: (2025)
by: Yang, Shuai, et al.
Published: (2025)
MoGA: Mixture-of-Groups Attention for End-to-End Long Video Generation
by: Jia, Weinan, et al.
Published: (2025)
by: Jia, Weinan, et al.
Published: (2025)
AlignVid: Training-Free Attention Scaling for Semantic Fidelity in Text-Guided Image-to-Video Generation
by: Liu, Yexin, et al.
Published: (2025)
by: Liu, Yexin, et al.
Published: (2025)
Wavelet-based Frame Selection by Detecting Semantic Boundary for Long Video Understanding
by: Chen, Wang, et al.
Published: (2026)
by: Chen, Wang, et al.
Published: (2026)
Similar Items
-
FairerCLIP: Debiasing CLIP's Zero-Shot Predictions using Functions in RKHSs
by: Dehdashtian, Sepehr, et al.
Published: (2024) -
Action Reimagined: Text-to-Pose Video Editing for Dynamic Human Actions
by: Wang, Lan, et al.
Published: (2024) -
CryptoFace: End-to-End Encrypted Face Recognition
by: Ao, Wei, et al.
Published: (2025) -
Utility-Fairness Trade-Offs and How to Find Them
by: Dehdashtian, Sepehr, et al.
Published: (2024) -
OASIS Uncovers: High-Quality T2I Models, Same Old Stereotypes
by: Dehdashtian, Sepehr, et al.
Published: (2025)