Movie Trailer Genre Classification Using Multimodal Pretrained Features
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sulun, Serkan, Viana, Paula, Davies, Matthew E. P. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VEMOCLAP: A video emotion classification web application
von: Sulun, Serkan, et al.
Veröffentlicht: (2024)
von: Sulun, Serkan, et al.
Veröffentlicht: (2024)
Video Soundtrack Generation by Aligning Emotions and Temporal Boundaries
von: Sulun, Serkan, et al.
Veröffentlicht: (2025)
von: Sulun, Serkan, et al.
Veröffentlicht: (2025)
d-Sketch: Improving Visual Fidelity of Sketch-to-Image Translation with Pretrained Latent Diffusion Models without Retraining
von: Roy, Prasun, et al.
Veröffentlicht: (2025)
von: Roy, Prasun, et al.
Veröffentlicht: (2025)
VineetVC: Adaptive Video Conferencing Under Severe Bandwidth Constraints Using Audio-Driven Talking-Head Reconstruction
von: Rakesh, Vineet Kumar, et al.
Veröffentlicht: (2026)
von: Rakesh, Vineet Kumar, et al.
Veröffentlicht: (2026)
DietDelta: A Vision-Language Approach for Dietary Assessment via Before-and-After Images
von: Vinod, Gautham, et al.
Veröffentlicht: (2026)
von: Vinod, Gautham, et al.
Veröffentlicht: (2026)
Towards Real-world Video Face Restoration: A New Benchmark
von: Chen, Ziyan, et al.
Veröffentlicht: (2024)
von: Chen, Ziyan, et al.
Veröffentlicht: (2024)
EVAN: Evolutional Video Streaming Adaptation via Neural Representation
von: Liu, Mufan, et al.
Veröffentlicht: (2024)
von: Liu, Mufan, et al.
Veröffentlicht: (2024)
Machine Perception-Driven Image Compression: A Layered Generative Approach
von: Zhang, Yuefeng, et al.
Veröffentlicht: (2023)
von: Zhang, Yuefeng, et al.
Veröffentlicht: (2023)
Change Detection Between Optical Remote Sensing Imagery and Map Data via Segment Anything Model (SAM)
von: Chen, Hongruixuan, et al.
Veröffentlicht: (2024)
von: Chen, Hongruixuan, et al.
Veröffentlicht: (2024)
SurFITR: A Dataset for Surveillance Image Forgery Detection and Localisation
von: Wang, Qizhou, et al.
Veröffentlicht: (2026)
von: Wang, Qizhou, et al.
Veröffentlicht: (2026)
Visual and Text Prompt Segmentation: A Novel Multi-Model Framework for Remote Sensing
von: Zi, Xing, et al.
Veröffentlicht: (2025)
von: Zi, Xing, et al.
Veröffentlicht: (2025)
Towards Blind Bitstream-corrupted Video Recovery via a Visual Foundation Model-driven Framework
von: Liu, Tianyi, et al.
Veröffentlicht: (2025)
von: Liu, Tianyi, et al.
Veröffentlicht: (2025)
PixelBoost: Leveraging Brownian Motion for Realistic-Image Super-Resolution
von: Mishra, Aradhana, et al.
Veröffentlicht: (2025)
von: Mishra, Aradhana, et al.
Veröffentlicht: (2025)
Food Portion Estimation: From Pixels to Calories
von: Vinod, Gautham, et al.
Veröffentlicht: (2026)
von: Vinod, Gautham, et al.
Veröffentlicht: (2026)
RAISE: Realness Assessment for Image Synthesis and Evaluation
von: Mukherjee, Aniruddha, et al.
Veröffentlicht: (2025)
von: Mukherjee, Aniruddha, et al.
Veröffentlicht: (2025)
From Darkness to Detail: Frequency-Aware SSMs for Low-Light Vision
von: Adhikarla, Eashan, et al.
Veröffentlicht: (2024)
von: Adhikarla, Eashan, et al.
Veröffentlicht: (2024)
CAMSIC: Content-aware Masked Image Modeling Transformer for Stereo Image Compression
von: Zhang, Xinjie, et al.
Veröffentlicht: (2024)
von: Zhang, Xinjie, et al.
Veröffentlicht: (2024)
Optimal Transcoding Resolution Prediction for Efficient Per-Title Bitrate Ladder Estimation
von: Yang, Jinhai, et al.
Veröffentlicht: (2024)
von: Yang, Jinhai, et al.
Veröffentlicht: (2024)
GaussianImage: 1000 FPS Image Representation and Compression by 2D Gaussian Splatting
von: Zhang, Xinjie, et al.
Veröffentlicht: (2024)
von: Zhang, Xinjie, et al.
Veröffentlicht: (2024)
LinMU: Multimodal Understanding Made Linear
von: Wang, Hongjie, et al.
Veröffentlicht: (2026)
von: Wang, Hongjie, et al.
Veröffentlicht: (2026)
Weakly-supervised Localization of Manipulated Image Regions Using Multi-resolution Learned Features
von: Wang, Ziyong, et al.
Veröffentlicht: (2025)
von: Wang, Ziyong, et al.
Veröffentlicht: (2025)
Video-based Music Generation
von: Sulun, Serkan
Veröffentlicht: (2026)
von: Sulun, Serkan
Veröffentlicht: (2026)
Tackle CSM in JPEG Steganalysis with Data Adaptation
von: Abecidan, Rony, et al.
Veröffentlicht: (2026)
von: Abecidan, Rony, et al.
Veröffentlicht: (2026)
DrawVideo: Generating Long Video from Storyboard Keyframe Sketches
von: Xu, Chuanzhi, et al.
Veröffentlicht: (2026)
von: Xu, Chuanzhi, et al.
Veröffentlicht: (2026)
Semantic-Aware Adaptive Video Streaming Using Latent Diffusion Models for Wireless Networks
von: Yan, Zijiang, et al.
Veröffentlicht: (2025)
von: Yan, Zijiang, et al.
Veröffentlicht: (2025)
Channel-wise Feature Decorrelation for Enhanced Learned Image Compression
von: Pakdaman, Farhad, et al.
Veröffentlicht: (2024)
von: Pakdaman, Farhad, et al.
Veröffentlicht: (2024)
Leveraging Compressed Frame Sizes For Ultra-Fast Video Classification
von: Han, Yuxing, et al.
Veröffentlicht: (2024)
von: Han, Yuxing, et al.
Veröffentlicht: (2024)
MSNeRV: Neural Video Representation with Multi-Scale Feature Fusion
von: Zhu, Jun, et al.
Veröffentlicht: (2025)
von: Zhu, Jun, et al.
Veröffentlicht: (2025)
Enhancing Blind Video Quality Assessment with Rich Quality-aware Features
von: Sun, Wei, et al.
Veröffentlicht: (2024)
von: Sun, Wei, et al.
Veröffentlicht: (2024)
R-Bench: Are your Large Multimodal Model Robust to Real-world Corruptions?
von: Li, Chunyi, et al.
Veröffentlicht: (2024)
von: Li, Chunyi, et al.
Veröffentlicht: (2024)
Toward Scalable Image Feature Compression: A Content-Adaptive and Diffusion-Based Approach
von: Guo, Sha, et al.
Veröffentlicht: (2024)
von: Guo, Sha, et al.
Veröffentlicht: (2024)
Opinion-Unaware Blind Image Quality Assessment using Multi-Scale Deep Feature Statistics
von: Ni, Zhangkai, et al.
Veröffentlicht: (2024)
von: Ni, Zhangkai, et al.
Veröffentlicht: (2024)
Learning Event-guided Exposure-agnostic Video Frame Interpolation via Adaptive Feature Blending
von: Jung, Junsik, et al.
Veröffentlicht: (2025)
von: Jung, Junsik, et al.
Veröffentlicht: (2025)
Beyond Feature Mapping GAP: Integrating Real HDRTV Priors for Superior SDRTV-to-HDRTV Conversion
von: He, Gang, et al.
Veröffentlicht: (2024)
von: He, Gang, et al.
Veröffentlicht: (2024)
Modeling Beyond MOS: Quality Assessment Models Must Integrate Context, Reasoning, and Multimodality
von: Kerkouri, Mohamed Amine, et al.
Veröffentlicht: (2025)
von: Kerkouri, Mohamed Amine, et al.
Veröffentlicht: (2025)
CAMP-VQA: Caption-Embedded Multimodal Perception for No-Reference Quality Assessment of Compressed Video
von: Wang, Xinyi, et al.
Veröffentlicht: (2025)
von: Wang, Xinyi, et al.
Veröffentlicht: (2025)
Optimizing Multimodal LLMs for Egocentric Video Understanding: A Solution for the HD-EPIC VQA Challenge
von: Yang, Sicheng, et al.
Veröffentlicht: (2026)
von: Yang, Sicheng, et al.
Veröffentlicht: (2026)
Low-Complexity Patch-based No-Reference Point Cloud Quality Metric exploiting Weighted Structure and Texture Features
von: Neri, Michael, et al.
Veröffentlicht: (2025)
von: Neri, Michael, et al.
Veröffentlicht: (2025)
Learned Compression of Point Cloud Geometry and Attributes in a Single Model through Multimodal Rate-Control
von: Rudolph, Michael, et al.
Veröffentlicht: (2024)
von: Rudolph, Michael, et al.
Veröffentlicht: (2024)
Evaluation of strategies for efficient rate-distortion NeRF streaming
von: Martin, Pedro, et al.
Veröffentlicht: (2024)
von: Martin, Pedro, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
VEMOCLAP: A video emotion classification web application
von: Sulun, Serkan, et al.
Veröffentlicht: (2024) -
Video Soundtrack Generation by Aligning Emotions and Temporal Boundaries
von: Sulun, Serkan, et al.
Veröffentlicht: (2025) -
d-Sketch: Improving Visual Fidelity of Sketch-to-Image Translation with Pretrained Latent Diffusion Models without Retraining
von: Roy, Prasun, et al.
Veröffentlicht: (2025) -
VineetVC: Adaptive Video Conferencing Under Severe Bandwidth Constraints Using Audio-Driven Talking-Head Reconstruction
von: Rakesh, Vineet Kumar, et al.
Veröffentlicht: (2026) -
DietDelta: A Vision-Language Approach for Dietary Assessment via Before-and-After Images
von: Vinod, Gautham, et al.
Veröffentlicht: (2026)