Gespeichert in:
| Hauptverfasser: | Abdoli, Mohsen, Youvalari, Ramin G., Plowman, Frank, Tissier, Alexandre |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2503.18679 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Video Compression Beyond VVC: Quantitative Analysis of Intra Coding Tools in Enhanced Compression Model (ECM)
von: Abdoli, Mohsen, et al.
Veröffentlicht: (2024)
von: Abdoli, Mohsen, et al.
Veröffentlicht: (2024)
Enhanced Template-based Intra Mode Derivation with Adaptive Block Vector Replacement
von: Zhang, Jiaqi, et al.
Veröffentlicht: (2025)
von: Zhang, Jiaqi, et al.
Veröffentlicht: (2025)
Retracted: The Development Strategy of the Multimedia Fusion Mode of Big Data Technology in Japanese Translation Teaching
von: Advances in Multimedia
Veröffentlicht: (2024)
von: Advances in Multimedia
Veröffentlicht: (2024)
Augmenting Intra-Modal Understanding in MLLMs for Robust Multimodal Keyphrase Generation
von: Cao, Jiajun, et al.
Veröffentlicht: (2025)
von: Cao, Jiajun, et al.
Veröffentlicht: (2025)
Let the Model Learn to Feel: Mode-Guided Tonality Injection for Symbolic Music Emotion Recognition
von: Xia, Haiying, et al.
Veröffentlicht: (2025)
von: Xia, Haiying, et al.
Veröffentlicht: (2025)
Audio-visual Event Localization on Portrait Mode Short Videos
von: Liu, Wuyang, et al.
Veröffentlicht: (2025)
von: Liu, Wuyang, et al.
Veröffentlicht: (2025)
MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation
von: Li, Haitian, et al.
Veröffentlicht: (2026)
von: Li, Haitian, et al.
Veröffentlicht: (2026)
Editing Away the Evidence: Diffusion-Based Image Manipulation and the Failure Modes of Robust Watermarking
von: Qi, Qian, et al.
Veröffentlicht: (2026)
von: Qi, Qian, et al.
Veröffentlicht: (2026)
Evaluation of Objective Image Quality Metrics for High-Fidelity Image Compression
von: Mohammadi, Shima, et al.
Veröffentlicht: (2025)
von: Mohammadi, Shima, et al.
Veröffentlicht: (2025)
In-place Double Stimulus Methodology for Subjective Assessment of High Quality Images
von: Mohammadi, Shima, et al.
Veröffentlicht: (2025)
von: Mohammadi, Shima, et al.
Veröffentlicht: (2025)
Multi Agents Semantic Emotion Aligned Music to Image Generation with Music Derived Captions
von: Shi, Junchang, et al.
Veröffentlicht: (2025)
von: Shi, Junchang, et al.
Veröffentlicht: (2025)
PIRA: Pan-CDN Intra-video Resource Adaptation for Short Video Streaming
von: Qiao, Chunyu, et al.
Veröffentlicht: (2025)
von: Qiao, Chunyu, et al.
Veröffentlicht: (2025)
Clustering Internet Memes Through Template Matching and Multi-Dimensional Similarity
von: Bloem, Tygo, et al.
Veröffentlicht: (2025)
von: Bloem, Tygo, et al.
Veröffentlicht: (2025)
Low Complexity Learning-based Lossless Event-based Compression
von: Sezavar, Ahmadreza, et al.
Veröffentlicht: (2024)
von: Sezavar, Ahmadreza, et al.
Veröffentlicht: (2024)
G4G:A Generic Framework for High Fidelity Talking Face Generation with Fine-grained Intra-modal Alignment
von: Zhang, Juan, et al.
Veröffentlicht: (2024)
von: Zhang, Juan, et al.
Veröffentlicht: (2024)
UVG-VPC: Voxelized Point Cloud Dataset for Visual Volumetric Video-based Coding
von: Gautier, Guillaume, et al.
Veröffentlicht: (2025)
von: Gautier, Guillaume, et al.
Veröffentlicht: (2025)
Learning-based Lossless Event Data Compression
von: Sezavar, Ahmadreza, et al.
Veröffentlicht: (2024)
von: Sezavar, Ahmadreza, et al.
Veröffentlicht: (2024)
Layer-wise Model Merging for Unsupervised Domain Adaptation in Segmentation Tasks
von: Alcover-Couso, Roberto, et al.
Veröffentlicht: (2024)
von: Alcover-Couso, Roberto, et al.
Veröffentlicht: (2024)
Multimodal LLM-based Query Paraphrasing for Video Search
von: Wu, Jiaxin, et al.
Veröffentlicht: (2024)
von: Wu, Jiaxin, et al.
Veröffentlicht: (2024)
A CLIP-based siamese approach for meme classification
von: Huertas-Tato, Javier, et al.
Veröffentlicht: (2024)
von: Huertas-Tato, Javier, et al.
Veröffentlicht: (2024)
Start from Video-Music Retrieval: An Inter-Intra Modal Loss for Cross Modal Retrieval
von: Chen, Zeyu, et al.
Veröffentlicht: (2024)
von: Chen, Zeyu, et al.
Veröffentlicht: (2024)
QuMATL: Query-based Multi-annotator Tendency Learning
von: Zhang, Liyun, et al.
Veröffentlicht: (2025)
von: Zhang, Liyun, et al.
Veröffentlicht: (2025)
Structured Image-based Coding for Efficient Gaussian Splatting Compression
von: Martin, Pedro, et al.
Veröffentlicht: (2026)
von: Martin, Pedro, et al.
Veröffentlicht: (2026)
Graph-based Interaction Augmentation Network for Robust Multimodal Sentiment Analysis
von: Zhangfeng, Hu, et al.
Veröffentlicht: (2025)
von: Zhangfeng, Hu, et al.
Veröffentlicht: (2025)
MMC: Iterative Refinement of VLM Reasoning via MCTS-based Multimodal Critique
von: Liu, Shuhang, et al.
Veröffentlicht: (2025)
von: Liu, Shuhang, et al.
Veröffentlicht: (2025)
Movement- and Traffic-based User Identification in Commercial Virtual Reality Applications: Threats and Opportunities
von: Baldoni, Sara, et al.
Veröffentlicht: (2025)
von: Baldoni, Sara, et al.
Veröffentlicht: (2025)
Gain of Grain: A Film Grain Handling Toolchain for VVC-based Open Implementations
von: Menon, Vignesh V, et al.
Veröffentlicht: (2024)
von: Menon, Vignesh V, et al.
Veröffentlicht: (2024)
VRAgent-R1: Boosting Video Recommendation with MLLM-based Agents via Reinforcement Learning
von: Chen, Siran, et al.
Veröffentlicht: (2025)
von: Chen, Siran, et al.
Veröffentlicht: (2025)
ProMSC-MIS: Prompt-based Multimodal Semantic Communication for Multi-Spectral Image Segmentation
von: Zhang, Haoshuo, et al.
Veröffentlicht: (2025)
von: Zhang, Haoshuo, et al.
Veröffentlicht: (2025)
Nagare Media Engine: A System for Cloud- and Edge-Native Network-based Multimedia Workflows
von: Neugebauer, Matthias
Veröffentlicht: (2025)
von: Neugebauer, Matthias
Veröffentlicht: (2025)
Multi-view Hypergraph-based Contrastive Learning Model for Cold-Start Micro-video Recommendation
von: Lyu, Sisuo, et al.
Veröffentlicht: (2024)
von: Lyu, Sisuo, et al.
Veröffentlicht: (2024)
Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark
von: Zhang, Han, et al.
Veröffentlicht: (2025)
von: Zhang, Han, et al.
Veröffentlicht: (2025)
Startup Delay Aware Short Video Ordering: Problem, Model, and A Reinforcement Learning based Algorithm
von: Gao, Zhipeng, et al.
Veröffentlicht: (2024)
von: Gao, Zhipeng, et al.
Veröffentlicht: (2024)
YTLive: A Dataset of Real-World YouTube Live Streaming Sessions
von: Mozhganfar, Mojtaba, et al.
Veröffentlicht: (2025)
von: Mozhganfar, Mojtaba, et al.
Veröffentlicht: (2025)
IIANet: An Intra- and Inter-Modality Attention Network for Audio-Visual Speech Separation
von: Li, Kai, et al.
Veröffentlicht: (2023)
von: Li, Kai, et al.
Veröffentlicht: (2023)
M3ST-DTI: A multi-task learning model for drug-target interactions based on multi-modal features and multi-stage alignment
von: Li, Xiangyu, et al.
Veröffentlicht: (2025)
von: Li, Xiangyu, et al.
Veröffentlicht: (2025)
ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
von: Zhang, Zhenxing, et al.
Veröffentlicht: (2024)
von: Zhang, Zhenxing, et al.
Veröffentlicht: (2024)
MMED: A Multimodal Micro-Expression Dataset based on Audio-Visual Fusion
von: Wang, Junbo, et al.
Veröffentlicht: (2025)
von: Wang, Junbo, et al.
Veröffentlicht: (2025)
Ensembling Synchronisation-based and Face-Voice Association Paradigms for Robust Active Speaker Detection in Egocentric Recordings
von: Clarke, Jason, et al.
Veröffentlicht: (2025)
von: Clarke, Jason, et al.
Veröffentlicht: (2025)
Inferencias causales durante la comprensión de textos expositivos en formato multimedia
von: Gastón Saux
Veröffentlicht: (2012)
von: Gastón Saux
Veröffentlicht: (2012)
Ähnliche Einträge
-
Video Compression Beyond VVC: Quantitative Analysis of Intra Coding Tools in Enhanced Compression Model (ECM)
von: Abdoli, Mohsen, et al.
Veröffentlicht: (2024) -
Enhanced Template-based Intra Mode Derivation with Adaptive Block Vector Replacement
von: Zhang, Jiaqi, et al.
Veröffentlicht: (2025) -
Retracted: The Development Strategy of the Multimedia Fusion Mode of Big Data Technology in Japanese Translation Teaching
von: Advances in Multimedia
Veröffentlicht: (2024) -
Augmenting Intra-Modal Understanding in MLLMs for Robust Multimodal Keyphrase Generation
von: Cao, Jiajun, et al.
Veröffentlicht: (2025) -
Let the Model Learn to Feel: Mode-Guided Tonality Injection for Symbolic Music Emotion Recognition
von: Xia, Haiying, et al.
Veröffentlicht: (2025)