The Future is Meta: Metadata, Formats and Perspectives towards Interactive and Personalized AV Content
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Weller, Alexander, Bleisteiner, Werner, Hufnagel, Christian, Iber, Michael |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Multi-modal and Metadata Capture Model for Micro Video Popularity Prediction
par: Lu, Jiacheng, et autres
Publié: (2025)
par: Lu, Jiacheng, et autres
Publié: (2025)
Compression Metadata-assisted RoI Extraction and Adaptive Inference for Efficient Video Analytics
par: Wang, Chengzhi, et autres
Publié: (2025)
par: Wang, Chengzhi, et autres
Publié: (2025)
Transform and Entropy Coding in AV2
par: Nalci, Alican, et autres
Publié: (2026)
par: Nalci, Alican, et autres
Publié: (2026)
LongAV-Compass: Towards Unified Evaluation of Minute-Scale Audio-Visual Generation Across T2AV, I2AV, and V2AV
par: Liu, Tengfei, et autres
Publié: (2026)
par: Liu, Tengfei, et autres
Publié: (2026)
Encoding Time and Energy Model for SVT-AV1 based on Video Complexity
par: Eichermüller, Lena, et autres
Publié: (2024)
par: Eichermüller, Lena, et autres
Publié: (2024)
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues
par: Zhou, Dingkun, et autres
Publié: (2025)
par: Zhou, Dingkun, et autres
Publié: (2025)
Improved Screen Content Coding in VVC Using Soft Context Formation
par: Och, Hannah, et autres
Publié: (2023)
par: Och, Hannah, et autres
Publié: (2023)
Leveraging User-Generated Metadata of Online Videos for Cover Song Identification
par: Hachmeier, Simon, et autres
Publié: (2024)
par: Hachmeier, Simon, et autres
Publié: (2024)
Balanced Multimodal Learning: An Unidirectional Dynamic Interaction Perspective
par: Wang, Shijie, et autres
Publié: (2025)
par: Wang, Shijie, et autres
Publié: (2025)
AIM 2024 Challenge on Efficient Video Super-Resolution for AV1 Compressed Content
par: Conde, Marcos V, et autres
Publié: (2024)
par: Conde, Marcos V, et autres
Publié: (2024)
AV1 Motion Vector Fidelity and Application for Efficient Optical Flow
par: Zouein, Julien, et autres
Publié: (2025)
par: Zouein, Julien, et autres
Publié: (2025)
Interoperable Provenance Authentication of Broadcast Media using Open Standards-based Metadata, Watermarking and Cryptography
par: Simmons, John C., et autres
Publié: (2024)
par: Simmons, John C., et autres
Publié: (2024)
EquiAV: Leveraging Equivariance for Audio-Visual Contrastive Learning
par: Kim, Jongsuk, et autres
Publié: (2024)
par: Kim, Jongsuk, et autres
Publié: (2024)
Human-centered Interactive Learning via MLLMs for Text-to-Image Person Re-identification
par: Qin, Yang, et autres
Publié: (2025)
par: Qin, Yang, et autres
Publié: (2025)
TOP:A New Target-Audience Oriented Content Paraphrase Task
par: Lin, Boda, et autres
Publié: (2024)
par: Lin, Boda, et autres
Publié: (2024)
Prototypical Prompting for Text-to-image Person Re-identification
par: Yan, Shuanglin, et autres
Publié: (2024)
par: Yan, Shuanglin, et autres
Publié: (2024)
Flexible Control in Symbolic Music Generation via Musical Metadata
par: Han, Sangjun, et autres
Publié: (2024)
par: Han, Sangjun, et autres
Publié: (2024)
Content-Adaptive Rate-Quality Curve Prediction Model in Media Processing System
par: Yin, Shibo, et autres
Publié: (2024)
par: Yin, Shibo, et autres
Publié: (2024)
Fully Automatic Content-Aware Tiling Pipeline for Pathology Whole Slide Images
par: Jabar, Falah, et autres
Publié: (2024)
par: Jabar, Falah, et autres
Publié: (2024)
ProAV-DiT: A Projected Latent Diffusion Transformer for Efficient Synchronized Audio-Video Generation
par: Sun, Jiahui, et autres
Publié: (2025)
par: Sun, Jiahui, et autres
Publié: (2025)
A Large-scale Dataset with Behavior, Attributes, and Content of Mobile Short-video Platform
par: Shang, Yu, et autres
Publié: (2025)
par: Shang, Yu, et autres
Publié: (2025)
CAMeL: Cross-modality Adaptive Meta-Learning for Text-based Person Retrieval
par: Yu, Hang, et autres
Publié: (2025)
par: Yu, Hang, et autres
Publié: (2025)
Think before You Leap: Content-Aware Low-Cost Edge-Assisted Video Semantic Segmentation
par: Yan, Mingxuan, et autres
Publié: (2024)
par: Yan, Mingxuan, et autres
Publié: (2024)
SPICA: Interactive Video Content Exploration through Augmented Audio Descriptions for Blind or Low-Vision Viewers
par: Ning, Zheng, et autres
Publié: (2024)
par: Ning, Zheng, et autres
Publié: (2024)
Personalized Playback Technology: How Short Video Services Create Excellent User Experience
par: Deng, Weihui, et autres
Publié: (2024)
par: Deng, Weihui, et autres
Publié: (2024)
Harnessing Multimodal Large Language Models for Personalized Product Search with Query-aware Refinement
par: Zhang, Beibei, et autres
Publié: (2025)
par: Zhang, Beibei, et autres
Publié: (2025)
Task Presentation and Human Perception in Interactive Video Retrieval
par: Willis, Nina, et autres
Publié: (2024)
par: Willis, Nina, et autres
Publié: (2024)
Integrated Semantic and Temporal Alignment for Interactive Video Retrieval
par: Luu, Thanh-Danh, et autres
Publié: (2025)
par: Luu, Thanh-Danh, et autres
Publié: (2025)
SpaceMeta: Global-Scale Massive Multi-User Virtual Interaction over LEO Satellite Constellations
par: Huang, Jiahe, et autres
Publié: (2024)
par: Huang, Jiahe, et autres
Publié: (2024)
Content-Driven Frame-Level Bit Prediction for Rate Control in Versatile Video Coding
par: Premkumar, Amritha, et autres
Publié: (2026)
par: Premkumar, Amritha, et autres
Publié: (2026)
Fostering Emotional Perspective-Taking: An Exploration of Affective Face-Tracking Interactions in the VR Narrative Rekindle
par: Fan, Hector, et autres
Publié: (2026)
par: Fan, Hector, et autres
Publié: (2026)
Dynamic Interaction-Aware and Causality-Disentangled Framework for Multimodal Sentiment Analysis
par: Dong, Guangyuan, et autres
Publié: (2026)
par: Dong, Guangyuan, et autres
Publié: (2026)
EmotionTalk: An Interactive Chinese Multimodal Emotion Dataset With Rich Annotations
par: Sun, Haoqin, et autres
Publié: (2025)
par: Sun, Haoqin, et autres
Publié: (2025)
Graph-based Interaction Augmentation Network for Robust Multimodal Sentiment Analysis
par: Zhangfeng, Hu, et autres
Publié: (2025)
par: Zhangfeng, Hu, et autres
Publié: (2025)
Dual-Stream Decoupled Learning for Temporal Consistency and Speaker Interaction in AVSD
par: Xiao, Junhao, et autres
Publié: (2025)
par: Xiao, Junhao, et autres
Publié: (2025)
Modeling Human Responses to Multimodal AI Content
par: Shen, Zhiqi, et autres
Publié: (2025)
par: Shen, Zhiqi, et autres
Publié: (2025)
Lossless 4:2:0 Screen Content Coding Using Luma-Guided Soft Context Formation
par: Och, Hannah, et autres
Publié: (2025)
par: Och, Hannah, et autres
Publié: (2025)
diveXplore 6.0: ITEC's Interactive Video Exploration System at VBS 2022
par: Leibetseder, Andreas, et autres
Publié: (2025)
par: Leibetseder, Andreas, et autres
Publié: (2025)
Beyond Isolated Utterances: Cue-Guided Interaction for Context-Dependent Conversational Multimodal Understanding
par: Pan, Zhaoyan, et autres
Publié: (2026)
par: Pan, Zhaoyan, et autres
Publié: (2026)
XGC-AVis: Towards Audio-Visual Content Understanding with a Multi-Agent Collaborative System
par: Cao, Yuqin, et autres
Publié: (2025)
par: Cao, Yuqin, et autres
Publié: (2025)
Documents similaires
-
Multi-modal and Metadata Capture Model for Micro Video Popularity Prediction
par: Lu, Jiacheng, et autres
Publié: (2025) -
Compression Metadata-assisted RoI Extraction and Adaptive Inference for Efficient Video Analytics
par: Wang, Chengzhi, et autres
Publié: (2025) -
Transform and Entropy Coding in AV2
par: Nalci, Alican, et autres
Publié: (2026) -
LongAV-Compass: Towards Unified Evaluation of Minute-Scale Audio-Visual Generation Across T2AV, I2AV, and V2AV
par: Liu, Tengfei, et autres
Publié: (2026) -
Encoding Time and Energy Model for SVT-AV1 based on Video Complexity
par: Eichermüller, Lena, et autres
Publié: (2024)