Stepwise Schema-Guided Prompting Framework with Parameter Efficient Instruction Tuning for Multimedia Event Extraction
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yuan, Xiang, Chen, Xinrong, Li, Haochen, Yang, Hang, Wang, Guanyu, Li, Weiping, Mo, Tong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MORE-R1: Guiding LVLM for Multimodal Object-Entity Relation Extraction via Stepwise Reasoning with Reinforcement Learning
von: Yuan, Xiang, et al.
Veröffentlicht: (2026)
von: Yuan, Xiang, et al.
Veröffentlicht: (2026)
Benchmarking and Improving LVLMs on Event Extraction from Multimedia Documents
von: Xing, Fuyu, et al.
Veröffentlicht: (2025)
von: Xing, Fuyu, et al.
Veröffentlicht: (2025)
Efficient Prompt Tuning for Hierarchical Ingredient Recognition
von: Gui, Yinxuan, et al.
Veröffentlicht: (2025)
von: Gui, Yinxuan, et al.
Veröffentlicht: (2025)
Knowledge-aware Diffusion-Enhanced Multimedia Recommendation
von: Mo, Xian, et al.
Veröffentlicht: (2025)
von: Mo, Xian, et al.
Veröffentlicht: (2025)
Towards Event Extraction from Speech with Contextual Clues
von: Kang, Jingqi, et al.
Veröffentlicht: (2024)
von: Kang, Jingqi, et al.
Veröffentlicht: (2024)
SpeechEE: A Novel Benchmark for Speech Event Extraction
von: Wang, Bin, et al.
Veröffentlicht: (2024)
von: Wang, Bin, et al.
Veröffentlicht: (2024)
Concept Drift Guided LayerNorm Tuning for Efficient Multimodal Metaphor Identification
von: Qian, Wenhao, et al.
Veröffentlicht: (2025)
von: Qian, Wenhao, et al.
Veröffentlicht: (2025)
SIDQL: An Efficient Keyframe Extraction and Motion Reconstruction Framework in Motion Capture
von: Zhang, Xuling, et al.
Veröffentlicht: (2024)
von: Zhang, Xuling, et al.
Veröffentlicht: (2024)
A Multimedia Framework for Continuum Robots: Systematic, Computational, and Control Perspectives
von: Hsieh, Po-Yu, et al.
Veröffentlicht: (2024)
von: Hsieh, Po-Yu, et al.
Veröffentlicht: (2024)
Guided Masked Self-Distillation Modeling for Distributed Multimedia Sensor Event Analysis
von: Yasuda, Masahiro, et al.
Veröffentlicht: (2024)
von: Yasuda, Masahiro, et al.
Veröffentlicht: (2024)
Identity-Driven Multimedia Forgery Detection via Reference Assistance
von: Xu, Junhao, et al.
Veröffentlicht: (2024)
von: Xu, Junhao, et al.
Veröffentlicht: (2024)
MAO: Efficient Model-Agnostic Optimization of Prompt Tuning for Vision-Language Models
von: Li, Haoyang, et al.
Veröffentlicht: (2025)
von: Li, Haoyang, et al.
Veröffentlicht: (2025)
Turing Patterns for Multimedia: Reaction-Diffusion Multi-Modal Fusion for Language-Guided Video Moment Retrieval
von: Fang, Xiang, et al.
Veröffentlicht: (2026)
von: Fang, Xiang, et al.
Veröffentlicht: (2026)
Digital Fingerprinting on Multimedia: A Survey
von: Chen, Wendi, et al.
Veröffentlicht: (2024)
von: Chen, Wendi, et al.
Veröffentlicht: (2024)
UMETTS: A Unified Framework for Emotional Text-to-Speech Synthesis with Multimodal Prompts
von: Cheng, Zhi-Qi, et al.
Veröffentlicht: (2024)
von: Cheng, Zhi-Qi, et al.
Veröffentlicht: (2024)
Can We Hear from Events? Generating Speech from Event Camera
von: Fang, Jingping, et al.
Veröffentlicht: (2026)
von: Fang, Jingping, et al.
Veröffentlicht: (2026)
Latent Feature-Guided Conditional Diffusion for Generative Image Semantic Communication
von: Chen, Zehao, et al.
Veröffentlicht: (2025)
von: Chen, Zehao, et al.
Veröffentlicht: (2025)
Language-oriented Semantic Communication for Image Transmission with Fine-Tuned Diffusion Model
von: Wei, Xinfeng, et al.
Veröffentlicht: (2024)
von: Wei, Xinfeng, et al.
Veröffentlicht: (2024)
Compression Metadata-assisted RoI Extraction and Adaptive Inference for Efficient Video Analytics
von: Wang, Chengzhi, et al.
Veröffentlicht: (2025)
von: Wang, Chengzhi, et al.
Veröffentlicht: (2025)
Nagare Media Ingest: A System for Multimedia Ingest Workflows
von: Neugebauer, Matthias
Veröffentlicht: (2025)
von: Neugebauer, Matthias
Veröffentlicht: (2025)
Reducing Latency for Multimedia Broadcast Services Over Mobile Networks
von: Lentisco, C. M., et al.
Veröffentlicht: (2024)
von: Lentisco, C. M., et al.
Veröffentlicht: (2024)
Q-Adapt: Adapting LMM for Visual Quality Assessment with Progressive Instruction Tuning
von: Lu, Yiting, et al.
Veröffentlicht: (2025)
von: Lu, Yiting, et al.
Veröffentlicht: (2025)
ARGUS: Defending Against Multimodal Indirect Prompt Injection via Steering Instruction-Following Behavior
von: Lu, Weikai, et al.
Veröffentlicht: (2025)
von: Lu, Weikai, et al.
Veröffentlicht: (2025)
CLAIP-Emo: Parameter-Efficient Adaptation of Language-supervised models for In-the-Wild Audiovisual Emotion Recognition
von: Chen, Yin, et al.
Veröffentlicht: (2025)
von: Chen, Yin, et al.
Veröffentlicht: (2025)
TAROT: Towards Optimization-Driven Adaptive FEC Parameter Tuning for Video Streaming
von: Sidhu, Jashanjot Singh, et al.
Veröffentlicht: (2026)
von: Sidhu, Jashanjot Singh, et al.
Veröffentlicht: (2026)
Performance Evaluation in Multimedia Retrieval
von: Sauter, Loris, et al.
Veröffentlicht: (2024)
von: Sauter, Loris, et al.
Veröffentlicht: (2024)
Block Erasure-Aware Semantic Multimedia Compression via JSCC Autoencoder
von: Esfahanizadeh, Homa, et al.
Veröffentlicht: (2026)
von: Esfahanizadeh, Homa, et al.
Veröffentlicht: (2026)
Privacy-Preserving Multimedia Mobile Cloud Computing Using Protective Perturbation
von: Tang, Zhongze, et al.
Veröffentlicht: (2024)
von: Tang, Zhongze, et al.
Veröffentlicht: (2024)
UniPT: Universal Parallel Tuning for Transfer Learning with Efficient Parameter and Memory
von: Diao, Haiwen, et al.
Veröffentlicht: (2023)
von: Diao, Haiwen, et al.
Veröffentlicht: (2023)
SACRED: A Faithful Annotated Multimedia Multimodal Multilingual Dataset for Classifying Connectedness Types in Online Spirituality
von: Guan, Qinghao, et al.
Veröffentlicht: (2026)
von: Guan, Qinghao, et al.
Veröffentlicht: (2026)
Generative AI-enabled Mobile Tactical Multimedia Networks: Distribution, Generation, and Perception
von: Xu, Minrui, et al.
Veröffentlicht: (2024)
von: Xu, Minrui, et al.
Veröffentlicht: (2024)
Teach Me Sign: Stepwise Prompting LLM for Sign Language Production
von: An, Zhaoyi, et al.
Veröffentlicht: (2025)
von: An, Zhaoyi, et al.
Veröffentlicht: (2025)
Mining the Social Fabric: Unveiling Communities for Fake News Detection in Short Videos
von: Gong, Haisong, et al.
Veröffentlicht: (2025)
von: Gong, Haisong, et al.
Veröffentlicht: (2025)
Characterizing Multimedia Information Environment through Multi-modal Clustering of YouTube Videos
von: Yousefi, Niloofar, et al.
Veröffentlicht: (2024)
von: Yousefi, Niloofar, et al.
Veröffentlicht: (2024)
Instruction-Guided Editing Controls for Images and Multimedia: A Survey in LLM era
von: Nguyen, Thanh Tam, et al.
Veröffentlicht: (2024)
von: Nguyen, Thanh Tam, et al.
Veröffentlicht: (2024)
Modeling the Popularity of Events on Web by Sparsity and Mutual-Excitation Guided Graph Neural Network
von: Deng, Jiaxin, et al.
Veröffentlicht: (2024)
von: Deng, Jiaxin, et al.
Veröffentlicht: (2024)
Design of a 5G Multimedia Broadcast Application Function Supporting Adaptive Error Recovery
von: Lentisco, C. M., et al.
Veröffentlicht: (2024)
von: Lentisco, C. M., et al.
Veröffentlicht: (2024)
Nagare Media Engine: A System for Cloud- and Edge-Native Network-based Multimedia Workflows
von: Neugebauer, Matthias
Veröffentlicht: (2025)
von: Neugebauer, Matthias
Veröffentlicht: (2025)
PDA: Text-Augmented Defense Framework for Robust Vision-Language Models against Adversarial Image Attacks
von: Xu, Jingning, et al.
Veröffentlicht: (2026)
von: Xu, Jingning, et al.
Veröffentlicht: (2026)
Multimedia Courseware and Multimedia "Telemedicine" at AMSIE '96.
von: Mihram, Danielle, et al.
Veröffentlicht: (1996)
von: Mihram, Danielle, et al.
Veröffentlicht: (1996)
Ähnliche Einträge
-
MORE-R1: Guiding LVLM for Multimodal Object-Entity Relation Extraction via Stepwise Reasoning with Reinforcement Learning
von: Yuan, Xiang, et al.
Veröffentlicht: (2026) -
Benchmarking and Improving LVLMs on Event Extraction from Multimedia Documents
von: Xing, Fuyu, et al.
Veröffentlicht: (2025) -
Efficient Prompt Tuning for Hierarchical Ingredient Recognition
von: Gui, Yinxuan, et al.
Veröffentlicht: (2025) -
Knowledge-aware Diffusion-Enhanced Multimedia Recommendation
von: Mo, Xian, et al.
Veröffentlicht: (2025) -
Towards Event Extraction from Speech with Contextual Clues
von: Kang, Jingqi, et al.
Veröffentlicht: (2024)