FortisAVQA and MAVEN: a Benchmark Dataset and Debiasing Framework for Robust Multimodal Reasoning
Fuente:
arXiv
Salvato in:
| Autori principali: | Ma, Jie, Gao, Zhitao, Chai, Qi, Liu, Jun, Wang, Pinghui, Tao, Jing, Su, Zhou |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Debate on Graph: a Flexible and Reliable Reasoning Framework for Large Language Models
di: Ma, Jie, et al.
Pubblicazione: (2024)
di: Ma, Jie, et al.
Pubblicazione: (2024)
Deliberation on Priors: Trustworthy Reasoning of Large Language Models on Knowledge Graphs
di: Ma, Jie, et al.
Pubblicazione: (2025)
di: Ma, Jie, et al.
Pubblicazione: (2025)
EXR: An Interactive Immersive EHR Visualization in Extended Reality
di: Marteau, Benoit, et al.
Pubblicazione: (2025)
di: Marteau, Benoit, et al.
Pubblicazione: (2025)
HOME-KGQA: A Benchmark Dataset for Multimodal Knowledge Graph Question Answering on Household Daily Activities
di: Egami, Shusaku, et al.
Pubblicazione: (2026)
di: Egami, Shusaku, et al.
Pubblicazione: (2026)
INDS: Incremental Named Data Streaming for Real-Time Point Cloud Video
di: Chai, Ruonan, et al.
Pubblicazione: (2025)
di: Chai, Ruonan, et al.
Pubblicazione: (2025)
A Real-Time, Vision-Based System for Badminton Smash Speed Estimation on Mobile Devices
di: Huang, Diwen
Pubblicazione: (2025)
di: Huang, Diwen
Pubblicazione: (2025)
FASH-iCNN: Making Editorial Fashion Identity Inspectable Through Multimodal CNN Probing
di: Adeyemi, Morayo Danielle, et al.
Pubblicazione: (2026)
di: Adeyemi, Morayo Danielle, et al.
Pubblicazione: (2026)
Scene Detection Policies and Keyframe Extraction Strategies for Large-Scale Video Analysis
di: Korolkov, Vasilii
Pubblicazione: (2025)
di: Korolkov, Vasilii
Pubblicazione: (2025)
Geo2Sound: A Scalable Geo-Aligned Framework for Soundscape Generation from Satellite Imagery
di: Wu, Kunlin, et al.
Pubblicazione: (2026)
di: Wu, Kunlin, et al.
Pubblicazione: (2026)
Generative AI for Video Translation: A Scalable Architecture for Multilingual Video Conferencing
di: Oskooei, Amirkia Rafiei, et al.
Pubblicazione: (2025)
di: Oskooei, Amirkia Rafiei, et al.
Pubblicazione: (2025)
StyleMM: Stylized 3D Morphable Face Model via Text-Driven Aligned Image Translation
di: Lee, Seungmi, et al.
Pubblicazione: (2025)
di: Lee, Seungmi, et al.
Pubblicazione: (2025)
EditIQ: Automated Cinematic Editing of Static Wide-Angle Videos via Dialogue Interpretation and Saliency Cues
di: Girmaji, Rohit, et al.
Pubblicazione: (2025)
di: Girmaji, Rohit, et al.
Pubblicazione: (2025)
Empowering VLMs for Few-Shot Multimodal Time Series Classification via Tailored Agentic Reasoning
di: Li, Lin, et al.
Pubblicazione: (2026)
di: Li, Lin, et al.
Pubblicazione: (2026)
U-SEG: Uncertainty in SEGmentation -- A systematic multi-variable exploration
di: Smith, Michael, et al.
Pubblicazione: (2026)
di: Smith, Michael, et al.
Pubblicazione: (2026)
A Human-Machine Collaboration Framework for the Development of Schemas
di: Isaak, Nicos
Pubblicazione: (2024)
di: Isaak, Nicos
Pubblicazione: (2024)
RCI: A Score for Evaluating Global and Local Reasoning in Multimodal Benchmarks
di: Agarwal, Amit, et al.
Pubblicazione: (2025)
di: Agarwal, Amit, et al.
Pubblicazione: (2025)
Automatic Detection of Intro and Credits in Video using CLIP and Multihead Attention
di: Korolkov, Vasilii, et al.
Pubblicazione: (2025)
di: Korolkov, Vasilii, et al.
Pubblicazione: (2025)
Motion Attribution for Video Generation
di: Wu, Xindi, et al.
Pubblicazione: (2026)
di: Wu, Xindi, et al.
Pubblicazione: (2026)
IDN Authoring -- a design case
di: Nack, Frank
Pubblicazione: (2023)
di: Nack, Frank
Pubblicazione: (2023)
Listen to the Unexpected: Self-Supervised Surprise Detection for Efficient Viewport Prediction
di: Khah, Arman Nik, et al.
Pubblicazione: (2026)
di: Khah, Arman Nik, et al.
Pubblicazione: (2026)
Predicting the descent into extremism and terrorism
di: Lane, R. O., et al.
Pubblicazione: (2025)
di: Lane, R. O., et al.
Pubblicazione: (2025)
Personalized QoE Prediction: A Demographic-Augmented Machine Learning Framework for 5G Video Streaming Networks
di: Ahmed, Syeda Zunaira, et al.
Pubblicazione: (2025)
di: Ahmed, Syeda Zunaira, et al.
Pubblicazione: (2025)
ADALog: Adaptive Unsupervised Anomaly detection in Logs with Self-attention Masked Language Model
di: Pospieszny, Przemek, et al.
Pubblicazione: (2025)
di: Pospieszny, Przemek, et al.
Pubblicazione: (2025)
FedMentalCare: Towards Privacy-Preserving Fine-Tuned LLMs to Analyze Mental Health Status Using Federated Learning Framework
di: Sarwar, Nobin
Pubblicazione: (2025)
di: Sarwar, Nobin
Pubblicazione: (2025)
Evaluation Before Generation: A Paradigm for Robust Multimodal Sentiment Analysis with Missing Modalities
di: Chen, Rongfei, et al.
Pubblicazione: (2026)
di: Chen, Rongfei, et al.
Pubblicazione: (2026)
Towards a Robust Framework for Multimodal Hate Detection: A Study on Video vs. Image-based Content
di: Koushik, Girish A., et al.
Pubblicazione: (2025)
di: Koushik, Girish A., et al.
Pubblicazione: (2025)
The Interaction Fidelity Model: A Taxonomy to Distinguish the Aspects of Fidelity in Virtual Reality
di: Bonfert, Michael, et al.
Pubblicazione: (2024)
di: Bonfert, Michael, et al.
Pubblicazione: (2024)
MIMIC-SR-ICD11: A Dataset for Narrative-Based Diagnosis
di: Wu, Yuexin, et al.
Pubblicazione: (2025)
di: Wu, Yuexin, et al.
Pubblicazione: (2025)
Contrastive Similarity Learning for Market Forecasting: The ContraSim Framework
di: Vinden, Nicholas, et al.
Pubblicazione: (2025)
di: Vinden, Nicholas, et al.
Pubblicazione: (2025)
A group-theoretic framework for machine learning in hyperbolic spaces
di: Jaćimović, Vladimir
Pubblicazione: (2025)
di: Jaćimović, Vladimir
Pubblicazione: (2025)
MemeCraft: Contextual and Stance-Driven Multimodal Meme Generation
di: Wang, Han, et al.
Pubblicazione: (2024)
di: Wang, Han, et al.
Pubblicazione: (2024)
VHAKG: A Multi-modal Knowledge Graph Based on Synchronized Multi-view Videos of Daily Activities
di: Egami, Shusaku, et al.
Pubblicazione: (2024)
di: Egami, Shusaku, et al.
Pubblicazione: (2024)
Causal Debiasing for Visual Commonsense Reasoning
di: Zou, Jiayi, et al.
Pubblicazione: (2025)
di: Zou, Jiayi, et al.
Pubblicazione: (2025)
CR-LT-KGQA: A Knowledge Graph Question Answering Dataset Requiring Commonsense Reasoning and Long-Tail Knowledge
di: Guo, Willis, et al.
Pubblicazione: (2024)
di: Guo, Willis, et al.
Pubblicazione: (2024)
A SHACL-based Data Consistency Solution for Contract Compliance Verification (Extended Paper)
di: David, Robert, et al.
Pubblicazione: (2025)
di: David, Robert, et al.
Pubblicazione: (2025)
MMRec: Simplifying Multimodal Recommendation
di: Zhou, Xin
Pubblicazione: (2023)
di: Zhou, Xin
Pubblicazione: (2023)
Benchmarking and Bridging Emotion Conflicts for Multimodal Emotion Reasoning
di: Han, Zhiyuan, et al.
Pubblicazione: (2025)
di: Han, Zhiyuan, et al.
Pubblicazione: (2025)
MultiHateClip: A Multilingual Benchmark Dataset for Hateful Video Detection on YouTube and Bilibili
di: Wang, Han, et al.
Pubblicazione: (2024)
di: Wang, Han, et al.
Pubblicazione: (2024)
Meaning over Motion: A Semantic-First Approach to 360° Viewport Prediction
di: Khah, Arman Nik, et al.
Pubblicazione: (2026)
di: Khah, Arman Nik, et al.
Pubblicazione: (2026)
A Self-explainable Model of Long Time Series by Extracting Informative Structured Causal Patterns
di: Wang, Ziqian, et al.
Pubblicazione: (2025)
di: Wang, Ziqian, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Debate on Graph: a Flexible and Reliable Reasoning Framework for Large Language Models
di: Ma, Jie, et al.
Pubblicazione: (2024) -
Deliberation on Priors: Trustworthy Reasoning of Large Language Models on Knowledge Graphs
di: Ma, Jie, et al.
Pubblicazione: (2025) -
EXR: An Interactive Immersive EHR Visualization in Extended Reality
di: Marteau, Benoit, et al.
Pubblicazione: (2025) -
HOME-KGQA: A Benchmark Dataset for Multimodal Knowledge Graph Question Answering on Household Daily Activities
di: Egami, Shusaku, et al.
Pubblicazione: (2026) -
INDS: Incremental Named Data Streaming for Real-Time Point Cloud Video
di: Chai, Ruonan, et al.
Pubblicazione: (2025)