VALSE: A Task-Independent Benchmark for Vision and Language Models Centered on Linguistic Phenomena
Fuente:
arXiv
Saved in:
| Main Authors: | Parcalabescu, Letitia, Cafagna, Michele, Muradjan, Lilitta, Frank, Anette, Calixto, Iacer, Gatt, Albert |
|---|---|
| Format: | Preprint |
| Published: |
2021
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MM-SHAP: A Performance-agnostic Metric for Measuring Multimodal Contributions in Vision and Language Models & Tasks
by: Parcalabescu, Letitia, et al.
Published: (2022)
by: Parcalabescu, Letitia, et al.
Published: (2022)
On Measuring Faithfulness or Self-consistency of Natural Language Explanations
by: Parcalabescu, Letitia, et al.
Published: (2023)
by: Parcalabescu, Letitia, et al.
Published: (2023)
Do Vision & Language Decoders use Images and Text equally? How Self-consistent are their Explanations?
by: Parcalabescu, Letitia, et al.
Published: (2024)
by: Parcalabescu, Letitia, et al.
Published: (2024)
When Audio Generators Become Good Listeners: Generative Features for Understanding Tasks
by: Xie, Zeyu, et al.
Published: (2025)
by: Xie, Zeyu, et al.
Published: (2025)
Bounding Hallucinations: Information-Theoretic Guarantees for RAG Systems via Merlin-Arthur Protocols
by: Deiseroth, Björn, et al.
Published: (2025)
by: Deiseroth, Björn, et al.
Published: (2025)
SemanticVocoder: Bridging Audio Generation and Audio Understanding via Semantic Latents
by: Xie, Zeyu, et al.
Published: (2026)
by: Xie, Zeyu, et al.
Published: (2026)
Verifiable Dropout: Turning Randomness into a Verifiable Claim
by: Lee, Kichang, et al.
Published: (2025)
by: Lee, Kichang, et al.
Published: (2025)
PairHuman: A High-Fidelity Photographic Dataset for Customized Dual-Person Generation
by: Pan, Ting, et al.
Published: (2025)
by: Pan, Ting, et al.
Published: (2025)
AudioTime: A Temporally-aligned Audio-text Benchmark Dataset
by: Xie, Zeyu, et al.
Published: (2024)
by: Xie, Zeyu, et al.
Published: (2024)
FakeSound2: A Benchmark for Explainable and Generalizable Deepfake Sound Detection
by: Xie, Zeyu, et al.
Published: (2025)
by: Xie, Zeyu, et al.
Published: (2025)
Distributional Drift Adaptation with Temporal Conditional Variational Autoencoder for Multivariate Time Series Forecasting
by: He, Hui, et al.
Published: (2022)
by: He, Hui, et al.
Published: (2022)
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
by: Raoufi, Behnam, et al.
Published: (2025)
by: Raoufi, Behnam, et al.
Published: (2025)
STAR: Speech-to-Audio Generation via Representation Learning
by: Xie, Zeyu, et al.
Published: (2025)
by: Xie, Zeyu, et al.
Published: (2025)
PicoAudio: Enabling Precise Timestamp and Frequency Controllability of Audio Events in Text-to-audio Generation
by: Xie, Zeyu, et al.
Published: (2024)
by: Xie, Zeyu, et al.
Published: (2024)
Contemporary Agent Technology: LLM-Driven Advancements vs Classic Multi-Agent Systems
by: Bădică, Costin, et al.
Published: (2025)
by: Bădică, Costin, et al.
Published: (2025)
CAST-TTS: A Simple Cross-Attention Framework for Unified Timbre Control in TTS
by: Zheng, Zihao, et al.
Published: (2026)
by: Zheng, Zihao, et al.
Published: (2026)
PicoAudio2: Temporal Controllable Text-to-Audio Generation with Natural Language Description
by: Zheng, Zihao, et al.
Published: (2025)
by: Zheng, Zihao, et al.
Published: (2025)
FakeSound: Deepfake General Audio Detection
by: Xie, Zeyu, et al.
Published: (2024)
by: Xie, Zeyu, et al.
Published: (2024)
Toward Storage-Aware Learning with Compressed Data An Empirical Exploratory Study on JPEG
by: Lee, Kichang, et al.
Published: (2025)
by: Lee, Kichang, et al.
Published: (2025)
VCRScore: Image captioning metric based on V\&L Transformers, CLIP, and precision-recall
by: Ruiz, Guillermo, et al.
Published: (2025)
by: Ruiz, Guillermo, et al.
Published: (2025)
Robust Multivariate Time Series Forecasting against Intra- and Inter-Series Transitional Shift
by: He, Hui, et al.
Published: (2024)
by: He, Hui, et al.
Published: (2024)
AI-Driven Innovations in Modern Cloud Computing
by: Kumar, Animesh
Published: (2024)
by: Kumar, Animesh
Published: (2024)
Enabling Trustworthy Federated Learning in Industrial IoT: Bridging the Gap Between Interpretability and Robustness
by: Jagatheesaperumal, Senthil Kumar, et al.
Published: (2024)
by: Jagatheesaperumal, Senthil Kumar, et al.
Published: (2024)
Redefining Finance: The Influence of Artificial Intelligence (AI) and Machine Learning (ML)
by: Kumar, Animesh
Published: (2024)
by: Kumar, Animesh
Published: (2024)
Spatially-Aware Speaker for Vision-and-Language Navigation Instruction Generation
by: Gopinathan, Muraleekrishna, et al.
Published: (2024)
by: Gopinathan, Muraleekrishna, et al.
Published: (2024)
Multimodal sensor fusion in the latent representation space
by: Piechocki, Robert J., et al.
Published: (2022)
by: Piechocki, Robert J., et al.
Published: (2022)
Dynamic Cooperative Strategies in Search Engine Advertising Market: With and Without Retail Competition
by: Li, Huiran, et al.
Published: (2025)
by: Li, Huiran, et al.
Published: (2025)
PhysicsArena: The First Multimodal Physics Reasoning Benchmark Exploring Variable, Process, and Solution Dimensions
by: Dai, Song, et al.
Published: (2025)
by: Dai, Song, et al.
Published: (2025)
Dance of the ADS: Orchestrating Failures through Historically-Informed Scenario Fuzzing
by: Wang, Tong, et al.
Published: (2024)
by: Wang, Tong, et al.
Published: (2024)
Learning the meanings of function words from grounded language using a visual question answering model
by: Portelance, Eva, et al.
Published: (2023)
by: Portelance, Eva, et al.
Published: (2023)
RoNFA: Robust Neural Field-based Approach for Few-Shot Image Classification with Noisy Labels
by: Xiang, Nan, et al.
Published: (2025)
by: Xiang, Nan, et al.
Published: (2025)
Methods to integrate multinormals and compute classification measures
by: Das, Abhranil, et al.
Published: (2020)
by: Das, Abhranil, et al.
Published: (2020)
Superior Scoring Rules for Probabilistic Evaluation of Single-Label Multi-Class Classification Tasks
by: Ahmadian, Rouhollah, et al.
Published: (2024)
by: Ahmadian, Rouhollah, et al.
Published: (2024)
U-Net-Like Spiking Neural Networks for Single Image Dehazing
by: Li, Huibin, et al.
Published: (2025)
by: Li, Huibin, et al.
Published: (2025)
Model-Based Soft Maximization of Suitable Metrics of Long-Term Human Power
by: Heitzig, Jobst, et al.
Published: (2025)
by: Heitzig, Jobst, et al.
Published: (2025)
Conversion rate prediction in online advertising: modeling techniques, performance evaluation and future directions
by: Xue, Tao, et al.
Published: (2025)
by: Xue, Tao, et al.
Published: (2025)
Toward a benchmark for CTR prediction in online advertising: datasets, evaluation protocols and perspectives
by: Gao, Shan, et al.
Published: (2025)
by: Gao, Shan, et al.
Published: (2025)
Temperature Scaling Attack Disrupting Model Confidence in Federated Learning
by: Lee, Kichang, et al.
Published: (2026)
by: Lee, Kichang, et al.
Published: (2026)
Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning
by: Tong, Jingqi, et al.
Published: (2025)
by: Tong, Jingqi, et al.
Published: (2025)
ICG: Improving Cover Image Generation via MLLM-based Prompting and Personalized Preference Alignment
by: Bian, Zhipeng, et al.
Published: (2026)
by: Bian, Zhipeng, et al.
Published: (2026)
Similar Items
-
MM-SHAP: A Performance-agnostic Metric for Measuring Multimodal Contributions in Vision and Language Models & Tasks
by: Parcalabescu, Letitia, et al.
Published: (2022) -
On Measuring Faithfulness or Self-consistency of Natural Language Explanations
by: Parcalabescu, Letitia, et al.
Published: (2023) -
Do Vision & Language Decoders use Images and Text equally? How Self-consistent are their Explanations?
by: Parcalabescu, Letitia, et al.
Published: (2024) -
When Audio Generators Become Good Listeners: Generative Features for Understanding Tasks
by: Xie, Zeyu, et al.
Published: (2025) -
Bounding Hallucinations: Information-Theoretic Guarantees for RAG Systems via Merlin-Arthur Protocols
by: Deiseroth, Björn, et al.
Published: (2025)