VISLA Benchmark: Evaluating Embedding Sensitivity to Semantic and Lexical Alterations
Fuente:
arXiv
Saved in:
| Main Authors: | Dumpala, Sri Harsha, Jaiswal, Aman, Sastry, Chandramouli, Milios, Evangelos, Oore, Sageev, Sajjad, Hassan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SUGARCREPE++ Dataset: Vision-Language Model Sensitivity to Semantic and Lexical Alterations
by: Dumpala, Sri Harsha, et al.
Published: (2024)
by: Dumpala, Sri Harsha, et al.
Published: (2024)
Sensitivity of Generative VLMs to Semantically and Lexically Altered Prompts
by: Dumpala, Sri Harsha, et al.
Published: (2024)
by: Dumpala, Sri Harsha, et al.
Published: (2024)
DiffAug: A Diffuse-and-Denoise Augmentation for Training Robust Classifiers
by: Sastry, Chandramouli, et al.
Published: (2023)
by: Sastry, Chandramouli, et al.
Published: (2023)
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models
by: Dumpala, Sri Harsha, et al.
Published: (2024)
by: Dumpala, Sri Harsha, et al.
Published: (2024)
Test-Time Training for Depression Detection
by: Dumpala, Sri Harsha, et al.
Published: (2024)
by: Dumpala, Sri Harsha, et al.
Published: (2024)
Predicting Individual Depression Symptoms from Acoustic Features During Speech
by: Rodriguez, Sebastian, et al.
Published: (2024)
by: Rodriguez, Sebastian, et al.
Published: (2024)
Exploring the features used for summary evaluation by Human and GPT
by: Sadeghi, Zahra, et al.
Published: (2025)
by: Sadeghi, Zahra, et al.
Published: (2025)
Stance Reasoner: Zero-Shot Stance Detection on Social Media with Explicit Reasoning
by: Taranukhin, Maksym, et al.
Published: (2024)
by: Taranukhin, Maksym, et al.
Published: (2024)
Self-Supervised Embeddings for Detecting Individual Symptoms of Depression
by: Dumpala, Sri Harsha, et al.
Published: (2024)
by: Dumpala, Sri Harsha, et al.
Published: (2024)
Empowering Air Travelers: A Chatbot for Canadian Air Passenger Rights
by: Taranukhin, Maksym, et al.
Published: (2024)
by: Taranukhin, Maksym, et al.
Published: (2024)
Same Meaning, Different Scores: Lexical and Syntactic Sensitivity in LLM Evaluation
by: Kostić, Bogdan, et al.
Published: (2026)
by: Kostić, Bogdan, et al.
Published: (2026)
Unveiling the Lexical Sensitivity of LLMs: Combinatorial Optimization for Prompt Enhancement
by: Zhan, Pengwei, et al.
Published: (2024)
by: Zhan, Pengwei, et al.
Published: (2024)
The Structured Output Benchmark: A Multi-Source Benchmark for Evaluating Structured Output Quality in Large Language Models
by: Singh, Abhinav Kumar, et al.
Published: (2026)
by: Singh, Abhinav Kumar, et al.
Published: (2026)
SMILE: A Composite Lexical-Semantic Metric for Question-Answering Evaluation
by: Kendre, Shrikant, et al.
Published: (2025)
by: Kendre, Shrikant, et al.
Published: (2025)
Unsupervised Candidate Ranking for Lexical Substitution via Holistic Sentence Semantics
by: Hu, Zhongyang, et al.
Published: (2025)
by: Hu, Zhongyang, et al.
Published: (2025)
LGSE: Lexically Grounded Subword Embedding Initialization for Low-Resource Language Adaptation
by: Teklehaymanot, Hailay, et al.
Published: (2026)
by: Teklehaymanot, Hailay, et al.
Published: (2026)
Between Century and Poet: Graph-Based Lexical Semantic Change in Persian Poetry
by: Shahnazari, Kourosh, et al.
Published: (2026)
by: Shahnazari, Kourosh, et al.
Published: (2026)
Interpreting the Effects of Quantization on LLMs
by: Singh, Manpreet, et al.
Published: (2025)
by: Singh, Manpreet, et al.
Published: (2025)
Quantifying the Capabilities of LLMs across Scale and Precision
by: Badshah, Sher, et al.
Published: (2024)
by: Badshah, Sher, et al.
Published: (2024)
TALE: A Tool-Augmented Framework for Reference-Free Evaluation of Large Language Models
by: Badshah, Sher, et al.
Published: (2025)
by: Badshah, Sher, et al.
Published: (2025)
GraphLSS: Integrating Lexical, Structural, and Semantic Features for Long Document Extractive Summarization
by: Bugueño, Margarita, et al.
Published: (2024)
by: Bugueño, Margarita, et al.
Published: (2024)
PragWorld: A Benchmark Evaluating LLMs' Local World Model under Minimal Linguistic Alterations and Conversational Dynamics
by: Vashistha, Sachin, et al.
Published: (2025)
by: Vashistha, Sachin, et al.
Published: (2025)
SiLVERScore: Semantically-Aware Embeddings for Sign Language Generation Evaluation
by: Imai, Saki, et al.
Published: (2025)
by: Imai, Saki, et al.
Published: (2025)
Word Sense Disambiguation in Native Spanish: A Comprehensive Lexical Evaluation Resource
by: Ortega, Pablo, et al.
Published: (2024)
by: Ortega, Pablo, et al.
Published: (2024)
LexInstructEval: Lexical Instruction Following Evaluation for Large Language Models
by: Ren, Huimin, et al.
Published: (2025)
by: Ren, Huimin, et al.
Published: (2025)
Enhancing ASR Performance in the Medical Domain for Dravidian Languages
by: Devarakonda, Sri Charan, et al.
Published: (2026)
by: Devarakonda, Sri Charan, et al.
Published: (2026)
LLMs Underperform Graph-Based Parsers on Supervised Relation Extraction for Complex Graphs
by: Gajo, Paolo, et al.
Published: (2026)
by: Gajo, Paolo, et al.
Published: (2026)
Spectral Tempering for Embedding Compression in Dense Passage Retrieval
by: Li, Yongkang, et al.
Published: (2026)
by: Li, Yongkang, et al.
Published: (2026)
ROSA: Random Subspace Adaptation for Efficient Fine-Tuning
by: Hameed, Marawan Gamal Abdel, et al.
Published: (2024)
by: Hameed, Marawan Gamal Abdel, et al.
Published: (2024)
BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation
by: Gisserot-Boukhlef, Hippolyte, et al.
Published: (2026)
by: Gisserot-Boukhlef, Hippolyte, et al.
Published: (2026)
DALDALL: Data Augmentation for Lexical and Semantic Diverse in Legal Domain by leveraging LLM-Persona
by: Choi, Janghyeok, et al.
Published: (2026)
by: Choi, Janghyeok, et al.
Published: (2026)
SaudiCulture: A Benchmark for Evaluating Large Language Models Cultural Competence within Saudi Arabia
by: Ayash, Lama, et al.
Published: (2025)
by: Ayash, Lama, et al.
Published: (2025)
German Text Embedding Clustering Benchmark
by: Wehrli, Silvan, et al.
Published: (2024)
by: Wehrli, Silvan, et al.
Published: (2024)
The Scandinavian Embedding Benchmarks: Comprehensive Assessment of Multilingual and Monolingual Text Embedding
by: Enevoldsen, Kenneth, et al.
Published: (2024)
by: Enevoldsen, Kenneth, et al.
Published: (2024)
Semantic Structure in Large Language Model Embeddings
by: Kozlowski, Austin C., et al.
Published: (2025)
by: Kozlowski, Austin C., et al.
Published: (2025)
Using Language Models to Disambiguate Lexical Choices in Translation
by: Barua, Josh, et al.
Published: (2024)
by: Barua, Josh, et al.
Published: (2024)
Polysemanticity or Polysemy? Lexical Identity Confounds Superposition Metrics
by: Hou, Iyad Ait, et al.
Published: (2026)
by: Hou, Iyad Ait, et al.
Published: (2026)
Language Model Re-rankers are Fooled by Lexical Similarities
by: Hagström, Lovisa, et al.
Published: (2025)
by: Hagström, Lovisa, et al.
Published: (2025)
A Systematic Investigation of Document Chunking Strategies and Embedding Sensitivity
by: Shaukat, Muhammad Arslan, et al.
Published: (2026)
by: Shaukat, Muhammad Arslan, et al.
Published: (2026)
Understanding AI Evaluation Patterns: How Different GPT Models Assess Vision-Language Descriptions
by: Abdoli, Sajjad, et al.
Published: (2025)
by: Abdoli, Sajjad, et al.
Published: (2025)
Similar Items
-
SUGARCREPE++ Dataset: Vision-Language Model Sensitivity to Semantic and Lexical Alterations
by: Dumpala, Sri Harsha, et al.
Published: (2024) -
Sensitivity of Generative VLMs to Semantically and Lexically Altered Prompts
by: Dumpala, Sri Harsha, et al.
Published: (2024) -
DiffAug: A Diffuse-and-Denoise Augmentation for Training Robust Classifiers
by: Sastry, Chandramouli, et al.
Published: (2023) -
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models
by: Dumpala, Sri Harsha, et al.
Published: (2024) -
Test-Time Training for Depression Detection
by: Dumpala, Sri Harsha, et al.
Published: (2024)