Evaluating Concurrent Robustness of Language Models Across Diverse Challenge Sets
Fuente:
arXiv
Saved in:
| Main Authors: | Gupta, Vatsal, Pandya, Pranshu, Kataria, Tushar, Gupta, Vivek, Roth, Dan |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
NTSEBENCH: Cognitive Reasoning Benchmark for Vision Language Models
by: Pandya, Pranshu, et al.
Published: (2024)
by: Pandya, Pranshu, et al.
Published: (2024)
FlowVQA: Mapping Multimodal Logic in Visual Question Answering with Flowcharts
by: Singh, Shubhankar, et al.
Published: (2024)
by: Singh, Shubhankar, et al.
Published: (2024)
CORE-T: COherent REtrieval of Tables for Text-to-SQL
by: Soliman, Hassan, et al.
Published: (2026)
by: Soliman, Hassan, et al.
Published: (2026)
TransientTables: Evaluating LLMs' Reasoning on Temporally Evolving Semi-structured Tables
by: Shankarampeta, Abhilash, et al.
Published: (2025)
by: Shankarampeta, Abhilash, et al.
Published: (2025)
Improving Robustness of Tabular Retrieval via Representational Stability
by: Bhandari, Kushal Raj, et al.
Published: (2026)
by: Bhandari, Kushal Raj, et al.
Published: (2026)
fact check AI at SemEval-2025 Task 7: Multilingual and Crosslingual Fact-checked Claim Retrieval
by: Rastogi, Pranshu
Published: (2025)
by: Rastogi, Pranshu
Published: (2025)
Evidence-Guided Schema Normalization for Temporal Tabular Reasoning
by: Thanga, Ashish, et al.
Published: (2025)
by: Thanga, Ashish, et al.
Published: (2025)
Weaver: Interweaving SQL and LLM for Table Reasoning
by: Khoja, Rohit, et al.
Published: (2025)
by: Khoja, Rohit, et al.
Published: (2025)
DENSE: Longitudinal Progress Note Generation with Temporal Modeling of Heterogeneous Clinical Notes Across Hospital Visits
by: Keerthana, Garapati, et al.
Published: (2025)
by: Keerthana, Garapati, et al.
Published: (2025)
DocQAC: Adaptive Trie-Guided Decoding for Effective In-Document Query Auto-Completion
by: Mehta, Rahul, et al.
Published: (2026)
by: Mehta, Rahul, et al.
Published: (2026)
Med-CoDE: Medical Critique based Disagreement Evaluation Framework
by: Gupta, Mohit, et al.
Published: (2025)
by: Gupta, Mohit, et al.
Published: (2025)
Model Editing at Scale leads to Gradual and Catastrophic Forgetting
by: Gupta, Akshat, et al.
Published: (2024)
by: Gupta, Akshat, et al.
Published: (2024)
Is Table Retrieval a Solved Problem? Exploring Join-Aware Multi-Table Retrieval
by: Chen, Peter Baile, et al.
Published: (2024)
by: Chen, Peter Baile, et al.
Published: (2024)
CLI-RAG: A Retrieval-Augmented Framework for Clinically Structured and Context Aware Text Generation with LLMs
by: Keerthana, Garapati, et al.
Published: (2025)
by: Keerthana, Garapati, et al.
Published: (2025)
Statements: Universal Information Extraction from Tables with Large Language Models for ESG KPIs
by: Mishra, Lokesh, et al.
Published: (2024)
by: Mishra, Lokesh, et al.
Published: (2024)
Can we Retrieve Everything All at Once? ARM: An Alignment-Oriented LLM-based Retrieval Method
by: Chen, Peter Baile, et al.
Published: (2025)
by: Chen, Peter Baile, et al.
Published: (2025)
EnrichIndex: Using LLMs to Enrich Retrieval Indices Offline
by: Chen, Peter Baile, et al.
Published: (2025)
by: Chen, Peter Baile, et al.
Published: (2025)
Multilingual Information Retrieval with a Monolingual Knowledge Base
by: Zhuang, Yingying, et al.
Published: (2025)
by: Zhuang, Yingying, et al.
Published: (2025)
Hierarchical Resolution Transformers: A Wavelet-Inspired Architecture for Multi-Scale Language Understanding
by: Sar, Ayan, et al.
Published: (2025)
by: Sar, Ayan, et al.
Published: (2025)
A Semantic Search Pipeline for Causality-driven Adhoc Information Retrieval
by: Dalal, Dhairya, et al.
Published: (2025)
by: Dalal, Dhairya, et al.
Published: (2025)
A Comprehensive Survey of Retrieval-Augmented Generation (RAG): Evolution, Current Landscape and Future Directions
by: Gupta, Shailja, et al.
Published: (2024)
by: Gupta, Shailja, et al.
Published: (2024)
When Large Language Models Meet Personalization: Perspectives of Challenges and Opportunities
by: Chen, Jin, et al.
Published: (2023)
by: Chen, Jin, et al.
Published: (2023)
CLARINET: Augmenting Language Models to Ask Clarification Questions for Retrieval
by: Chi, Yizhou, et al.
Published: (2024)
by: Chi, Yizhou, et al.
Published: (2024)
CiteEval: Principle-Driven Citation Evaluation for Source Attribution
by: Xu, Yumo, et al.
Published: (2025)
by: Xu, Yumo, et al.
Published: (2025)
Retrieval-Augmented Reasoning for Chartered Accountancy
by: Gupta, Jatin, et al.
Published: (2026)
by: Gupta, Jatin, et al.
Published: (2026)
Exploring the Implicit Semantic Ability of Multimodal Large Language Models: A Pilot Study on Entity Set Expansion
by: Wang, Hebin, et al.
Published: (2024)
by: Wang, Hebin, et al.
Published: (2024)
Efficient Evaluation of Large Language Models via Collaborative Filtering
by: Zhong, Xu-Xiang, et al.
Published: (2025)
by: Zhong, Xu-Xiang, et al.
Published: (2025)
Evaluating the External and Parametric Knowledge Fusion of Large Language Models
by: Zhang, Hao, et al.
Published: (2024)
by: Zhang, Hao, et al.
Published: (2024)
Beyond Words: Evaluating Large Language Models in Transportation Planning
by: Ying, Shaowei, et al.
Published: (2024)
by: Ying, Shaowei, et al.
Published: (2024)
Inducing Sustained Creativity and Diversity in Large Language Models
by: Luo, Queenie, et al.
Published: (2026)
by: Luo, Queenie, et al.
Published: (2026)
Leveraging LLMs for Synthesizing Training Data Across Many Languages in Multilingual Dense Retrieval
by: Thakur, Nandan, et al.
Published: (2023)
by: Thakur, Nandan, et al.
Published: (2023)
Building FKG.in: a Knowledge Graph for Indian Food
by: Gupta, Saransh Kumar, et al.
Published: (2024)
by: Gupta, Saransh Kumar, et al.
Published: (2024)
Evaluating Large Language Models as Generative User Simulators for Conversational Recommendation
by: Yoon, Se-eun, et al.
Published: (2024)
by: Yoon, Se-eun, et al.
Published: (2024)
Evaluating Multimodal Large Language Models on Educational Textbook Question Answering
by: Alawwad, Hessa A., et al.
Published: (2025)
by: Alawwad, Hessa A., et al.
Published: (2025)
Deep Learning Based Named Entity Recognition Models for Recipes
by: Goel, Mansi, et al.
Published: (2024)
by: Goel, Mansi, et al.
Published: (2024)
Evaluating Robustness of Generative Search Engine on Adversarial Factual Questions
by: Hu, Xuming, et al.
Published: (2024)
by: Hu, Xuming, et al.
Published: (2024)
Dynamic Reasoning Chains through Depth-Specialized Mixture-of-Experts in Transformer Architectures
by: Roy, Sampurna, et al.
Published: (2025)
by: Roy, Sampurna, et al.
Published: (2025)
Autofocus Retrieval: An Effective Pipeline for Multi-Hop Question Answering With Semi-Structured Knowledge
by: Boer, Derian, et al.
Published: (2025)
by: Boer, Derian, et al.
Published: (2025)
A Survey of LLM-based Deep Search Agents: Paradigm, Optimization, Evaluation, and Challenges
by: Xi, Yunjia, et al.
Published: (2025)
by: Xi, Yunjia, et al.
Published: (2025)
Enhancing FKG.in: automating Indian food composition analysis
by: Gupta, Saransh Kumar, et al.
Published: (2024)
by: Gupta, Saransh Kumar, et al.
Published: (2024)
Similar Items
-
NTSEBENCH: Cognitive Reasoning Benchmark for Vision Language Models
by: Pandya, Pranshu, et al.
Published: (2024) -
FlowVQA: Mapping Multimodal Logic in Visual Question Answering with Flowcharts
by: Singh, Shubhankar, et al.
Published: (2024) -
CORE-T: COherent REtrieval of Tables for Text-to-SQL
by: Soliman, Hassan, et al.
Published: (2026) -
TransientTables: Evaluating LLMs' Reasoning on Temporally Evolving Semi-structured Tables
by: Shankarampeta, Abhilash, et al.
Published: (2025) -
Improving Robustness of Tabular Retrieval via Representational Stability
by: Bhandari, Kushal Raj, et al.
Published: (2026)