GSM-SEM: Benchmark and Framework for Generating Semantically Variant Augmentations
Fuente:
arXiv
Salvato in:
| Autori principali: | Singh, Jyotika, Tu, Fang, Mirsaidova, Aziza, Agarwal, Amit, Patel, Hitesh Laxmichand, Ghoshal, Sandip, Ballesteros, Miguel, Dua, Karan, Benajiba, Yassine, Sun, Weiyi, Sheng, Tao, Horwood, Graham, Ravi, Sujith, Roth, Dan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MT-OSC: Path for LLMs that Get Lost in Multi-Turn Conversation
di: Singh, Jyotika, et al.
Pubblicazione: (2026)
di: Singh, Jyotika, et al.
Pubblicazione: (2026)
Can LLMs Narrate Tabular Data? An Evaluation Framework for Natural Language Representations of Text-to-SQL System Outputs
di: Singh, Jyotika, et al.
Pubblicazione: (2025)
di: Singh, Jyotika, et al.
Pubblicazione: (2025)
RCI: A Score for Evaluating Global and Local Reasoning in Multimodal Benchmarks
di: Agarwal, Amit, et al.
Pubblicazione: (2025)
di: Agarwal, Amit, et al.
Pubblicazione: (2025)
JTPRO: A Joint Tool-Prompt Reflective Optimization Framework for Language Agents
di: Ghoshal, Sandip, et al.
Pubblicazione: (2026)
di: Ghoshal, Sandip, et al.
Pubblicazione: (2026)
PCRI: Measuring Context Robustness in Multimodal Models for Enterprise Applications
di: Patel, Hitesh Laxmichand, et al.
Pubblicazione: (2025)
di: Patel, Hitesh Laxmichand, et al.
Pubblicazione: (2025)
Do Image-Text Metrics Respect Semantic Invariances?
di: Agarwal, Amit, et al.
Pubblicazione: (2026)
di: Agarwal, Amit, et al.
Pubblicazione: (2026)
Aligning LLMs for Multilingual Consistency in Enterprise Applications
di: Agarwal, Amit, et al.
Pubblicazione: (2025)
di: Agarwal, Amit, et al.
Pubblicazione: (2025)
DiffuMask: Diffusion Language Model for Token-level Prompt Pruning
di: Zheng, Caleb, et al.
Pubblicazione: (2026)
di: Zheng, Caleb, et al.
Pubblicazione: (2026)
SPENCE: A Syntactic Probe for Detecting Contamination in NL2SQL Benchmarks
di: Safarzadeh, Mohammadtaher, et al.
Pubblicazione: (2026)
di: Safarzadeh, Mohammadtaher, et al.
Pubblicazione: (2026)
Barriers to Discrete Reasoning with Transformers: A Survey Across Depth, Exactness, and Bandwidth
di: Yuan, Michelle, et al.
Pubblicazione: (2026)
di: Yuan, Michelle, et al.
Pubblicazione: (2026)
FlexDoc: Parameterized Sampling for Diverse Multilingual Synthetic Documents for Training Document Understanding Models
di: Dua, Karan, et al.
Pubblicazione: (2025)
di: Dua, Karan, et al.
Pubblicazione: (2025)
SpeechWeave: Diverse Multilingual Synthetic Text & Audio Data Generation Pipeline for Training Text to Speech Models
di: Dua, Karan, et al.
Pubblicazione: (2025)
di: Dua, Karan, et al.
Pubblicazione: (2025)
Active Evaluation Acquisition for Efficient LLM Benchmarking
di: Li, Yang, et al.
Pubblicazione: (2024)
di: Li, Yang, et al.
Pubblicazione: (2024)
AccessEval: Benchmarking Disability Bias in Large Language Models
di: Panda, Srikant, et al.
Pubblicazione: (2025)
di: Panda, Srikant, et al.
Pubblicazione: (2025)
Tokenization Matters: Improving Zero-Shot NER for Indic Languages
di: Pattnayak, Priyaranjan, et al.
Pubblicazione: (2025)
di: Pattnayak, Priyaranjan, et al.
Pubblicazione: (2025)
LLM for Barcodes: Generating Diverse Synthetic Data for Identity Documents
di: Patel, Hitesh Laxmichand, et al.
Pubblicazione: (2024)
di: Patel, Hitesh Laxmichand, et al.
Pubblicazione: (2024)
LLM-Guided Lifecycle-Aware Clustering of Multi-Turn Customer Support Conversations
di: Pattnayak, Priyaranjan, et al.
Pubblicazione: (2026)
di: Pattnayak, Priyaranjan, et al.
Pubblicazione: (2026)
Hybrid AI for Responsive Multi-Turn Online Conversations with Novel Dynamic Routing and Feedback Adaptation
di: Pattnayak, Priyaranjan, et al.
Pubblicazione: (2025)
di: Pattnayak, Priyaranjan, et al.
Pubblicazione: (2025)
Hard Negative Mining for Domain-Specific Retrieval in Enterprise Systems
di: Meghwani, Hansa, et al.
Pubblicazione: (2025)
di: Meghwani, Hansa, et al.
Pubblicazione: (2025)
Who's Asking? Investigating Bias Through the Lens of Disability Framed Queries in LLMs
di: Hari, Vishnu, et al.
Pubblicazione: (2025)
di: Hari, Vishnu, et al.
Pubblicazione: (2025)
RECOR: Reasoning-focused Multi-turn Conversational Retrieval Benchmark
di: Ali, Mohammed, et al.
Pubblicazione: (2026)
di: Ali, Mohammed, et al.
Pubblicazione: (2026)
Robust Audio-Text Retrieval via Cross-Modal Attention and Hybrid Loss
di: Liu, Meizhu, et al.
Pubblicazione: (2026)
di: Liu, Meizhu, et al.
Pubblicazione: (2026)
Arabic Named Entity Recognition
di: Yassine Benajiba
Pubblicazione: (2010)
di: Yassine Benajiba
Pubblicazione: (2010)
Think Twice Before You Write -- an Entropy-based Decoding Strategy to Enhance LLM Reasoning
di: He, Jiashu, et al.
Pubblicazione: (2026)
di: He, Jiashu, et al.
Pubblicazione: (2026)
Clinical QA 2.0: Multi-Task Learning for Answer Extraction and Categorization
di: Pattnayak, Priyaranjan, et al.
Pubblicazione: (2025)
di: Pattnayak, Priyaranjan, et al.
Pubblicazione: (2025)
Survey of Large Multimodal Model Datasets, Application Categories and Taxonomy
di: Pattnayak, Priyaranjan, et al.
Pubblicazione: (2024)
di: Pattnayak, Priyaranjan, et al.
Pubblicazione: (2024)
Budget-Aware Anytime Reasoning with LLM-Synthesized Preference Data
di: Zhang, Xuanming, et al.
Pubblicazione: (2026)
di: Zhang, Xuanming, et al.
Pubblicazione: (2026)
DAIQ: Auditing Demographic Attribute Inference from Question in LLMs
di: Panda, Srikant, et al.
Pubblicazione: (2025)
di: Panda, Srikant, et al.
Pubblicazione: (2025)
Judging What We Cannot Solve: A Consequence-Based Approach for Oracle-Free Evaluation of Research-Level Math
di: Son, Guijin, et al.
Pubblicazione: (2026)
di: Son, Guijin, et al.
Pubblicazione: (2026)
MetaSynth: Meta-Prompting-Driven Agentic Scaffolds for Diverse Synthetic Data Generation
di: Riaz, Haris, et al.
Pubblicazione: (2025)
di: Riaz, Haris, et al.
Pubblicazione: (2025)
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use
di: Patel, Hitesh Laxmichand, et al.
Pubblicazione: (2025)
di: Patel, Hitesh Laxmichand, et al.
Pubblicazione: (2025)
Open Agent Specification (Agent Spec): A Unified Representation for AI Agents
di: Amini, Soufiane, et al.
Pubblicazione: (2025)
di: Amini, Soufiane, et al.
Pubblicazione: (2025)
Self-supervised Analogical Learning using Language Models
di: Zhou, Ben, et al.
Pubblicazione: (2025)
di: Zhou, Ben, et al.
Pubblicazione: (2025)
INGLIZ VA O'ZBEK TILLARIDA JOY NOMLARINING LINGVOKULTUROLOGIK XUSUSIYATLARI
di: O.S. Ahmedov, N.M. Mirsaidova
Pubblicazione: (2024)
di: O.S. Ahmedov, N.M. Mirsaidova
Pubblicazione: (2024)
Enhancing Document AI Data Generation Through Graph-Based Synthetic Layouts
di: Agarwal, Amit, et al.
Pubblicazione: (2024)
di: Agarwal, Amit, et al.
Pubblicazione: (2024)
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation
di: Kim, Eunsu, et al.
Pubblicazione: (2025)
di: Kim, Eunsu, et al.
Pubblicazione: (2025)
Post-Vaccination COVID-19 Data Analysis: Privacy and Ethics
di: Das, Sankha, et al.
Pubblicazione: (2024)
di: Das, Sankha, et al.
Pubblicazione: (2024)
Continuous Spiking Graph Neural Networks
di: Yin, Nan, et al.
Pubblicazione: (2024)
di: Yin, Nan, et al.
Pubblicazione: (2024)
Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge
di: Spiliopoulou, Evangelia, et al.
Pubblicazione: (2025)
di: Spiliopoulou, Evangelia, et al.
Pubblicazione: (2025)
Pushing on Multilingual Reasoning Models with Language-Mixed Chain-of-Thought
di: Son, Guijin, et al.
Pubblicazione: (2025)
di: Son, Guijin, et al.
Pubblicazione: (2025)
Documenti analoghi
-
MT-OSC: Path for LLMs that Get Lost in Multi-Turn Conversation
di: Singh, Jyotika, et al.
Pubblicazione: (2026) -
Can LLMs Narrate Tabular Data? An Evaluation Framework for Natural Language Representations of Text-to-SQL System Outputs
di: Singh, Jyotika, et al.
Pubblicazione: (2025) -
RCI: A Score for Evaluating Global and Local Reasoning in Multimodal Benchmarks
di: Agarwal, Amit, et al.
Pubblicazione: (2025) -
JTPRO: A Joint Tool-Prompt Reflective Optimization Framework for Language Agents
di: Ghoshal, Sandip, et al.
Pubblicazione: (2026) -
PCRI: Measuring Context Robustness in Multimodal Models for Enterprise Applications
di: Patel, Hitesh Laxmichand, et al.
Pubblicazione: (2025)