IntentGrasp: A Comprehensive Benchmark for Intent Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Yin, Yuwei, Li, Chuyuan, Carenini, Giuseppe |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SWI: Speaking with Intent in Large Language Models
by: Yin, Yuwei, et al.
Published: (2025)
by: Yin, Yuwei, et al.
Published: (2025)
Improving Language Models with Intentional Analysis
by: Yin, Yuwei, et al.
Published: (2025)
by: Yin, Yuwei, et al.
Published: (2025)
Neural Multimodal Topic Modeling: A Comprehensive Evaluation
by: González-Pizarro, Felipe, et al.
Published: (2024)
by: González-Pizarro, Felipe, et al.
Published: (2024)
Harmful Intent as a Geometrically Recoverable Feature of LLM Residual Streams
by: Llorente-Saguer, Isaac
Published: (2026)
by: Llorente-Saguer, Isaac
Published: (2026)
The Geometry of Harmful Intent: Training-Free Anomaly Detection via Angular Deviation in LLM Residual Streams
by: Llorente-Saguer, Isaac
Published: (2026)
by: Llorente-Saguer, Isaac
Published: (2026)
Measuring Intent Comprehension in LLMs
by: Kunievsky, Nadav, et al.
Published: (2025)
by: Kunievsky, Nadav, et al.
Published: (2025)
The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants
by: Bandarkar, Lucas, et al.
Published: (2023)
by: Bandarkar, Lucas, et al.
Published: (2023)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
by: Oketunji, Abiodun Finbarrs
Published: (2023)
by: Oketunji, Abiodun Finbarrs
Published: (2023)
Knowledge Graph Embeddings: A Comprehensive Survey on Capturing Relation Properties
by: Niu, Guanglin
Published: (2024)
by: Niu, Guanglin
Published: (2024)
Positional Failures in Long-Context LLMs: A Blind Spot in Reasoning Benchmarks
by: Zhang, Chuyifei, et al.
Published: (2026)
by: Zhang, Chuyifei, et al.
Published: (2026)
Mathador-LM: A Dynamic Benchmark for Mathematical Reasoning on Large Language Models
by: Kurtic, Eldar, et al.
Published: (2024)
by: Kurtic, Eldar, et al.
Published: (2024)
IdentifyMe: A Challenging Long-Context Mention Resolution Benchmark for LLMs
by: Manikantan, Kawshik, et al.
Published: (2024)
by: Manikantan, Kawshik, et al.
Published: (2024)
Large Language Model (LLM) Bias Index -- LLMBI
by: Oketunji, Abiodun Finbarrs, et al.
Published: (2023)
by: Oketunji, Abiodun Finbarrs, et al.
Published: (2023)
ALBA: A European Portuguese Benchmark for Evaluating Language and Linguistic Dimensions in Generative LLMs
by: Vieira, Inês, et al.
Published: (2026)
by: Vieira, Inês, et al.
Published: (2026)
Benchmarking Cognitive Biases in Large Language Models as Evaluators
by: Koo, Ryan, et al.
Published: (2023)
by: Koo, Ryan, et al.
Published: (2023)
Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy
by: Lan, Guangchen, et al.
Published: (2026)
by: Lan, Guangchen, et al.
Published: (2026)
EasyMath: A 0-shot Math Benchmark for SLMs
by: Karki, Drishya, et al.
Published: (2025)
by: Karki, Drishya, et al.
Published: (2025)
Cognitive Load Limits in Large Language Models: Benchmarking Multi-Hop Reasoning
by: Adapala, Sai Teja Reddy
Published: (2025)
by: Adapala, Sai Teja Reddy
Published: (2025)
Diagnosing and Addressing Pitfalls in KG-RAG Datasets: Toward More Reliable Benchmarking
by: Zhang, Liangliang, et al.
Published: (2025)
by: Zhang, Liangliang, et al.
Published: (2025)
MaPPO: Maximum a Posteriori Preference Optimization with Prior Knowledge
by: Lan, Guangchen, et al.
Published: (2025)
by: Lan, Guangchen, et al.
Published: (2025)
PARAPHRASUS : A Comprehensive Benchmark for Evaluating Paraphrase Detection Models
by: Michail, Andrianos, et al.
Published: (2024)
by: Michail, Andrianos, et al.
Published: (2024)
AMALIA Technical Report: A Fully Open Source Large Language Model for European Portuguese
by: Simplício, Afonso, et al.
Published: (2026)
by: Simplício, Afonso, et al.
Published: (2026)
RAC: Efficient LLM Factuality Correction with Retrieval Augmentation
by: Li, Changmao, et al.
Published: (2024)
by: Li, Changmao, et al.
Published: (2024)
ObfusQAte: A Proposed Framework to Evaluate LLM Robustness on Obfuscated Factual Question Answering
by: Ghosh, Shubhra, et al.
Published: (2025)
by: Ghosh, Shubhra, et al.
Published: (2025)
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
by: Fadli, Samih
Published: (2025)
by: Fadli, Samih
Published: (2025)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
by: Peters, Sydney, et al.
Published: (2025)
by: Peters, Sydney, et al.
Published: (2025)
ProSparse: Introducing and Enhancing Intrinsic Activation Sparsity within Large Language Models
by: Song, Chenyang, et al.
Published: (2024)
by: Song, Chenyang, et al.
Published: (2024)
Flash Multi-Head Feed-Forward Network
by: Zhang, Minshen, et al.
Published: (2025)
by: Zhang, Minshen, et al.
Published: (2025)
HAMMER: Hamiltonian Curiosity Augmented Large Language Model Reinforcement
by: Yang, Ming, et al.
Published: (2025)
by: Yang, Ming, et al.
Published: (2025)
A Survey of the State of Explainable AI for Natural Language Processing
by: Danilevsky, Marina, et al.
Published: (2020)
by: Danilevsky, Marina, et al.
Published: (2020)
Major Entity Identification: A Generalizable Alternative to Coreference Resolution
by: Manikantan, Kawshik, et al.
Published: (2024)
by: Manikantan, Kawshik, et al.
Published: (2024)
Dealing with Annotator Disagreement in Hate Speech Classification
by: Dehghan, Somaiyeh, et al.
Published: (2025)
by: Dehghan, Somaiyeh, et al.
Published: (2025)
Improving Discrete Diffusion Unmasking Policies Beyond Explicit Reference Policies
by: Hong, Chunsan, et al.
Published: (2025)
by: Hong, Chunsan, et al.
Published: (2025)
Three Regimes of Context-Parametric Conflict: A Predictive Framework and Empirical Validation
by: Venkata, Pruthvinath Jeripity
Published: (2026)
by: Venkata, Pruthvinath Jeripity
Published: (2026)
Detecting Hallucinations in Large Language Model Generation: A Token Probability Approach
by: Quevedo, Ernesto, et al.
Published: (2024)
by: Quevedo, Ernesto, et al.
Published: (2024)
Path-Lock Expert: Separating Reasoning Mode in Hybrid Thinking via Architecture-Level Separation
by: Wang, Shouren, et al.
Published: (2026)
by: Wang, Shouren, et al.
Published: (2026)
Every Sample Matters: Leveraging Mixture-of-Experts and High-Quality Data for Efficient and Accurate Code LLM
by: Codefuse, et al.
Published: (2025)
by: Codefuse, et al.
Published: (2025)
Towards Intrinsic Interpretability of Large Language Models:A Survey of Design Principles and Architectures
by: Gao, Yutong, et al.
Published: (2026)
by: Gao, Yutong, et al.
Published: (2026)
Learning from Supervision with Semantic and Episodic Memory: A Reflective Approach to Agent Adaptation
by: Hassell, Jackson, et al.
Published: (2025)
by: Hassell, Jackson, et al.
Published: (2025)
SODA: A Natural Language Processing Package to Extract Social Determinants of Health for Cancer Studies
by: Yu, Zehao, et al.
Published: (2022)
by: Yu, Zehao, et al.
Published: (2022)
Similar Items
-
SWI: Speaking with Intent in Large Language Models
by: Yin, Yuwei, et al.
Published: (2025) -
Improving Language Models with Intentional Analysis
by: Yin, Yuwei, et al.
Published: (2025) -
Neural Multimodal Topic Modeling: A Comprehensive Evaluation
by: González-Pizarro, Felipe, et al.
Published: (2024) -
Harmful Intent as a Geometrically Recoverable Feature of LLM Residual Streams
by: Llorente-Saguer, Isaac
Published: (2026) -
The Geometry of Harmful Intent: Training-Free Anomaly Detection via Angular Deviation in LLM Residual Streams
by: Llorente-Saguer, Isaac
Published: (2026)