What Makes a Good Query? Measuring the Impact of Human-Confusing Linguistic Features on LLM Performance
Fuente:
arXiv
Saved in:
| Main Authors: | Watson, William, Cho, Nicole, Ganesh, Sumitra, Veloso, Manuela |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
QueryBandits for Hallucination Mitigation: Exploiting Semantic Features for No-Regret Rewriting
by: Cho, Nicole, et al.
Published: (2025)
by: Cho, Nicole, et al.
Published: (2025)
No One Size Fits All: QueryBandits for Hallucination Mitigation
by: Cho, Nicole, et al.
Published: (2026)
by: Cho, Nicole, et al.
Published: (2026)
TASER: Table Agents for Schema-guided Extraction and Recommendation
by: Cho, Nicole, et al.
Published: (2025)
by: Cho, Nicole, et al.
Published: (2025)
HiddenTables & PyQTax: A Cooperative Game and Dataset For TableQA to Ensure Scale and Data Privacy Across a Myriad of Taxonomies
by: Watson, William, et al.
Published: (2024)
by: Watson, William, et al.
Published: (2024)
What Makes Good Multilingual Reasoning? Disentangling Reasoning Traces with Measurable Features
by: Ki, Dayeon, et al.
Published: (2026)
by: Ki, Dayeon, et al.
Published: (2026)
MultiQ&A: An Analysis in Measuring Robustness via Automated Crowdsourcing of Question Perturbations and Answers
by: Cho, Nicole, et al.
Published: (2025)
by: Cho, Nicole, et al.
Published: (2025)
FlowMind: Automatic Workflow Generation with LLMs
by: Zeng, Zhen, et al.
Published: (2024)
by: Zeng, Zhen, et al.
Published: (2024)
ChartAgent: A Multimodal Agent for Visually Grounded Reasoning in Complex Chart Question Answering
by: Kaur, Rachneet, et al.
Published: (2025)
by: Kaur, Rachneet, et al.
Published: (2025)
O3D: Offline Data-driven Discovery and Distillation for Sequential Decision-Making with Large Language Models
by: Xiao, Yuchen, et al.
Published: (2023)
by: Xiao, Yuchen, et al.
Published: (2023)
When Safety Blocks Sense: Measuring Semantic Confusion in LLM Refusals
by: Anonto, Riad Ahmed, et al.
Published: (2025)
by: Anonto, Riad Ahmed, et al.
Published: (2025)
Is There No Such Thing as a Bad Question? H4R: HalluciBot For Ratiocination, Rewriting, Ranking, and Routing
by: Watson, William, et al.
Published: (2024)
by: Watson, William, et al.
Published: (2024)
LAW: Legal Agentic Workflows for Custody and Fund Services Contracts
by: Watson, William, et al.
Published: (2024)
by: Watson, William, et al.
Published: (2024)
LIME: Making LLM Data More Efficient with Linguistic Metadata Embeddings
by: Sztwiertnia, Sebastian, et al.
Published: (2025)
by: Sztwiertnia, Sebastian, et al.
Published: (2025)
Improve LLM-based Automatic Essay Scoring with Linguistic Features
by: Hou, Zhaoyi Joey, et al.
Published: (2025)
by: Hou, Zhaoyi Joey, et al.
Published: (2025)
What Makes a Good Response? An Empirical Analysis of Quality in Qualitative Interviews
by: Ivey, Jonathan, et al.
Published: (2026)
by: Ivey, Jonathan, et al.
Published: (2026)
Confusion-Aware Rubric Optimization for LLM-based Automated Grading
by: Chu, Yucheng, et al.
Published: (2026)
by: Chu, Yucheng, et al.
Published: (2026)
Collab: Controlled Decoding using Mixture of Agents for LLM Alignment
by: Chakraborty, Souradip, et al.
Published: (2025)
by: Chakraborty, Souradip, et al.
Published: (2025)
Experiences Build Characters: The Linguistic Origins and Functional Impact of LLM Personality
by: Wang, Xi, et al.
Published: (2026)
by: Wang, Xi, et al.
Published: (2026)
What Makes a Good Speech Tokenizer for LLM-Centric Speech Generation? A Systematic Study
by: Fan, Xiaoran, et al.
Published: (2025)
by: Fan, Xiaoran, et al.
Published: (2025)
Continual Learning of Domain Knowledge from Human Feedback in Text-to-SQL
by: Cook, Thomas, et al.
Published: (2025)
by: Cook, Thomas, et al.
Published: (2025)
What Makes a Reward Model a Good Teacher? An Optimization Perspective
by: Razin, Noam, et al.
Published: (2025)
by: Razin, Noam, et al.
Published: (2025)
KatFishNet: Detecting LLM-Generated Korean Text through Linguistic Feature Analysis
by: Park, Shinwoo, et al.
Published: (2025)
by: Park, Shinwoo, et al.
Published: (2025)
Differentiating Between Human-Written and AI-Generated Texts Using Automatically Extracted Linguistic Features
by: Georgiou, Georgios P.
Published: (2024)
by: Georgiou, Georgios P.
Published: (2024)
FISHNET: Financial Intelligence from Sub-querying, Harmonizing, Neural-Conditioning, Expert Swarms, and Task Planning
by: Cho, Nicole, et al.
Published: (2024)
by: Cho, Nicole, et al.
Published: (2024)
Impatient Users Confuse AI Agents: High-fidelity Simulations of Human Traits for Testing Agents
by: He, Muyu, et al.
Published: (2025)
by: He, Muyu, et al.
Published: (2025)
What Kind of Reasoning (if any) is an LLM actually doing? On the Stochastic Nature and Abductive Appearance of Large Language Models
by: Floridi, Luciano, et al.
Published: (2025)
by: Floridi, Luciano, et al.
Published: (2025)
The Death of Feature Engineering? BERT with Linguistic Features on SQuAD 2.0
by: Li, Jiawei, et al.
Published: (2024)
by: Li, Jiawei, et al.
Published: (2024)
Evaluating Causal Explanation in Medical Reports with LLM-Based and Human-Aligned Metrics
by: Cho, Yousang, et al.
Published: (2025)
by: Cho, Yousang, et al.
Published: (2025)
Query Disambiguation via Answer-Free Context: Doubling Performance on Humanity's Last Exam
by: Majurski, Michael, et al.
Published: (2026)
by: Majurski, Michael, et al.
Published: (2026)
An Investigation of Linguistic Biases in LLM-Based Recommendations
by: Venkateswaran, Nitin, et al.
Published: (2026)
by: Venkateswaran, Nitin, et al.
Published: (2026)
Linguistics-Aware Non-Distortionary LLM Watermarking
by: Park, Shinwoo, et al.
Published: (2026)
by: Park, Shinwoo, et al.
Published: (2026)
Bridging Linguistic Gaps: Cross-Lingual Mapping in Pre-Training and Dataset for Enhanced Multilingual LLM Performance
by: Zheng, Weihua, et al.
Published: (2026)
by: Zheng, Weihua, et al.
Published: (2026)
AdaptAgent: Adapting Multimodal Web Agents with Few-Shot Learning from Human Demonstrations
by: Verma, Gaurav, et al.
Published: (2024)
by: Verma, Gaurav, et al.
Published: (2024)
Linguistic Comparison of AI- and Human-Written Responses to Online Mental Health Queries
by: Saha, Koustuv, et al.
Published: (2025)
by: Saha, Koustuv, et al.
Published: (2025)
What Makes Good Data for Alignment? A Comprehensive Study of Automatic Data Selection in Instruction Tuning
by: Liu, Wei, et al.
Published: (2023)
by: Liu, Wei, et al.
Published: (2023)
LLMs can be easily Confused by Instructional Distractions
by: Hwang, Yerin, et al.
Published: (2025)
by: Hwang, Yerin, et al.
Published: (2025)
Good Idea or Not, Representation of LLM Could Tell
by: Xu, Yi, et al.
Published: (2024)
by: Xu, Yi, et al.
Published: (2024)
What Makes a Sale? Rethinking End-to-End Seller--Buyer Retail Dynamics with LLM Agents
by: Choi, Jeonghwan, et al.
Published: (2026)
by: Choi, Jeonghwan, et al.
Published: (2026)
A Linguistics-Aware LLM Watermarking via Syntactic Predictability
by: Park, Shinwoo, et al.
Published: (2025)
by: Park, Shinwoo, et al.
Published: (2025)
Measuring What Matters -- or What's Convenient?: Robustness of LLM-Based Scoring Systems to Construct-Irrelevant Factors
by: Walsh, Cole, et al.
Published: (2026)
by: Walsh, Cole, et al.
Published: (2026)
Similar Items
-
QueryBandits for Hallucination Mitigation: Exploiting Semantic Features for No-Regret Rewriting
by: Cho, Nicole, et al.
Published: (2025) -
No One Size Fits All: QueryBandits for Hallucination Mitigation
by: Cho, Nicole, et al.
Published: (2026) -
TASER: Table Agents for Schema-guided Extraction and Recommendation
by: Cho, Nicole, et al.
Published: (2025) -
HiddenTables & PyQTax: A Cooperative Game and Dataset For TableQA to Ensure Scale and Data Privacy Across a Myriad of Taxonomies
by: Watson, William, et al.
Published: (2024) -
What Makes Good Multilingual Reasoning? Disentangling Reasoning Traces with Measurable Features
by: Ki, Dayeon, et al.
Published: (2026)