Implicit Intelligence -- Evaluating Agents on What Users Don't Say
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sirdeshmukh, Ved, Wetter, Marc |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Reasoning Models Don't Always Say What They Think
von: Chen, Yanda, et al.
Veröffentlicht: (2025)
von: Chen, Yanda, et al.
Veröffentlicht: (2025)
R-ConstraintBench: Evaluating LLMs on NP-Complete Scheduling
von: Jain, Raj, et al.
Veröffentlicht: (2025)
von: Jain, Raj, et al.
Veröffentlicht: (2025)
Intent Laundering: AI Safety Datasets Are Not What They Seem
von: Golchin, Shahriar, et al.
Veröffentlicht: (2026)
von: Golchin, Shahriar, et al.
Veröffentlicht: (2026)
Measuring What VLMs Don't Say: Validation Metrics Hide Clinical Terminology Erasure in Radiology Report Generation
von: Parikh, Aditya, et al.
Veröffentlicht: (2026)
von: Parikh, Aditya, et al.
Veröffentlicht: (2026)
What Models Know, How Well They Know It: Knowledge-Weighted Fine-Tuning for Learning When to Say "I Don't Know"
von: Lee, Joosung, et al.
Veröffentlicht: (2026)
von: Lee, Joosung, et al.
Veröffentlicht: (2026)
What We Don't C: Manifold Disentanglement for Structured Discovery
von: Rogers, Brian, et al.
Veröffentlicht: (2025)
von: Rogers, Brian, et al.
Veröffentlicht: (2025)
Don't Just Fine-tune the Agent, Tune the Environment
von: Lu, Siyuan, et al.
Veröffentlicht: (2025)
von: Lu, Siyuan, et al.
Veröffentlicht: (2025)
Can AI Assistants Know What They Don't Know?
von: Cheng, Qinyuan, et al.
Veröffentlicht: (2024)
von: Cheng, Qinyuan, et al.
Veröffentlicht: (2024)
Don't Start What You Can't Finish: A Counterfactual Audit of Support-State Triage in LLM Agents
von: Unlu, Eren
Veröffentlicht: (2026)
von: Unlu, Eren
Veröffentlicht: (2026)
Show, Don't Tell: Uncovering Implicit Character Portrayal using LLMs
von: Jaipersaud, Brandon, et al.
Veröffentlicht: (2024)
von: Jaipersaud, Brandon, et al.
Veröffentlicht: (2024)
Honest AI: Fine-Tuning "Small" Language Models to Say "I Don't Know", and Reducing Hallucination in RAG
von: Chen, Xinxi, et al.
Veröffentlicht: (2024)
von: Chen, Xinxi, et al.
Veröffentlicht: (2024)
Towards Agents That Know When They Don't Know: Uncertainty as a Control Signal for Structured Reasoning
von: Stoisser, Josefa Lia, et al.
Veröffentlicht: (2025)
von: Stoisser, Josefa Lia, et al.
Veröffentlicht: (2025)
Don't Pay Attention
von: Hammoud, Mohammad, et al.
Veröffentlicht: (2025)
von: Hammoud, Mohammad, et al.
Veröffentlicht: (2025)
What LLMs Think When You Don't Tell Them What to Think About?
von: Kwon, Yongchan, et al.
Veröffentlicht: (2026)
von: Kwon, Yongchan, et al.
Veröffentlicht: (2026)
xAI-Drop: Don't Use What You Cannot Explain
von: De Luca, Vincenzo Marco, et al.
Veröffentlicht: (2024)
von: De Luca, Vincenzo Marco, et al.
Veröffentlicht: (2024)
What Papers Don't Tell You: Recovering Tacit Knowledge for Automated Paper Reproduction
von: Li, Lehui, et al.
Veröffentlicht: (2026)
von: Li, Lehui, et al.
Veröffentlicht: (2026)
Don't Click That: Teaching Web Agents to Resist Deceptive Interfaces
von: Zhang, Yilin, et al.
Veröffentlicht: (2026)
von: Zhang, Yilin, et al.
Veröffentlicht: (2026)
Doing What They Say, Not What They Reason: Locating the Faithfulness Gap in LLM Agents
von: Wang, Yufeng
Veröffentlicht: (2026)
von: Wang, Yufeng
Veröffentlicht: (2026)
MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs
von: Sirdeshmukh, Ved, et al.
Veröffentlicht: (2025)
von: Sirdeshmukh, Ved, et al.
Veröffentlicht: (2025)
When Robots Should Say "I Don't Know": Benchmarking Abstention in Embodied Question Answering
von: Wu, Tao, et al.
Veröffentlicht: (2025)
von: Wu, Tao, et al.
Veröffentlicht: (2025)
Prediction Bottlenecks Don't Discover Causal Structure (But Here's What They Actually Do)
von: Lade, Ankit Hemant, et al.
Veröffentlicht: (2026)
von: Lade, Ankit Hemant, et al.
Veröffentlicht: (2026)
Know What You Don't Know: Uncertainty Calibration of Process Reward Models
von: Park, Young-Jin, et al.
Veröffentlicht: (2025)
von: Park, Young-Jin, et al.
Veröffentlicht: (2025)
Know What You Don't Know: Selective Prediction for Early Exit DNNs
von: Bajpai, Divya Jyoti, et al.
Veröffentlicht: (2025)
von: Bajpai, Divya Jyoti, et al.
Veröffentlicht: (2025)
Tell Me What You Don't Know: Enhancing Refusal Capabilities of Role-Playing Agents via Representation Space Analysis and Editing
von: Liu, Wenhao, et al.
Veröffentlicht: (2024)
von: Liu, Wenhao, et al.
Veröffentlicht: (2024)
Large Language Models Must Be Taught to Know What They Don't Know
von: Kapoor, Sanyam, et al.
Veröffentlicht: (2024)
von: Kapoor, Sanyam, et al.
Veröffentlicht: (2024)
Don't Freeze, Don't Crash: Extending the Safe Operating Range of Neural Navigation in Dense Crowds
von: Zhang, Jiefu, et al.
Veröffentlicht: (2026)
von: Zhang, Jiefu, et al.
Veröffentlicht: (2026)
Do Large Language Models Know What They Don't Know? Kalshibench: A New Benchmark for Evaluating Epistemic Calibration via Prediction Markets
von: Nel, Lukas
Veröffentlicht: (2025)
von: Nel, Lukas
Veröffentlicht: (2025)
Don't Blink: Evidence Collapse during Multimodal Reasoning
von: Raghu, Suresh, et al.
Veröffentlicht: (2026)
von: Raghu, Suresh, et al.
Veröffentlicht: (2026)
Don't Get Too Excited -- Eliciting Emotions in LLMs
von: Fazzi, Gino Franco, et al.
Veröffentlicht: (2025)
von: Fazzi, Gino Franco, et al.
Veröffentlicht: (2025)
PrivacyPeek: Auditing What LLM-Based Agents Acquire, Not Just What They Say
von: Zhang, Mingxuan, et al.
Veröffentlicht: (2026)
von: Zhang, Mingxuan, et al.
Veröffentlicht: (2026)
Please Don't Kill My Vibe: Empowering Agents with Data Flow Control
von: Summers, Charlie, et al.
Veröffentlicht: (2025)
von: Summers, Charlie, et al.
Veröffentlicht: (2025)
Trust, or Don't Predict: Introducing the CWSA Family for Confidence-Aware Model Evaluation
von: Shahnazari, Kourosh, et al.
Veröffentlicht: (2025)
von: Shahnazari, Kourosh, et al.
Veröffentlicht: (2025)
Don't Kill the Baby: The Case for AI in Arbitration
von: Broyde, Michael, et al.
Veröffentlicht: (2024)
von: Broyde, Michael, et al.
Veröffentlicht: (2024)
Reasoning Models Reason Well, Until They Don't
von: Rameshkumar, Revanth, et al.
Veröffentlicht: (2025)
von: Rameshkumar, Revanth, et al.
Veröffentlicht: (2025)
Don't Make the LLM Read the Graph: Make the Graph Think
von: Sun, Yuqi, et al.
Veröffentlicht: (2026)
von: Sun, Yuqi, et al.
Veröffentlicht: (2026)
LLM Cyber Evaluations Don't Capture Real-World Risk
von: Lukošiūtė, Kamilė, et al.
Veröffentlicht: (2025)
von: Lukošiūtė, Kamilė, et al.
Veröffentlicht: (2025)
s3: You Don't Need That Much Data to Train a Search Agent via RL
von: Jiang, Pengcheng, et al.
Veröffentlicht: (2025)
von: Jiang, Pengcheng, et al.
Veröffentlicht: (2025)
Don't Forget Imagination!
von: Vityaev, Evgenii E., et al.
Veröffentlicht: (2025)
von: Vityaev, Evgenii E., et al.
Veröffentlicht: (2025)
Language Models Don't Learn the Physical Manifestation of Language
von: Lee, Bruce W., et al.
Veröffentlicht: (2024)
von: Lee, Bruce W., et al.
Veröffentlicht: (2024)
Transformers Don't In-Context Learn Least Squares Regression
von: Hill, Joshua, et al.
Veröffentlicht: (2025)
von: Hill, Joshua, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Reasoning Models Don't Always Say What They Think
von: Chen, Yanda, et al.
Veröffentlicht: (2025) -
R-ConstraintBench: Evaluating LLMs on NP-Complete Scheduling
von: Jain, Raj, et al.
Veröffentlicht: (2025) -
Intent Laundering: AI Safety Datasets Are Not What They Seem
von: Golchin, Shahriar, et al.
Veröffentlicht: (2026) -
Measuring What VLMs Don't Say: Validation Metrics Hide Clinical Terminology Erasure in Radiology Report Generation
von: Parikh, Aditya, et al.
Veröffentlicht: (2026) -
What Models Know, How Well They Know It: Knowledge-Weighted Fine-Tuning for Learning When to Say "I Don't Know"
von: Lee, Joosung, et al.
Veröffentlicht: (2026)