Confidence is Not Competence
Fuente:
arXiv
Saved in:
| Main Authors: | Sanyal, Debdeep, Pandey, Manya, Kumar, Dhruv, Deshpande, Saurabh, Mandal, Murari |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Agents Are All You Need for LLM Unlearning
by: Sanyal, Debdeep, et al.
Published: (2025)
by: Sanyal, Debdeep, et al.
Published: (2025)
LLM-as-a-Judge for Time Series Explanations
by: Sivalingam, Preetham, et al.
Published: (2026)
by: Sivalingam, Preetham, et al.
Published: (2026)
time2time: Causal Intervention in Hidden States to Simulate Rare Events in Time Series Foundation Models
by: Sanyal, Debdeep, et al.
Published: (2025)
by: Sanyal, Debdeep, et al.
Published: (2025)
Policy Optimization Prefers The Path of Least Resistance
by: Sanyal, Debdeep, et al.
Published: (2025)
by: Sanyal, Debdeep, et al.
Published: (2025)
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs
by: Sanyal, Debdeep, et al.
Published: (2025)
by: Sanyal, Debdeep, et al.
Published: (2025)
Nine Ways to Break Copyright Law and Why Our LLM Won't: A Fair Use Aligned Generation Framework
by: Sharma, Aakash Sen, et al.
Published: (2025)
by: Sharma, Aakash Sen, et al.
Published: (2025)
Measuring Representation Robustness in Large Language Models for Geometry
by: Jawandhia, Vedant, et al.
Published: (2026)
by: Jawandhia, Vedant, et al.
Published: (2026)
NagaNLP: Bootstrapping NLP for Low-Resource Nagamese Creole with Human-in-the-Loop Synthetic Data
by: Maiti, Agniva, et al.
Published: (2025)
by: Maiti, Agniva, et al.
Published: (2025)
Meta-Cultural Competence: Climbing the Right Hill of Cultural Awareness
by: Saha, Sougata, et al.
Published: (2025)
by: Saha, Sougata, et al.
Published: (2025)
ReviewEval: An Evaluation Framework for AI-Generated Reviews
by: Garg, Madhav Krishan, et al.
Published: (2025)
by: Garg, Madhav Krishan, et al.
Published: (2025)
Investigating Pedagogical Teacher and Student LLM Agents: Genetic Adaptation Meets Retrieval Augmented Generation Across Learning Style
by: Sanyal, Debdeep, et al.
Published: (2025)
by: Sanyal, Debdeep, et al.
Published: (2025)
CricBench: A Multilingual Benchmark for Evaluating LLMs in Cricket Analytics
by: Agarwal, Parth, et al.
Published: (2025)
by: Agarwal, Parth, et al.
Published: (2025)
REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM
by: Jindal, Madhur, et al.
Published: (2025)
by: Jindal, Madhur, et al.
Published: (2025)
When Reject Turns into Accept: Quantifying the Vulnerability of LLM-Based Scientific Reviewers to Indirect Prompt Injection
by: Sahoo, Devanshu, et al.
Published: (2025)
by: Sahoo, Devanshu, et al.
Published: (2025)
UnStar: Unlearning with Self-Taught Anti-Sample Reasoning for LLMs
by: Sinha, Yash, et al.
Published: (2024)
by: Sinha, Yash, et al.
Published: (2024)
SwissNYF: Tool Grounded LLM Agents for Black Box Setting
by: Kumar, Somnath Sendhil, et al.
Published: (2024)
by: Kumar, Somnath Sendhil, et al.
Published: (2024)
The Realignment Problem: When Right becomes Wrong in LLMs
by: Sharma, Aakash Sen, et al.
Published: (2025)
by: Sharma, Aakash Sen, et al.
Published: (2025)
Reading between the Lines: Can LLMs Identify Cross-Cultural Communication Gaps?
by: Saha, Sougata, et al.
Published: (2025)
by: Saha, Sougata, et al.
Published: (2025)
The Compliance Paradox: Semantic-Instruction Decoupling in Automated Academic Code Evaluation
by: Sahoo, Devanshu, et al.
Published: (2026)
by: Sahoo, Devanshu, et al.
Published: (2026)
Can pre-trained language models generate titles for research papers?
by: Rehman, Tohida, et al.
Published: (2024)
by: Rehman, Tohida, et al.
Published: (2024)
OrgAccess: A Benchmark for Role Based Access Control in Organization Scale LLMs
by: Sanyal, Debdeep, et al.
Published: (2025)
by: Sanyal, Debdeep, et al.
Published: (2025)
S2WTM: Spherical Sliced-Wasserstein Autoencoder for Topic Modeling
by: Adhya, Suman, et al.
Published: (2025)
by: Adhya, Suman, et al.
Published: (2025)
DTECT: Dynamic Topic Explorer & Context Tracker
by: Adhya, Suman, et al.
Published: (2025)
by: Adhya, Suman, et al.
Published: (2025)
Overview of the SciHigh Track at FIRE 2025: Research Highlight Generation from Scientific Papers
by: Rehman, Tohida, et al.
Published: (2026)
by: Rehman, Tohida, et al.
Published: (2026)
Evaluating Novelty in AI-Generated Research Plans Using Multi-Workflow LLM Pipelines
by: Saraogi, Devesh, et al.
Published: (2025)
by: Saraogi, Devesh, et al.
Published: (2025)
Automated Analysis of Learning Outcomes and Exam Questions Based on Bloom's Taxonomy
by: Kumar, Ramya, et al.
Published: (2025)
by: Kumar, Ramya, et al.
Published: (2025)
Infinite Problem Generator: Verifiably Scaling Physics Reasoning Data with Agentic Workflows
by: Sharan, Aditya, et al.
Published: (2026)
by: Sharan, Aditya, et al.
Published: (2026)
Beyond Accuracy: Diagnosing Algebraic Reasoning Failures in LLMs Across Nine Complexity Dimensions
by: Patil, Parth, et al.
Published: (2026)
by: Patil, Parth, et al.
Published: (2026)
Evaluating LLMs for Zeolite Synthesis Event Extraction (ZSEE): A Systematic Analysis of Prompting Strategies
by: Rathore, Charan Prakash, et al.
Published: (2025)
by: Rathore, Charan Prakash, et al.
Published: (2025)
Diagnosing LLM Judge Reliability: Conformal Prediction Sets and Transitivity Violations
by: Gupta, Manan, et al.
Published: (2026)
by: Gupta, Manan, et al.
Published: (2026)
Latent Phase-Shift Rollback: Inference-Time Error Correction via Residual Stream Monitoring and KV-Cache Steering
by: Gupta, Manan, et al.
Published: (2026)
by: Gupta, Manan, et al.
Published: (2026)
RAGBench: Explainable Benchmark for Retrieval-Augmented Generation Systems
by: Friel, Robert, et al.
Published: (2024)
by: Friel, Robert, et al.
Published: (2024)
A Hybrid Supervised-LLM Pipeline for Actionable Suggestion Mining in Unstructured Customer Reviews
by: Trivedi, Aakash, et al.
Published: (2026)
by: Trivedi, Aakash, et al.
Published: (2026)
Confidence Under the Hood: An Investigation into the Confidence-Probability Alignment in Large Language Models
by: Kumar, Abhishek, et al.
Published: (2024)
by: Kumar, Abhishek, et al.
Published: (2024)
SkillFactory: Self-Distillation For Learning Cognitive Behaviors
by: Sprague, Zayne, et al.
Published: (2025)
by: Sprague, Zayne, et al.
Published: (2025)
Agentic Confidence Calibration
by: Zhang, Jiaxin, et al.
Published: (2026)
by: Zhang, Jiaxin, et al.
Published: (2026)
Low-Confidence Gold: Refining Low-Confidence Samples for Efficient Instruction Tuning
by: Cai, Hongyi, et al.
Published: (2025)
by: Cai, Hongyi, et al.
Published: (2025)
Luna: An Evaluation Foundation Model to Catch Language Model Hallucinations with High Accuracy and Low Cost
by: Belyi, Masha, et al.
Published: (2024)
by: Belyi, Masha, et al.
Published: (2024)
Beyond Sentiment: A Multi-Agent Pipeline for Actionable Business Advice from Reviews
by: Bhandari, Kartikey Singh, et al.
Published: (2026)
by: Bhandari, Kartikey Singh, et al.
Published: (2026)
IndicDB -- Benchmarking Multilingual Text-to-SQL Capabilities in Indian Languages
by: Dawar, Aviral, et al.
Published: (2026)
by: Dawar, Aviral, et al.
Published: (2026)
Similar Items
-
Agents Are All You Need for LLM Unlearning
by: Sanyal, Debdeep, et al.
Published: (2025) -
LLM-as-a-Judge for Time Series Explanations
by: Sivalingam, Preetham, et al.
Published: (2026) -
time2time: Causal Intervention in Hidden States to Simulate Rare Events in Time Series Foundation Models
by: Sanyal, Debdeep, et al.
Published: (2025) -
Policy Optimization Prefers The Path of Least Resistance
by: Sanyal, Debdeep, et al.
Published: (2025) -
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs
by: Sanyal, Debdeep, et al.
Published: (2025)