Whispers of Doubt Amidst Echoes of Triumph in NLP Robustness
Fuente:
arXiv
Saved in:
| Main Authors: | Gupta, Ashim, Rajendhran, Rishanth, Stringham, Nathan, Srikumar, Vivek, Marasović, Ana |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Chain-of-Thought Unfaithfulness as Disguised Accuracy
by: Bentham, Oliver, et al.
Published: (2024)
by: Bentham, Oliver, et al.
Published: (2024)
Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation
by: Gupta, Ashim, et al.
Published: (2025)
by: Gupta, Ashim, et al.
Published: (2025)
Found in Translation: Measuring Multilingual LLM Consistency as Simple as Translate then Evaluate
by: Gupta, Ashim, et al.
Published: (2025)
by: Gupta, Ashim, et al.
Published: (2025)
VeriFastScore: Speeding up long-form factuality evaluation
by: Rajendhran, Rishanth, et al.
Published: (2025)
by: Rajendhran, Rishanth, et al.
Published: (2025)
State Space Models are Strong Text Rerankers
by: Xu, Zhichao, et al.
Published: (2024)
by: Xu, Zhichao, et al.
Published: (2024)
An Empirical Investigation of Matrix Factorization Methods for Pre-trained Transformers
by: Gupta, Ashim, et al.
Published: (2024)
by: Gupta, Ashim, et al.
Published: (2024)
StoryScope: Investigating idiosyncrasies in AI fiction
by: Russell, Jenna, et al.
Published: (2026)
by: Russell, Jenna, et al.
Published: (2026)
Teaching People LLM's Errors and Getting it Right
by: Stringham, Nathan, et al.
Published: (2025)
by: Stringham, Nathan, et al.
Published: (2025)
Beyond Perplexity: Multi-dimensional Safety Evaluation of LLM Compression
by: Xu, Zhichao, et al.
Published: (2024)
by: Xu, Zhichao, et al.
Published: (2024)
Defragmenting Language Models: An Interpretability-based Approach for Vocabulary Expansion
by: Mehta, Maitrey, et al.
Published: (2026)
by: Mehta, Maitrey, et al.
Published: (2026)
LLM-Symbolic Integration for Robust Temporal Tabular Reasoning
by: Kulkarni, Atharv, et al.
Published: (2025)
by: Kulkarni, Atharv, et al.
Published: (2025)
BriefMe: A Legal NLP Benchmark for Assisting with Legal Briefs
by: Woo, Jesse, et al.
Published: (2025)
by: Woo, Jesse, et al.
Published: (2025)
Unequal Voices: How LLMs Construct Constrained Queer Narratives
by: Ghosal, Atreya, et al.
Published: (2025)
by: Ghosal, Atreya, et al.
Published: (2025)
Named Entity Recognition for Payment Data Using NLP
by: Nayak, Srikumar
Published: (2026)
by: Nayak, Srikumar
Published: (2026)
On Evaluating Explanation Utility for Human-AI Decision Making in NLP
by: Chaleshtori, Fateme Hashemi, et al.
Published: (2024)
by: Chaleshtori, Fateme Hashemi, et al.
Published: (2024)
Reinforcing Code Generation: Improving Text-to-SQL with Execution-Based Learning
by: Kulkarni, Atharv, et al.
Published: (2025)
by: Kulkarni, Atharv, et al.
Published: (2025)
InfiniteScienceGym: An Unbounded, Procedurally-Generated Benchmark for Scientific Analysis
by: Bentham, Oliver, et al.
Published: (2026)
by: Bentham, Oliver, et al.
Published: (2026)
Understanding the Logic of Direct Preference Alignment through Logic
by: Richardson, Kyle, et al.
Published: (2024)
by: Richardson, Kyle, et al.
Published: (2024)
Promptly Predicting Structures: The Return of Inference
by: Mehta, Maitrey, et al.
Published: (2024)
by: Mehta, Maitrey, et al.
Published: (2024)
Enhancing Question Answering on Charts Through Effective Pre-training Tasks
by: Gupta, Ashim, et al.
Published: (2024)
by: Gupta, Ashim, et al.
Published: (2024)
Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth
by: Gur-Arieh, Yoav, et al.
Published: (2026)
by: Gur-Arieh, Yoav, et al.
Published: (2026)
What Has Been Lost with Synthetic Evaluation?
by: Gill, Alexander, et al.
Published: (2025)
by: Gill, Alexander, et al.
Published: (2025)
In-Context Example Ordering Guided by Label Distributions
by: Xu, Zhichao, et al.
Published: (2024)
by: Xu, Zhichao, et al.
Published: (2024)
Measuring Chain of Thought Faithfulness by Unlearning Reasoning Steps
by: Tutek, Martin, et al.
Published: (2025)
by: Tutek, Martin, et al.
Published: (2025)
Distillation versus Contrastive Learning: How to Train Your Rerankers
by: Xu, Zhichao, et al.
Published: (2025)
by: Xu, Zhichao, et al.
Published: (2025)
Robust Privacy Amidst Innovation with Large Language Models Through a Critical Assessment of the Risks
by: Chuang, Yao-Shun, et al.
Published: (2024)
by: Chuang, Yao-Shun, et al.
Published: (2024)
Speaking of Language: Reflections on Metalanguage Research in NLP
by: Schneider, Nathan, et al.
Published: (2026)
by: Schneider, Nathan, et al.
Published: (2026)
HQFS: Hybrid Quantum Classical Financial Security with VQC Forecasting, QUBO Annealing, and Audit-Ready Post-Quantum Signing
by: Nayak, Srikumar
Published: (2026)
by: Nayak, Srikumar
Published: (2026)
Measuring the Robustness of NLP Models to Domain Shifts
by: Calderon, Nitay, et al.
Published: (2023)
by: Calderon, Nitay, et al.
Published: (2023)
When in Doubt, Cascade: Towards Building Efficient and Capable Guardrails
by: Nagireddy, Manish, et al.
Published: (2024)
by: Nagireddy, Manish, et al.
Published: (2024)
RLShield: Practical Multi-Agent RL for Financial Cyber Defense with Attack-Surface MDPs and Real-Time Response Orchestration
by: Nayak, Srikumar
Published: (2026)
by: Nayak, Srikumar
Published: (2026)
Calm-Whisper: Reduce Whisper Hallucination On Non-Speech By Calming Crazy Heads Down
by: Wang, Yingzhi, et al.
Published: (2025)
by: Wang, Yingzhi, et al.
Published: (2025)
Whispering in Amharic: Fine-tuning Whisper for Low-resource Language
by: Gete, Dawit Ketema, et al.
Published: (2025)
by: Gete, Dawit Ketema, et al.
Published: (2025)
Revisiting Overthinking in Long Chain-of-Thought from the Perspective of Self-Doubt
by: Peng, Keqin, et al.
Published: (2025)
by: Peng, Keqin, et al.
Published: (2025)
Indian Legal NLP Benchmarks : A Survey
by: Kalamkar, Prathamesh, et al.
Published: (2021)
by: Kalamkar, Prathamesh, et al.
Published: (2021)
Echoes of Discord: Forecasting Hater Reactions to Counterspeech
by: Song, Xiaoying, et al.
Published: (2025)
by: Song, Xiaoying, et al.
Published: (2025)
Mapping Clinical Doubt: Locating Linguistic Uncertainty in LLMs
by: Sridhar, Srivarshinee, et al.
Published: (2025)
by: Sridhar, Srivarshinee, et al.
Published: (2025)
Evaluating Concurrent Robustness of Language Models Across Diverse Challenge Sets
by: Gupta, Vatsal, et al.
Published: (2023)
by: Gupta, Vatsal, et al.
Published: (2023)
TurkicNLP: An NLP Toolkit for Turkic Languages
by: Hakimov, Sherzod
Published: (2026)
by: Hakimov, Sherzod
Published: (2026)
The Nature of NLP: Analyzing Contributions in NLP Papers
by: Pramanick, Aniket, et al.
Published: (2024)
by: Pramanick, Aniket, et al.
Published: (2024)
Similar Items
-
Chain-of-Thought Unfaithfulness as Disguised Accuracy
by: Bentham, Oliver, et al.
Published: (2024) -
Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation
by: Gupta, Ashim, et al.
Published: (2025) -
Found in Translation: Measuring Multilingual LLM Consistency as Simple as Translate then Evaluate
by: Gupta, Ashim, et al.
Published: (2025) -
VeriFastScore: Speeding up long-form factuality evaluation
by: Rajendhran, Rishanth, et al.
Published: (2025) -
State Space Models are Strong Text Rerankers
by: Xu, Zhichao, et al.
Published: (2024)