Robustness in Large Language Models: A Survey of Mitigation Strategies and Evaluation Metrics
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Kumar, Pankaj, Mishra, Subhankar |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
A Hybrid Framework for Real-Time Data Drift and Anomaly Identification Using Hierarchical Temporal Memory and Statistical Tests
par: Bandyopadhyay, Subhadip, et autres
Publié: (2025)
par: Bandyopadhyay, Subhadip, et autres
Publié: (2025)
Progressive Training for Explainable Citation-Grounded Dialogue: Reducing Hallucination to Zero in English-Hindi LLMs
par: Pandya, Vedant
Publié: (2026)
par: Pandya, Vedant
Publié: (2026)
SAGE: Sign-Adaptive Gradient for Memory-Efficient LLM Optimization
par: Lee, Wooin, et autres
Publié: (2026)
par: Lee, Wooin, et autres
Publié: (2026)
Language processing in humans and computers
par: Pavlovic, Dusko
Publié: (2024)
par: Pavlovic, Dusko
Publié: (2024)
SQuARE: Structured Query & Adaptive Retrieval Engine For Tabular Formats
par: Gondhalekar, Chinmay, et autres
Publié: (2025)
par: Gondhalekar, Chinmay, et autres
Publié: (2025)
When Are Two RLHF Objectives the Same?
par: Gaikwad, Madhava
Publié: (2025)
par: Gaikwad, Madhava
Publié: (2025)
Heterogeneous LLM Methods for Ontology Learning (Few-Shot Prompting, Ensemble Typing, and Attention-Based Taxonomies)
par: Beliaeva, Aleksandra, et autres
Publié: (2025)
par: Beliaeva, Aleksandra, et autres
Publié: (2025)
Thinking Machines: Mathematical Reasoning in the Age of LLMs
par: Asperti, Andrea, et autres
Publié: (2025)
par: Asperti, Andrea, et autres
Publié: (2025)
Challenges and Applications of Large Language Models: A Comparison of GPT and DeepSeek family of models
par: Sharma, Shubham, et autres
Publié: (2025)
par: Sharma, Shubham, et autres
Publié: (2025)
Understanding Reinforcement Learning for Model Training, and future directions with GRAPE
par: Patel, Rohit
Publié: (2025)
par: Patel, Rohit
Publié: (2025)
EdgeJury: Cross-Reviewed Small-Model Ensembles for Truthful Question Answering on Serverless Edge Inference
par: Kumar, Aayush
Publié: (2025)
par: Kumar, Aayush
Publié: (2025)
LinkedOut: Linking World Knowledge Representation Out of Video LLM for Next-Generation Video Recommendation
par: Zhang, Haichao, et autres
Publié: (2025)
par: Zhang, Haichao, et autres
Publié: (2025)
U-STS-LLM A Unified Spatio-Temporal Steered Large Language Model for Traffic Prediction and Imputation
par: Zhang, Yichen, et autres
Publié: (2026)
par: Zhang, Yichen, et autres
Publié: (2026)
Can Agentic AI Match the Performance of Human Data Scientists?
par: Luo, An, et autres
Publié: (2025)
par: Luo, An, et autres
Publié: (2025)
AgentDS Technical Report: Benchmarking the Future of Human-AI Collaboration in Domain-Specific Data Science
par: Luo, An, et autres
Publié: (2026)
par: Luo, An, et autres
Publié: (2026)
Ice Cream Doesn't Cause Drowning: Benchmarking LLMs Against Statistical Pitfalls in Causal Inference
par: Du, Jin, et autres
Publié: (2025)
par: Du, Jin, et autres
Publié: (2025)
EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta
par: Bernard, Raymond, et autres
Publié: (2024)
par: Bernard, Raymond, et autres
Publié: (2024)
ProactBench: Beyond What The User Asked For
par: Harfi, Sepehr, et autres
Publié: (2026)
par: Harfi, Sepehr, et autres
Publié: (2026)
Fact Grounded Attention: Eliminating Hallucination in Large Language Models Through Attention Level Knowledge Integration
par: Gupta, Aayush
Publié: (2025)
par: Gupta, Aayush
Publié: (2025)
Representing LLMs in Prompt Semantic Task Space
par: Kashani, Idan, et autres
Publié: (2025)
par: Kashani, Idan, et autres
Publié: (2025)
AVATAAR: Agentic Video Answering via Temporal Adaptive Alignment and Reasoning
par: Patel, Urjitkumar, et autres
Publié: (2025)
par: Patel, Urjitkumar, et autres
Publié: (2025)
Can Large Language Models Revolutionize Survey Research? Experiments with Disaster Preparedness Responses
par: Wang, Yan, et autres
Publié: (2026)
par: Wang, Yan, et autres
Publié: (2026)
Optimizing Retrieval-Augmented Generation (RAG) for Colloquial Cantonese: A LoRA-Based Systematic Review
par: Calonge, David Santandreu, et autres
Publié: (2025)
par: Calonge, David Santandreu, et autres
Publié: (2025)
AssistedDS: Benchmarking How External Domain Knowledge Assists LLMs in Automated Data Science
par: Luo, An, et autres
Publié: (2025)
par: Luo, An, et autres
Publié: (2025)
Improving Large-Scale k-Nearest Neighbor Text Categorization with Label Autoencoders
par: Ribadas-Pena, Francisco J., et autres
Publié: (2024)
par: Ribadas-Pena, Francisco J., et autres
Publié: (2024)
AI Agents-as-Judge: Automated Assessment of Accuracy, Consistency, Completeness and Clarity for Enterprise Documents
par: Dasgupta, Sudip, et autres
Publié: (2025)
par: Dasgupta, Sudip, et autres
Publié: (2025)
ORPHEAS: A Cross-Lingual Greek-English Embedding Model for Retrieval-Augmented Generation
par: Livieris, Ioannis E., et autres
Publié: (2026)
par: Livieris, Ioannis E., et autres
Publié: (2026)
Reducing Labeling Costs in Sentiment Analysis via Semi-Supervised Learning
par: Jafarlou, Minoo, et autres
Publié: (2024)
par: Jafarlou, Minoo, et autres
Publié: (2024)
WebMap -- Large Language Model-assisted Semantic Link Induction in the Web
par: Pokharel, Shiraj, et autres
Publié: (2025)
par: Pokharel, Shiraj, et autres
Publié: (2025)
On Self-improving Token Embeddings
par: Kubek, Mario M., et autres
Publié: (2025)
par: Kubek, Mario M., et autres
Publié: (2025)
Parameter-Efficient Neuroevolution for Diverse LLM Generation: Quality-Diversity Optimization via Prompt Embedding Evolution
par: Guo, Dongxin, et autres
Publié: (2026)
par: Guo, Dongxin, et autres
Publié: (2026)
Structure and Destructure: Dual Forces in the Making of Knowledge Engines
par: Chen, Yihong
Publié: (2025)
par: Chen, Yihong
Publié: (2025)
The OCON model: an old but green solution for distributable supervised classification for acoustic monitoring in smart cities
par: Giacomelli, Stefano, et autres
Publié: (2024)
par: Giacomelli, Stefano, et autres
Publié: (2024)
Beyond Long Context: When Semantics Matter More than Tokens
par: Chawdhury, Tarun Kumar, et autres
Publié: (2025)
par: Chawdhury, Tarun Kumar, et autres
Publié: (2025)
Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark
par: Bayram, M. Ali, et autres
Publié: (2025)
par: Bayram, M. Ali, et autres
Publié: (2025)
Variance-Aware LLM Annotation for Strategy Research: Sources, Diagnostics, and a Protocol for Reliable Measurement
par: Camuffo, Arnaldo, et autres
Publié: (2025)
par: Camuffo, Arnaldo, et autres
Publié: (2025)
Enhancing Diversity in Multi-objective Feature Selection
par: Miyandoab, Sevil Zanjani, et autres
Publié: (2024)
par: Miyandoab, Sevil Zanjani, et autres
Publié: (2024)
Combating data scarcity in recommendation services: Integrating cognitive types of VARK and neural network technologies (LLM)
par: Zmanovskii, Nikita
Publié: (2026)
par: Zmanovskii, Nikita
Publié: (2026)
LLM-supported document separation for printed reviews from zbMATH Open
par: Pluzhnikov, Ivan, et autres
Publié: (2026)
par: Pluzhnikov, Ivan, et autres
Publié: (2026)
Vis-CoT: A Human-in-the-Loop Framework for Interactive Visualization and Intervention in LLM Chain-of-Thought Reasoning
par: Pather, Kaviraj, et autres
Publié: (2025)
par: Pather, Kaviraj, et autres
Publié: (2025)
Documents similaires
-
A Hybrid Framework for Real-Time Data Drift and Anomaly Identification Using Hierarchical Temporal Memory and Statistical Tests
par: Bandyopadhyay, Subhadip, et autres
Publié: (2025) -
Progressive Training for Explainable Citation-Grounded Dialogue: Reducing Hallucination to Zero in English-Hindi LLMs
par: Pandya, Vedant
Publié: (2026) -
SAGE: Sign-Adaptive Gradient for Memory-Efficient LLM Optimization
par: Lee, Wooin, et autres
Publié: (2026) -
Language processing in humans and computers
par: Pavlovic, Dusko
Publié: (2024) -
SQuARE: Structured Query & Adaptive Retrieval Engine For Tabular Formats
par: Gondhalekar, Chinmay, et autres
Publié: (2025)