Certifying Knowledge Comprehension in LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Chaudhary, Isha, Jain, Vedaant V., Singh, Gagandeep |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Revealing Interpretable Failure Modes of VLMs
di: Chaudhary, Isha, et al.
Pubblicazione: (2026)
di: Chaudhary, Isha, et al.
Pubblicazione: (2026)
Bypassing the Safety Training of Open-Source LLMs with Priming Attacks
di: Vega, Jason, et al.
Pubblicazione: (2023)
di: Vega, Jason, et al.
Pubblicazione: (2023)
Certifying Counterfactual Bias in LLMs
di: Chaudhary, Isha, et al.
Pubblicazione: (2024)
di: Chaudhary, Isha, et al.
Pubblicazione: (2024)
Lumos: Let there be Language Model System Certification
di: Chaudhary, Isha, et al.
Pubblicazione: (2025)
di: Chaudhary, Isha, et al.
Pubblicazione: (2025)
How Catastrophic is Your LLM? Certifying Risk in Conversation
di: Wang, Chengxiao, et al.
Pubblicazione: (2025)
di: Wang, Chengxiao, et al.
Pubblicazione: (2025)
Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs
di: Wang, Ganghua, et al.
Pubblicazione: (2025)
di: Wang, Ganghua, et al.
Pubblicazione: (2025)
SafetyNet: Detecting Harmful Outputs in LLMs by Modeling and Monitoring Deceptive Behaviors
di: Chaudhary, Maheep, et al.
Pubblicazione: (2025)
di: Chaudhary, Maheep, et al.
Pubblicazione: (2025)
When Shallow Wins: Silent Failures and the Depth-Accuracy Paradox in Latent Reasoning
di: Sahoo, Subramanyam, et al.
Pubblicazione: (2026)
di: Sahoo, Subramanyam, et al.
Pubblicazione: (2026)
SAHOO: Safeguarded Alignment for High-Order Optimization Objectives in Recursive Self-Improvement
di: Sahoo, Subramanyam, et al.
Pubblicazione: (2026)
di: Sahoo, Subramanyam, et al.
Pubblicazione: (2026)
The Reasoning Trap -- Logical Reasoning as a Mechanistic Pathway to Situational Awareness
di: Sahoo, Subramanyam, et al.
Pubblicazione: (2026)
di: Sahoo, Subramanyam, et al.
Pubblicazione: (2026)
Deep Knowledge-Infusion For Explainable Depression Detection
di: Dalal, Sumit, et al.
Pubblicazione: (2024)
di: Dalal, Sumit, et al.
Pubblicazione: (2024)
KSOD: Knowledge Supplement for LLMs On Demand
di: Li, Haoran, et al.
Pubblicazione: (2025)
di: Li, Haoran, et al.
Pubblicazione: (2025)
Demystifying Hybrid Thinking: Can LLMs Truly Switch Between Think and No-Think?
di: Wang, Shouren, et al.
Pubblicazione: (2025)
di: Wang, Shouren, et al.
Pubblicazione: (2025)
LLMs Are Not Intelligent Thinkers: Introducing Mathematical Topic Tree Benchmark for Comprehensive Evaluation of LLMs
di: Davoodi, Arash Gholami, et al.
Pubblicazione: (2024)
di: Davoodi, Arash Gholami, et al.
Pubblicazione: (2024)
A Comprehensive Evaluation framework of Alignment Techniques for LLMs
di: Azmat, Muneeza, et al.
Pubblicazione: (2025)
di: Azmat, Muneeza, et al.
Pubblicazione: (2025)
Interpreting the Effects of Quantization on LLMs
di: Singh, Manpreet, et al.
Pubblicazione: (2025)
di: Singh, Manpreet, et al.
Pubblicazione: (2025)
Certified Robustness Under Bounded Levenshtein Distance
di: Rocamora, Elias Abad, et al.
Pubblicazione: (2025)
di: Rocamora, Elias Abad, et al.
Pubblicazione: (2025)
Layerwise Recall and the Geometry of Interwoven Knowledge in LLMs
di: Lei, Ge, et al.
Pubblicazione: (2025)
di: Lei, Ge, et al.
Pubblicazione: (2025)
PropMEND: Hypernetworks for Knowledge Propagation in LLMs
di: Liu, Zeyu Leo, et al.
Pubblicazione: (2025)
di: Liu, Zeyu Leo, et al.
Pubblicazione: (2025)
Time-R1: Towards Comprehensive Temporal Reasoning in LLMs
di: Liu, Zijia, et al.
Pubblicazione: (2025)
di: Liu, Zijia, et al.
Pubblicazione: (2025)
Beyond Answers: Transferring Reasoning Capabilities to Smaller LLMs Using Multi-Teacher Knowledge Distillation
di: Tian, Yijun, et al.
Pubblicazione: (2024)
di: Tian, Yijun, et al.
Pubblicazione: (2024)
BlockCert: Certified Blockwise Extraction of Transformer Mechanisms
di: Andric, Sandro
Pubblicazione: (2025)
di: Andric, Sandro
Pubblicazione: (2025)
Parameter Efficient Fine Tuning: A Comprehensive Analysis Across Applications
di: Balne, Charith Chandra Sai, et al.
Pubblicazione: (2024)
di: Balne, Charith Chandra Sai, et al.
Pubblicazione: (2024)
Task-Aware LoRA Adapter Composition via Similarity Retrieval in Vector Databases
di: Adsul, Riya, et al.
Pubblicazione: (2026)
di: Adsul, Riya, et al.
Pubblicazione: (2026)
Memorization vs. Reasoning: Updating LLMs with New Knowledge
di: Li, Aochong Oliver, et al.
Pubblicazione: (2025)
di: Li, Aochong Oliver, et al.
Pubblicazione: (2025)
Fine-Tuning or Retrieval? Comparing Knowledge Injection in LLMs
di: Ovadia, Oded, et al.
Pubblicazione: (2023)
di: Ovadia, Oded, et al.
Pubblicazione: (2023)
Alignment-Constrained Dynamic Pruning for LLMs: Identifying and Preserving Alignment-Critical Circuits
di: Patel, Dev, et al.
Pubblicazione: (2025)
di: Patel, Dev, et al.
Pubblicazione: (2025)
Dialogue Without Limits: Constant-Sized KV Caches for Extended Responses in LLMs
di: Ghadia, Ravi, et al.
Pubblicazione: (2025)
di: Ghadia, Ravi, et al.
Pubblicazione: (2025)
Can GPT Redefine Medical Understanding? Evaluating GPT on Biomedical Machine Reading Comprehension
di: Vatsal, Shubham, et al.
Pubblicazione: (2024)
di: Vatsal, Shubham, et al.
Pubblicazione: (2024)
Comprehensive Modeling and Question Answering of Cancer Clinical Practice Guidelines using LLMs
di: Gupta, Bhumika, et al.
Pubblicazione: (2025)
di: Gupta, Bhumika, et al.
Pubblicazione: (2025)
Towards a Holistic Evaluation of LLMs on Factual Knowledge Recall
di: Yuan, Jiaqing, et al.
Pubblicazione: (2024)
di: Yuan, Jiaqing, et al.
Pubblicazione: (2024)
RAG in the Wild: On the (In)effectiveness of LLMs with Mixture-of-Knowledge Retrieval Augmentation
di: Xu, Ran, et al.
Pubblicazione: (2025)
di: Xu, Ran, et al.
Pubblicazione: (2025)
Censored LLMs as a Natural Testbed for Secret Knowledge Elicitation
di: Casademunt, Helena, et al.
Pubblicazione: (2026)
di: Casademunt, Helena, et al.
Pubblicazione: (2026)
Do LLMs Encode Functional Importance of Reasoning Tokens?
di: Singh, Janvijay, et al.
Pubblicazione: (2026)
di: Singh, Janvijay, et al.
Pubblicazione: (2026)
Agribot: agriculture-specific question answer system
di: Jain, Naman, et al.
Pubblicazione: (2025)
di: Jain, Naman, et al.
Pubblicazione: (2025)
Comprehensive Reassessment of Large-Scale Evaluation Outcomes in LLMs: A Multifaceted Statistical Approach
di: Sun, Kun, et al.
Pubblicazione: (2024)
di: Sun, Kun, et al.
Pubblicazione: (2024)
PII-Scope: A Comprehensive Study on Training Data PII Extraction Attacks in LLMs
di: Nakka, Krishna Kanth, et al.
Pubblicazione: (2024)
di: Nakka, Krishna Kanth, et al.
Pubblicazione: (2024)
Retrieval Augmented Generation or Long-Context LLMs? A Comprehensive Study and Hybrid Approach
di: Li, Zhuowan, et al.
Pubblicazione: (2024)
di: Li, Zhuowan, et al.
Pubblicazione: (2024)
KS-Lottery: Finding Certified Lottery Tickets for Multilingual Language Models
di: Yuan, Fei, et al.
Pubblicazione: (2024)
di: Yuan, Fei, et al.
Pubblicazione: (2024)
Code Comprehension then Auditing for Unsupervised LLM Evaluation
di: Patel, Bhrij, et al.
Pubblicazione: (2024)
di: Patel, Bhrij, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Revealing Interpretable Failure Modes of VLMs
di: Chaudhary, Isha, et al.
Pubblicazione: (2026) -
Bypassing the Safety Training of Open-Source LLMs with Priming Attacks
di: Vega, Jason, et al.
Pubblicazione: (2023) -
Certifying Counterfactual Bias in LLMs
di: Chaudhary, Isha, et al.
Pubblicazione: (2024) -
Lumos: Let there be Language Model System Certification
di: Chaudhary, Isha, et al.
Pubblicazione: (2025) -
How Catastrophic is Your LLM? Certifying Risk in Conversation
di: Wang, Chengxiao, et al.
Pubblicazione: (2025)