PertEval: Unveiling Real Knowledge Capacity of LLMs with Knowledge-Invariant Perturbations
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Jiatong, Hu, Renjun, Huang, Kunzhe, Zhuang, Yan, Liu, Qi, Zhu, Mengxiao, Shi, Xing, Lin, Wei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Unveiling LLMs: The Evolution of Latent Representations in a Dynamic Knowledge Graph
von: Bronzini, Marco, et al.
Veröffentlicht: (2024)
von: Bronzini, Marco, et al.
Veröffentlicht: (2024)
IntelliChain: An Integrated Framework for Enhanced Socratic Method Dialogue with LLMs and Knowledge Graphs
von: Qi, Changyong, et al.
Veröffentlicht: (2025)
von: Qi, Changyong, et al.
Veröffentlicht: (2025)
ArabLegalEval: A Multitask Benchmark for Assessing Arabic Legal Knowledge in Large Language Models
von: Hijazi, Faris, et al.
Veröffentlicht: (2024)
von: Hijazi, Faris, et al.
Veröffentlicht: (2024)
Knowledge Acquisition on Mass-shooting Events via LLMs for AI-Driven Justice
von: Ihugba, Benign John, et al.
Veröffentlicht: (2025)
von: Ihugba, Benign John, et al.
Veröffentlicht: (2025)
Explainable Few-shot Knowledge Tracing
von: Li, Haoxuan, et al.
Veröffentlicht: (2024)
von: Li, Haoxuan, et al.
Veröffentlicht: (2024)
Survey of Computerized Adaptive Testing: A Machine Learning Perspective
von: Zhuang, Yan, et al.
Veröffentlicht: (2024)
von: Zhuang, Yan, et al.
Veröffentlicht: (2024)
Disentangling Heterogeneous Knowledge Concept Embedding for Cognitive Diagnosis on Untested Knowledge
von: Zhang, Miao, et al.
Veröffentlicht: (2024)
von: Zhang, Miao, et al.
Veröffentlicht: (2024)
Using LLMs for Knowledge Component-level Correctness Labeling in Open-ended Coding Problems
von: Duan, Zhangqi, et al.
Veröffentlicht: (2026)
von: Duan, Zhangqi, et al.
Veröffentlicht: (2026)
Persuasion Dynamics in LLMs: Investigating Robustness and Adaptability in Knowledge and Safety with DuET-PD
von: Tan, Bryan Chen Zhengyu, et al.
Veröffentlicht: (2025)
von: Tan, Bryan Chen Zhengyu, et al.
Veröffentlicht: (2025)
Exploring Knowledge Tracing in Tutor-Student Dialogues using LLMs
von: Scarlatos, Alexander, et al.
Veröffentlicht: (2024)
von: Scarlatos, Alexander, et al.
Veröffentlicht: (2024)
Global Patterns of Knowledge: Language, Genre, and the Geography of Knowledge
von: Matsui, Akira, et al.
Veröffentlicht: (2025)
von: Matsui, Akira, et al.
Veröffentlicht: (2025)
LocalBench: Benchmarking LLMs on County-Level Local Knowledge and Reasoning
von: Gao, Zihan, et al.
Veröffentlicht: (2025)
von: Gao, Zihan, et al.
Veröffentlicht: (2025)
Few-Shot Knowledge Distillation of LLMs With Counterfactual Explanations
von: Hamman, Faisal, et al.
Veröffentlicht: (2025)
von: Hamman, Faisal, et al.
Veröffentlicht: (2025)
NumPert: Numerical Perturbations to Probe Language Models for Veracity Prediction
von: Aarnes, Peter Røysland, et al.
Veröffentlicht: (2025)
von: Aarnes, Peter Røysland, et al.
Veröffentlicht: (2025)
FoundationalASSIST: An Educational Dataset for Foundational Knowledge Tracing and Pedagogical Grounding of LLMs
von: Worden, Eamon, et al.
Veröffentlicht: (2026)
von: Worden, Eamon, et al.
Veröffentlicht: (2026)
The Ethical Knowledge Gap: Dispersed Knowledge, Sensemaking Failures, and Epistemic Dependence
von: Gogoll, Jan
Veröffentlicht: (2026)
von: Gogoll, Jan
Veröffentlicht: (2026)
A Survey of Knowledge Tracing: Models, Variants, and Applications
von: Shen, Shuanghong, et al.
Veröffentlicht: (2021)
von: Shen, Shuanghong, et al.
Veröffentlicht: (2021)
CollabEdit: Towards Non-destructive Collaborative Knowledge Editing
von: Zheng, Jiamu, et al.
Veröffentlicht: (2024)
von: Zheng, Jiamu, et al.
Veröffentlicht: (2024)
Knowledge-Augmented Reasoning for EUAIA Compliance and Adversarial Robustness of LLMs
von: Momcilovic, Tomas Bueno, et al.
Veröffentlicht: (2024)
von: Momcilovic, Tomas Bueno, et al.
Veröffentlicht: (2024)
Evaluating Reliability Asymmetries in Chinese Factual Search and AI Answers
von: Liu, Geng, et al.
Veröffentlicht: (2025)
von: Liu, Geng, et al.
Veröffentlicht: (2025)
Content Knowledge Identification with Multi-Agent Large Language Models (LLMs)
von: Yang, Kaiqi, et al.
Veröffentlicht: (2024)
von: Yang, Kaiqi, et al.
Veröffentlicht: (2024)
A Knowledge Graph Informing Soil Carbon Modeling
von: Shirvani-Mahdavi, Nasim, et al.
Veröffentlicht: (2025)
von: Shirvani-Mahdavi, Nasim, et al.
Veröffentlicht: (2025)
PLawBench: A Rubric-Based Benchmark for Evaluating LLMs in Real-World Legal Practice
von: Shi, Yuzhen, et al.
Veröffentlicht: (2026)
von: Shi, Yuzhen, et al.
Veröffentlicht: (2026)
Evaluating Chinese Large Language Models: The Influence of Persona Assignment on Stereotypes and Safeguards
von: Liu, Geng, et al.
Veröffentlicht: (2025)
von: Liu, Geng, et al.
Veröffentlicht: (2025)
ZPD-SCA: Unveiling the Blind Spots of LLMs in Assessing Students' Cognitive Abilities
von: Dong, Wenhan, et al.
Veröffentlicht: (2025)
von: Dong, Wenhan, et al.
Veröffentlicht: (2025)
Unveiling Scoring Processes: Dissecting the Differences between LLMs and Human Graders in Automatic Scoring
von: Wu, Xuansheng, et al.
Veröffentlicht: (2024)
von: Wu, Xuansheng, et al.
Veröffentlicht: (2024)
Transdisciplinary Multi Modal Approach to Knowledge
von: Casado, Johanna, et al.
Veröffentlicht: (2024)
von: Casado, Johanna, et al.
Veröffentlicht: (2024)
Recontextualized Knowledge and Narrative Coalitions on Telegram
von: Willaert, Tom
Veröffentlicht: (2024)
von: Willaert, Tom
Veröffentlicht: (2024)
Leveraging Knowledge Graphs and Large Language Models to Track and Analyze Learning Trajectories
von: Chen, Yu-Hxiang, et al.
Veröffentlicht: (2025)
von: Chen, Yu-Hxiang, et al.
Veröffentlicht: (2025)
Leveraging LLMs to Assess Tutor Moves in Real-Life Dialogues: A Feasibility Study
von: Thomas, Danielle R., et al.
Veröffentlicht: (2025)
von: Thomas, Danielle R., et al.
Veröffentlicht: (2025)
Measuring Large Language Models Capacity to Annotate Journalistic Sourcing
von: Vincent, Subramaniam, et al.
Veröffentlicht: (2024)
von: Vincent, Subramaniam, et al.
Veröffentlicht: (2024)
Life Histories of Taboo Knowledge Artifacts
von: Champion, Kaylea, et al.
Veröffentlicht: (2024)
von: Champion, Kaylea, et al.
Veröffentlicht: (2024)
Opening Knowledge Gaps Drives Scientific Progress
von: Kedrick, Kara, et al.
Veröffentlicht: (2025)
von: Kedrick, Kara, et al.
Veröffentlicht: (2025)
DietGlance: Dietary Monitoring and Personalized Analysis at a Glance with Knowledge-Empowered AI Assistant
von: Jiang, Zhihan, et al.
Veröffentlicht: (2025)
von: Jiang, Zhihan, et al.
Veröffentlicht: (2025)
SOTOPIA-$Ω$: Dynamic Strategy Injection Learning and Social Instruction Following Evaluation for Social Agents
von: Zhang, Wenyuan, et al.
Veröffentlicht: (2025)
von: Zhang, Wenyuan, et al.
Veröffentlicht: (2025)
Toward Open Science in the AEC Community: An Ecosystem for Sustainable Digital Knowledge Sharing and Reuse
von: Xiong, Ruoxin, et al.
Veröffentlicht: (2026)
von: Xiong, Ruoxin, et al.
Veröffentlicht: (2026)
Designing Knowledge Tools: How Students Transition from Using to Creating Generative AI in STEAM classroom
von: Huang, Qian, et al.
Veröffentlicht: (2025)
von: Huang, Qian, et al.
Veröffentlicht: (2025)
Fact-checking with Generative AI: A Systematic Cross-Topic Examination of LLMs Capacity to Detect Veracity of Political Information
von: Kuznetsova, Elizaveta, et al.
Veröffentlicht: (2025)
von: Kuznetsova, Elizaveta, et al.
Veröffentlicht: (2025)
Whose ChatGPT? Unveiling Real-World Educational Inequalities Introduced by Large Language Models
von: Yu, Renzhe, et al.
Veröffentlicht: (2024)
von: Yu, Renzhe, et al.
Veröffentlicht: (2024)
PICKT: Practical Interlinked Concept Knowledge Tracing for Personalized Learning using Knowledge Map Concept Relations
von: Lee, Wonbeen, et al.
Veröffentlicht: (2025)
von: Lee, Wonbeen, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Unveiling LLMs: The Evolution of Latent Representations in a Dynamic Knowledge Graph
von: Bronzini, Marco, et al.
Veröffentlicht: (2024) -
IntelliChain: An Integrated Framework for Enhanced Socratic Method Dialogue with LLMs and Knowledge Graphs
von: Qi, Changyong, et al.
Veröffentlicht: (2025) -
ArabLegalEval: A Multitask Benchmark for Assessing Arabic Legal Knowledge in Large Language Models
von: Hijazi, Faris, et al.
Veröffentlicht: (2024) -
Knowledge Acquisition on Mass-shooting Events via LLMs for AI-Driven Justice
von: Ihugba, Benign John, et al.
Veröffentlicht: (2025) -
Explainable Few-shot Knowledge Tracing
von: Li, Haoxuan, et al.
Veröffentlicht: (2024)