Long-form evaluation of model editing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Rosati, Domenic, Gonzales, Robie, Chen, Jinkun, Yu, Xuemin, Erkan, Melis, Kayani, Yahya, Chavatapalli, Satya Deepika, Rudzicz, Frank, Sajjad, Hassan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Representation Noising: A Defence Mechanism Against Harmful Finetuning
von: Rosati, Domenic, et al.
Veröffentlicht: (2024)
von: Rosati, Domenic, et al.
Veröffentlicht: (2024)
Immunization against harmful fine-tuning attacks
von: Rosati, Domenic, et al.
Veröffentlicht: (2024)
von: Rosati, Domenic, et al.
Veröffentlicht: (2024)
Evaluating Defences against Unsafe Feedback in RLHF
von: Rosati, Domenic, et al.
Veröffentlicht: (2024)
von: Rosati, Domenic, et al.
Veröffentlicht: (2024)
Resolving Lexical Bias in Model Editing
von: Rizwan, Hammad, et al.
Veröffentlicht: (2024)
von: Rizwan, Hammad, et al.
Veröffentlicht: (2024)
Dependency Parsing is More Parameter-Efficient with Normalization
von: Gajo, Paolo, et al.
Veröffentlicht: (2025)
von: Gajo, Paolo, et al.
Veröffentlicht: (2025)
LLMs Underperform Graph-Based Parsers on Supervised Relation Extraction for Complex Graphs
von: Gajo, Paolo, et al.
Veröffentlicht: (2026)
von: Gajo, Paolo, et al.
Veröffentlicht: (2026)
Limits of Convergence-Rate Control for Open-Weight Safety
von: Rosati, Domenic, et al.
Veröffentlicht: (2026)
von: Rosati, Domenic, et al.
Veröffentlicht: (2026)
Exploring the features used for summary evaluation by Human and GPT
von: Sadeghi, Zahra, et al.
Veröffentlicht: (2025)
von: Sadeghi, Zahra, et al.
Veröffentlicht: (2025)
Vector Quantized Latent Concepts: A Scalable Alternative to Clustering-Based Concept Discovery
von: Yu, Xuemin, et al.
Veröffentlicht: (2026)
von: Yu, Xuemin, et al.
Veröffentlicht: (2026)
Latent Concept-based Explanation of NLP Models
von: Yu, Xuemin, et al.
Veröffentlicht: (2024)
von: Yu, Xuemin, et al.
Veröffentlicht: (2024)
Cross-Layer Discrete Concept Discovery for Interpreting Language Models
von: Garg, Ankur, et al.
Veröffentlicht: (2025)
von: Garg, Ankur, et al.
Veröffentlicht: (2025)
Semantic Consistency for Assuring Reliability of Large Language Models
von: Raj, Harsh, et al.
Veröffentlicht: (2023)
von: Raj, Harsh, et al.
Veröffentlicht: (2023)
The GPT-WritingPrompts Dataset: A Comparative Analysis of Character Portrayal in Short Stories
von: Huang, Xi Yu, et al.
Veröffentlicht: (2024)
von: Huang, Xi Yu, et al.
Veröffentlicht: (2024)
Improving Consistency in Large Language Models through Chain of Guidance
von: Raj, Harsh, et al.
Veröffentlicht: (2025)
von: Raj, Harsh, et al.
Veröffentlicht: (2025)
Graph-tree Fusion Model with Bidirectional Information Propagation for Long Document Classification
von: Roy, Sudipta Singha, et al.
Veröffentlicht: (2024)
von: Roy, Sudipta Singha, et al.
Veröffentlicht: (2024)
Understanding Language Model Circuits through Knowledge Editing
von: Ge, Huaizhi, et al.
Veröffentlicht: (2024)
von: Ge, Huaizhi, et al.
Veröffentlicht: (2024)
How Well Can Knowledge Edit Methods Edit Perplexing Knowledge?
von: Ge, Huaizhi, et al.
Veröffentlicht: (2024)
von: Ge, Huaizhi, et al.
Veröffentlicht: (2024)
Scenarios and Approaches for Situated Natural Language Explanations
von: Qiu, Pengshuo, et al.
Veröffentlicht: (2024)
von: Qiu, Pengshuo, et al.
Veröffentlicht: (2024)
Auxiliary Knowledge-Induced Learning for Automatic Multi-Label Medical Document Classification
von: Wang, Xindi, et al.
Veröffentlicht: (2024)
von: Wang, Xindi, et al.
Veröffentlicht: (2024)
Consistency in Language Models: Current Landscape, Challenges, and Future Directions
von: Novikova, Jekaterina, et al.
Veröffentlicht: (2025)
von: Novikova, Jekaterina, et al.
Veröffentlicht: (2025)
LLM Library Learning Fails: A LEGO-Prover Case Study
von: Berlot-Attwell, Ian, et al.
Veröffentlicht: (2025)
von: Berlot-Attwell, Ian, et al.
Veröffentlicht: (2025)
Multi-stage Retrieve and Re-rank Model for Automatic Medical Coding Recommendation
von: Wang, Xindi, et al.
Veröffentlicht: (2024)
von: Wang, Xindi, et al.
Veröffentlicht: (2024)
Library Learning Doesn't: The Curious Case of the Single-Use "Library"
von: Berlot-Attwell, Ian, et al.
Veröffentlicht: (2024)
von: Berlot-Attwell, Ian, et al.
Veröffentlicht: (2024)
LLM-Generated Black-box Explanations Can Be Adversarially Helpful
von: Ajwani, Rohan, et al.
Veröffentlicht: (2024)
von: Ajwani, Rohan, et al.
Veröffentlicht: (2024)
Show, Don't Tell: Uncovering Implicit Character Portrayal using LLMs
von: Jaipersaud, Brandon, et al.
Veröffentlicht: (2024)
von: Jaipersaud, Brandon, et al.
Veröffentlicht: (2024)
Interpreting the Effects of Quantization on LLMs
von: Singh, Manpreet, et al.
Veröffentlicht: (2025)
von: Singh, Manpreet, et al.
Veröffentlicht: (2025)
Quantifying the Capabilities of LLMs across Scale and Precision
von: Badshah, Sher, et al.
Veröffentlicht: (2024)
von: Badshah, Sher, et al.
Veröffentlicht: (2024)
Plug and Play with Prompts: A Prompt Tuning Approach for Controlling Text Generation
von: Ajwani, Rohan Deepak, et al.
Veröffentlicht: (2024)
von: Ajwani, Rohan Deepak, et al.
Veröffentlicht: (2024)
Understanding Syntactic Generalization in Structure-inducing Language Models
von: Arps, David, et al.
Veröffentlicht: (2025)
von: Arps, David, et al.
Veröffentlicht: (2025)
Discovering Salient Neurons in Deep NLP Models
von: Durrani, Nadir, et al.
Veröffentlicht: (2022)
von: Durrani, Nadir, et al.
Veröffentlicht: (2022)
Long-form RewardBench: Evaluating Reward Models for Long-form Generation
von: Huang, Hui, et al.
Veröffentlicht: (2026)
von: Huang, Hui, et al.
Veröffentlicht: (2026)
MedSynth: Realistic, Synthetic Medical Dialogue-Note Pairs
von: Mianroodi, Ahmad Rezaie, et al.
Veröffentlicht: (2025)
von: Mianroodi, Ahmad Rezaie, et al.
Veröffentlicht: (2025)
ACCORD: Closing the Commonsense Measurability Gap
von: Roewer-Després, François, et al.
Veröffentlicht: (2024)
von: Roewer-Després, François, et al.
Veröffentlicht: (2024)
UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu
von: Adeeba, Farah, et al.
Veröffentlicht: (2025)
von: Adeeba, Farah, et al.
Veröffentlicht: (2025)
Multilingual Nonce Dependency Treebanks: Understanding how Language Models represent and process syntactic structure
von: Arps, David, et al.
Veröffentlicht: (2023)
von: Arps, David, et al.
Veröffentlicht: (2023)
Introduction and Analysis of an Interlinear Qur'an Translation with Unknown Old Anatolian Turkish
von: Süleyman Aksu, et al.
Veröffentlicht: (2023)
von: Süleyman Aksu, et al.
Veröffentlicht: (2023)
Long-form factuality in large language models
von: Wei, Jerry, et al.
Veröffentlicht: (2024)
von: Wei, Jerry, et al.
Veröffentlicht: (2024)
TALE: A Tool-Augmented Framework for Reference-Free Evaluation of Large Language Models
von: Badshah, Sher, et al.
Veröffentlicht: (2025)
von: Badshah, Sher, et al.
Veröffentlicht: (2025)
Which Words Matter Most in Zero-Shot Prompts?
von: Sadr, Nikta Gohari, et al.
Veröffentlicht: (2025)
von: Sadr, Nikta Gohari, et al.
Veröffentlicht: (2025)
How Do Lexical Senses Correspond Between Spoken German and German Sign Language?
von: Çelikkol, Melis, et al.
Veröffentlicht: (2026)
von: Çelikkol, Melis, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Representation Noising: A Defence Mechanism Against Harmful Finetuning
von: Rosati, Domenic, et al.
Veröffentlicht: (2024) -
Immunization against harmful fine-tuning attacks
von: Rosati, Domenic, et al.
Veröffentlicht: (2024) -
Evaluating Defences against Unsafe Feedback in RLHF
von: Rosati, Domenic, et al.
Veröffentlicht: (2024) -
Resolving Lexical Bias in Model Editing
von: Rizwan, Hammad, et al.
Veröffentlicht: (2024) -
Dependency Parsing is More Parameter-Efficient with Normalization
von: Gajo, Paolo, et al.
Veröffentlicht: (2025)