Evaluating Input Feature Explanations through a Unified Diagnostic Evaluation Framework
Fuente:
arXiv
Salvato in:
| Autori principali: | Sun, Jingyi, Atanasova, Pepa, Augenstein, Isabelle |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Revealing the Parametric Knowledge of Language Models: A Unified Framework for Attribution Methods
di: Yu, Haeun, et al.
Pubblicazione: (2024)
di: Yu, Haeun, et al.
Pubblicazione: (2024)
DYNAMICQA: Tracing Internal Knowledge Conflicts in Language Models
di: Marjanović, Sara Vera, et al.
Pubblicazione: (2024)
di: Marjanović, Sara Vera, et al.
Pubblicazione: (2024)
The Unlikely Duel: Evaluating Creative Writing in LLMs through a Unique Scenario
di: Gómez-Rodríguez, Carlos, et al.
Pubblicazione: (2024)
di: Gómez-Rodríguez, Carlos, et al.
Pubblicazione: (2024)
Multiplication in Multimodal LLMs: Computation with Text, Image, and Audio Inputs
di: Balter, Samuel G., et al.
Pubblicazione: (2026)
di: Balter, Samuel G., et al.
Pubblicazione: (2026)
Trusted Uncertainty in Large Language Models: A Unified Framework for Confidence Calibration and Risk-Controlled Refusal
di: Oehri, Markus, et al.
Pubblicazione: (2025)
di: Oehri, Markus, et al.
Pubblicazione: (2025)
Evaluating Pixel Language Models on Non-Standardized Languages
di: Muñoz-Ortiz, Alberto, et al.
Pubblicazione: (2024)
di: Muñoz-Ortiz, Alberto, et al.
Pubblicazione: (2024)
Setting Standards in Turkish NLP: TR-MMLU for Large Language Model Evaluation
di: Bayram, M. Ali, et al.
Pubblicazione: (2024)
di: Bayram, M. Ali, et al.
Pubblicazione: (2024)
Omni-SafetyBench: A Benchmark for Safety Evaluation of Audio-Visual Large Language Models
di: Pan, Leyi, et al.
Pubblicazione: (2025)
di: Pan, Leyi, et al.
Pubblicazione: (2025)
The Paradox of Poetic Intent in Back-Translation: Evaluating the Quality of Large Language Models in Chinese Translation
di: Weigang, Li, et al.
Pubblicazione: (2025)
di: Weigang, Li, et al.
Pubblicazione: (2025)
NurValues: Real-World Nursing Values Evaluation for Large Language Models in Clinical Context
di: Yao, Ben, et al.
Pubblicazione: (2025)
di: Yao, Ben, et al.
Pubblicazione: (2025)
How Human-Like Are Large Language Models? A Register-Aware Linguistic Evaluation Framework
di: Nieth, Björn, et al.
Pubblicazione: (2026)
di: Nieth, Björn, et al.
Pubblicazione: (2026)
Unifying Uniform and Binary-coding Quantization for Accurate Compression of Large Language Models
di: Park, Seungcheol, et al.
Pubblicazione: (2025)
di: Park, Seungcheol, et al.
Pubblicazione: (2025)
Exploring Italian sentence embeddings properties through multi-tasking
di: Nastase, Vivi, et al.
Pubblicazione: (2024)
di: Nastase, Vivi, et al.
Pubblicazione: (2024)
Tracking linguistic information in transformer-based sentence embeddings through targeted sparsification
di: Nastase, Vivi, et al.
Pubblicazione: (2024)
di: Nastase, Vivi, et al.
Pubblicazione: (2024)
Exploring syntactic information in sentence embeddings through multilingual subject-verb agreement
di: Nastase, Vivi, et al.
Pubblicazione: (2024)
di: Nastase, Vivi, et al.
Pubblicazione: (2024)
SentiCSE: A Sentiment-aware Contrastive Sentence Embedding Framework with Sentiment-guided Textual Similarity
di: Kim, Jaemin, et al.
Pubblicazione: (2024)
di: Kim, Jaemin, et al.
Pubblicazione: (2024)
ScoreRAG: A Retrieval-Augmented Generation Framework with Consistency-Relevance Scoring and Structured Summarization for News Generation
di: Lin, Pei-Yun, et al.
Pubblicazione: (2025)
di: Lin, Pei-Yun, et al.
Pubblicazione: (2025)
Thinking Longer, Not Always Smarter: Evaluating LLM Capabilities in Hierarchical Legal Reasoning
di: Zhang, Li, et al.
Pubblicazione: (2025)
di: Zhang, Li, et al.
Pubblicazione: (2025)
TRUE: A Trustworthy Unified Explanation Framework for Large Language Model Reasoning
di: Yang, Yujiao
Pubblicazione: (2026)
di: Yang, Yujiao
Pubblicazione: (2026)
PaperAudit-Bench: Benchmarking Error Detection in Research Papers for Critical Automated Peer Review
di: Tu, Songjun, et al.
Pubblicazione: (2026)
di: Tu, Songjun, et al.
Pubblicazione: (2026)
VietMed-MCQ: A Consistency-Filtered Data Synthesis Framework for Vietnamese Traditional Medicine Evaluation
di: Kiet, Huynh Trung, et al.
Pubblicazione: (2026)
di: Kiet, Huynh Trung, et al.
Pubblicazione: (2026)
TIS-DPO: Token-level Importance Sampling for Direct Preference Optimization With Estimated Weights
di: Liu, Aiwei, et al.
Pubblicazione: (2024)
di: Liu, Aiwei, et al.
Pubblicazione: (2024)
Co-NAML-LSTUR: A Combined Model with Attentive Multi-View Learning and Long- and Short-term User Representations for News Recommendation
di: Nguyen, Minh Hoang, et al.
Pubblicazione: (2025)
di: Nguyen, Minh Hoang, et al.
Pubblicazione: (2025)
Charting a Decade of Computational Linguistics in Italy: The CLiC-it Corpus
di: Alzetta, Chiara, et al.
Pubblicazione: (2025)
di: Alzetta, Chiara, et al.
Pubblicazione: (2025)
UnifiedCrawl: Aggregated Common Crawl for Affordable Adaptation of LLMs on Low-Resource Languages
di: Tessema, Bethel Melesse, et al.
Pubblicazione: (2024)
di: Tessema, Bethel Melesse, et al.
Pubblicazione: (2024)
NOTAI.AI: Explainable Detection of Machine-Generated Text via Curvature and Feature Attribution
di: Breneur, Oleksandr Marchenko, et al.
Pubblicazione: (2026)
di: Breneur, Oleksandr Marchenko, et al.
Pubblicazione: (2026)
Textual Data Bias Detection and Mitigation -- An Extensible Pipeline with Experimental Evaluation
di: Görge, Rebekka, et al.
Pubblicazione: (2025)
di: Görge, Rebekka, et al.
Pubblicazione: (2025)
The Knesset Corpus: An Annotated Corpus of Hebrew Parliamentary Proceedings
di: Goldin, Gili, et al.
Pubblicazione: (2024)
di: Goldin, Gili, et al.
Pubblicazione: (2024)
Towards Effective and Efficient Continual Pre-training of Large Language Models
di: Chen, Jie, et al.
Pubblicazione: (2024)
di: Chen, Jie, et al.
Pubblicazione: (2024)
Direct Large Language Model Alignment Through Self-Rewarding Contrastive Prompt Distillation
di: Liu, Aiwei, et al.
Pubblicazione: (2024)
di: Liu, Aiwei, et al.
Pubblicazione: (2024)
The Superalignment of Superhuman Intelligence with Large Language Models
di: Huang, Minlie, et al.
Pubblicazione: (2024)
di: Huang, Minlie, et al.
Pubblicazione: (2024)
Train-Attention: Meta-Learning Where to Focus in Continual Knowledge Learning
di: Seo, Yeongbin, et al.
Pubblicazione: (2024)
di: Seo, Yeongbin, et al.
Pubblicazione: (2024)
Experimentation in Content Moderation using RWKV
di: Yildirim, Umut, et al.
Pubblicazione: (2024)
di: Yildirim, Umut, et al.
Pubblicazione: (2024)
Branching Narratives: Character Decision Points Detection
di: Tikhonov, Alexey
Pubblicazione: (2024)
di: Tikhonov, Alexey
Pubblicazione: (2024)
Dancing in the syntax forest: fast, accurate and explainable sentiment analysis with SALSA
di: Gómez-Rodríguez, Carlos, et al.
Pubblicazione: (2024)
di: Gómez-Rodríguez, Carlos, et al.
Pubblicazione: (2024)
On the Robustness of Document-Level Relation Extraction Models to Entity Name Variations
di: Meng, Shiao, et al.
Pubblicazione: (2024)
di: Meng, Shiao, et al.
Pubblicazione: (2024)
Are there identifiable structural parts in the sentence embedding whole?
di: Nastase, Vivi, et al.
Pubblicazione: (2024)
di: Nastase, Vivi, et al.
Pubblicazione: (2024)
Can LLMs Compute with Reasons?
di: Sandilya, Harshit, et al.
Pubblicazione: (2024)
di: Sandilya, Harshit, et al.
Pubblicazione: (2024)
A Survey on Natural Language Counterfactual Generation
di: Wang, Yongjie, et al.
Pubblicazione: (2024)
di: Wang, Yongjie, et al.
Pubblicazione: (2024)
Large Language Models(LLMs) on Tabular Data: Prediction, Generation, and Understanding -- A Survey
di: Fang, Xi, et al.
Pubblicazione: (2024)
di: Fang, Xi, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Revealing the Parametric Knowledge of Language Models: A Unified Framework for Attribution Methods
di: Yu, Haeun, et al.
Pubblicazione: (2024) -
DYNAMICQA: Tracing Internal Knowledge Conflicts in Language Models
di: Marjanović, Sara Vera, et al.
Pubblicazione: (2024) -
The Unlikely Duel: Evaluating Creative Writing in LLMs through a Unique Scenario
di: Gómez-Rodríguez, Carlos, et al.
Pubblicazione: (2024) -
Multiplication in Multimodal LLMs: Computation with Text, Image, and Audio Inputs
di: Balter, Samuel G., et al.
Pubblicazione: (2026) -
Trusted Uncertainty in Large Language Models: A Unified Framework for Confidence Calibration and Risk-Controlled Refusal
di: Oehri, Markus, et al.
Pubblicazione: (2025)