Generative Active Testing: Efficient LLM Evaluation via Proxy Task Adaptation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ramakrishnan, Aashish Anantha, Saeedi, Ardavan, Hassanzadeh, Hamid Reza, Mohaghegh, Fazlolah, Lee, Dongwon |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CORDIAL: Can Multimodal Large Language Models Effectively Understand Coherence Relationships?
von: Ramakrishnan, Aashish Anantha, et al.
Veröffentlicht: (2025)
von: Ramakrishnan, Aashish Anantha, et al.
Veröffentlicht: (2025)
IRONIC: Coherence-Aware Reasoning Chains for Multi-Modal Sarcasm Detection
von: Ramakrishnan, Aashish Anantha, et al.
Veröffentlicht: (2025)
von: Ramakrishnan, Aashish Anantha, et al.
Veröffentlicht: (2025)
RONA: Pragmatically Diverse Image Captioning with Coherence Relations
von: Ramakrishnan, Aashish Anantha, et al.
Veröffentlicht: (2025)
von: Ramakrishnan, Aashish Anantha, et al.
Veröffentlicht: (2025)
ANCHOR: LLM-driven Subject Conditioning for Text-to-Image Synthesis
von: Ramakrishnan, Aashish Anantha, et al.
Veröffentlicht: (2024)
von: Ramakrishnan, Aashish Anantha, et al.
Veröffentlicht: (2024)
FIM-LoRA: Task-Informative Rank Allocation for LoRA via Calibration-Time Gradient-Variance Estimation
von: Sathyavageeswaran, Ramakrishnan
Veröffentlicht: (2026)
von: Sathyavageeswaran, Ramakrishnan
Veröffentlicht: (2026)
Classification of descriptions and summary using multiple passes of statistical and natural language toolkits
von: Banthia, Saumya, et al.
Veröffentlicht: (2020)
von: Banthia, Saumya, et al.
Veröffentlicht: (2020)
Toward Architecture-Aware Evaluation Metrics for LLM Agents
von: Souza, Débora, et al.
Veröffentlicht: (2026)
von: Souza, Débora, et al.
Veröffentlicht: (2026)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
von: Ashuach, Tomer, et al.
Veröffentlicht: (2025)
von: Ashuach, Tomer, et al.
Veröffentlicht: (2025)
Automatic Task Detection and Heterogeneous LLM Speculative Decoding
von: Ge, Danying, et al.
Veröffentlicht: (2025)
von: Ge, Danying, et al.
Veröffentlicht: (2025)
SPRInG: Continual LLM Personalization via Selective Parametric Adaptation and Retrieval-Interpolated Generation
von: Kim, Seoyeon, et al.
Veröffentlicht: (2026)
von: Kim, Seoyeon, et al.
Veröffentlicht: (2026)
AsyncTLS: Efficient Generative LLM Inference with Asynchronous Two-level Sparse Attention
von: Hu, Yuxuan, et al.
Veröffentlicht: (2026)
von: Hu, Yuxuan, et al.
Veröffentlicht: (2026)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
von: Oketunji, Abiodun Finbarrs
Veröffentlicht: (2023)
von: Oketunji, Abiodun Finbarrs
Veröffentlicht: (2023)
Enhancing Paraphrase Type Generation: The Impact of DPO and RLHF Evaluated with Human-Ranked Data
von: Lübbers, Christopher Lee
Veröffentlicht: (2025)
von: Lübbers, Christopher Lee
Veröffentlicht: (2025)
Emergent Lexical Semantics in Neural Language Models: Testing Martin's Law on LLM-Generated Text
von: Kugler, Kai
Veröffentlicht: (2025)
von: Kugler, Kai
Veröffentlicht: (2025)
Quantization-Robust LLM Unlearning via Low-Rank Adaptation
von: Abitante, João Vitor Boer, et al.
Veröffentlicht: (2026)
von: Abitante, João Vitor Boer, et al.
Veröffentlicht: (2026)
Evaluating the efficacy of LLM Safety Solutions : The Palit Benchmark Dataset
von: Palit, Sayon, et al.
Veröffentlicht: (2025)
von: Palit, Sayon, et al.
Veröffentlicht: (2025)
Active Context Compression: Autonomous Memory Management in LLM Agents
von: Verma, Nikhil
Veröffentlicht: (2026)
von: Verma, Nikhil
Veröffentlicht: (2026)
LoRS: Efficient Low-Rank Adaptation for Sparse Large Language Model
von: Hu, Yuxuan, et al.
Veröffentlicht: (2025)
von: Hu, Yuxuan, et al.
Veröffentlicht: (2025)
LLM-GLOBE: A Benchmark Evaluating the Cultural Values Embedded in LLM Output
von: Karinshak, Elise, et al.
Veröffentlicht: (2024)
von: Karinshak, Elise, et al.
Veröffentlicht: (2024)
Guided Speculative Inference for Efficient Test-Time Alignment of LLMs
von: Geuter, Jonathan, et al.
Veröffentlicht: (2025)
von: Geuter, Jonathan, et al.
Veröffentlicht: (2025)
ALISON: Fast and Effective Stylometric Authorship Obfuscation
von: Xing, Eric, et al.
Veröffentlicht: (2024)
von: Xing, Eric, et al.
Veröffentlicht: (2024)
Evaluating LLM Metrics Through Real-World Capabilities
von: Miller, Justin K, et al.
Veröffentlicht: (2025)
von: Miller, Justin K, et al.
Veröffentlicht: (2025)
Overview of the Sensemaking Task at the ELOQUENT 2025 Lab: LLMs as Teachers, Students and Evaluators
von: Šindelář, Pavel, et al.
Veröffentlicht: (2025)
von: Šindelář, Pavel, et al.
Veröffentlicht: (2025)
Understanding LLM Evaluator Behavior: A Structured Multi-Evaluator Framework for Merchant Risk Assessment
von: Wang, Liang, et al.
Veröffentlicht: (2026)
von: Wang, Liang, et al.
Veröffentlicht: (2026)
Identifying Fairness Issues in Automatically Generated Testing Content
von: Stowe, Kevin, et al.
Veröffentlicht: (2024)
von: Stowe, Kevin, et al.
Veröffentlicht: (2024)
HumanLLM: Benchmarking and Improving LLM Anthropomorphism via Human Cognitive Patterns
von: Wang, Xintao, et al.
Veröffentlicht: (2026)
von: Wang, Xintao, et al.
Veröffentlicht: (2026)
Efficient LLM Safety Evaluation through Multi-Agent Debate
von: Lin, Dachuan, et al.
Veröffentlicht: (2025)
von: Lin, Dachuan, et al.
Veröffentlicht: (2025)
Rubrik's Cube: Testing a New Rubric for Evaluating Explanations on the CUBE dataset
von: Galvan-Sosa, Diana, et al.
Veröffentlicht: (2025)
von: Galvan-Sosa, Diana, et al.
Veröffentlicht: (2025)
Adapting Multilingual Models to Code-Mixed Tasks via Model Merging
von: Kodali, Prashant, et al.
Veröffentlicht: (2025)
von: Kodali, Prashant, et al.
Veröffentlicht: (2025)
Xinyu: An Efficient LLM-based System for Commentary Generation
von: Wu, Yiquan, et al.
Veröffentlicht: (2024)
von: Wu, Yiquan, et al.
Veröffentlicht: (2024)
Decoding-Free Sampling Strategies for LLM Marginalization
von: Pohl, David, et al.
Veröffentlicht: (2025)
von: Pohl, David, et al.
Veröffentlicht: (2025)
LUCID: LLM-Generated Utterances for Complex and Interesting Dialogues
von: Stacey, Joe, et al.
Veröffentlicht: (2024)
von: Stacey, Joe, et al.
Veröffentlicht: (2024)
A Library of LLM Intrinsics for Retrieval-Augmented Generation
von: Danilevsky, Marina, et al.
Veröffentlicht: (2025)
von: Danilevsky, Marina, et al.
Veröffentlicht: (2025)
Text Generation: A Systematic Literature Review of Tasks, Evaluation, and Challenges
von: Becker, Jonas, et al.
Veröffentlicht: (2024)
von: Becker, Jonas, et al.
Veröffentlicht: (2024)
Select or Project? Evaluating Lower-dimensional Vectors for LLM Training Data Explanations
von: Hinterleitner, Lukas, et al.
Veröffentlicht: (2026)
von: Hinterleitner, Lukas, et al.
Veröffentlicht: (2026)
Can LLM Graph Reasoning Generalize beyond Pattern Memorization?
von: Zhang, Yizhuo, et al.
Veröffentlicht: (2024)
von: Zhang, Yizhuo, et al.
Veröffentlicht: (2024)
SAGE: Hierarchical LLM-Based Literary Evaluation through Ontology-Grounded Interpretive Dimensions
von: Wang, Tianyu, et al.
Veröffentlicht: (2026)
von: Wang, Tianyu, et al.
Veröffentlicht: (2026)
Clinical Document Corpora -- Real Ones, Translated and Synthetic Substitutes, and Assorted Domain Proxies: A Survey of Diversity in Corpus Design, with Focus on German Text Data
von: Hahn, Udo
Veröffentlicht: (2024)
von: Hahn, Udo
Veröffentlicht: (2024)
A Multi-Task Benchmark for Abusive Language Detection in Low-Resource Settings
von: Gaim, Fitsum, et al.
Veröffentlicht: (2025)
von: Gaim, Fitsum, et al.
Veröffentlicht: (2025)
VeriGuard: Enhancing LLM Agent Safety via Verified Code Generation
von: Miculicich, Lesly, et al.
Veröffentlicht: (2025)
von: Miculicich, Lesly, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
CORDIAL: Can Multimodal Large Language Models Effectively Understand Coherence Relationships?
von: Ramakrishnan, Aashish Anantha, et al.
Veröffentlicht: (2025) -
IRONIC: Coherence-Aware Reasoning Chains for Multi-Modal Sarcasm Detection
von: Ramakrishnan, Aashish Anantha, et al.
Veröffentlicht: (2025) -
RONA: Pragmatically Diverse Image Captioning with Coherence Relations
von: Ramakrishnan, Aashish Anantha, et al.
Veröffentlicht: (2025) -
ANCHOR: LLM-driven Subject Conditioning for Text-to-Image Synthesis
von: Ramakrishnan, Aashish Anantha, et al.
Veröffentlicht: (2024) -
FIM-LoRA: Task-Informative Rank Allocation for LoRA via Calibration-Time Gradient-Variance Estimation
von: Sathyavageeswaran, Ramakrishnan
Veröffentlicht: (2026)