Beyond Next Word Prediction: Developing Comprehensive Evaluation Frameworks for measuring LLM performance on real world applications
Fuente:
arXiv
Saved in:
| Main Authors: | Agrawal, Vishakha, Chaudhury, Archie, Agrawal, Shreya |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Addressing LLM Diversity by Infusing Random Concepts
by: Agrawal, Pulin, et al.
Published: (2026)
by: Agrawal, Pulin, et al.
Published: (2026)
Alignment is Localized: A Causal Probe into Preference Layers
by: Chaudhury, Archie
Published: (2025)
by: Chaudhury, Archie
Published: (2025)
Illuminate: A novel approach for depression detection with explainable analysis and proactive therapy using prompt engineering
by: Agrawal, Aryan
Published: (2024)
by: Agrawal, Aryan
Published: (2024)
Joint Detection of Fraud and Concept Drift inOnline Conversations with LLM-Assisted Judgment
by: Senol, Ali, et al.
Published: (2025)
by: Senol, Ali, et al.
Published: (2025)
Counting Clues: A Lightweight Probabilistic Baseline Can Match an LLM
by: Jia, Furong, et al.
Published: (2025)
by: Jia, Furong, et al.
Published: (2025)
Vidur: A Large-Scale Simulation Framework For LLM Inference
by: Agrawal, Amey, et al.
Published: (2024)
by: Agrawal, Amey, et al.
Published: (2024)
An Examination of the Robustness of Reference-Free Image Captioning Evaluation Metrics
by: Ahmadi, Saba, et al.
Published: (2023)
by: Ahmadi, Saba, et al.
Published: (2023)
Etalon: Holistic Performance Evaluation Framework for LLM Inference Systems
by: Agrawal, Amey, et al.
Published: (2024)
by: Agrawal, Amey, et al.
Published: (2024)
Spiffy: Multiplying Diffusion LLM Acceleration via Lossless Speculative Decoding
by: Agrawal, Sudhanshu, et al.
Published: (2025)
by: Agrawal, Sudhanshu, et al.
Published: (2025)
NextLocLLM: Location Semantics Modeling and Coordinate-Based Next Location Prediction with LLMs
by: Liu, Shuai, et al.
Published: (2024)
by: Liu, Shuai, et al.
Published: (2024)
Beyond Next Token Prediction: Patch-Level Training for Large Language Models
by: Shao, Chenze, et al.
Published: (2024)
by: Shao, Chenze, et al.
Published: (2024)
CCRS: A Zero-Shot LLM-as-a-Judge Framework for Comprehensive RAG Evaluation
by: Muhamed, Aashiq
Published: (2025)
by: Muhamed, Aashiq
Published: (2025)
Improving Automatic VQA Evaluation Using Large Language Models
by: Mañas, Oscar, et al.
Published: (2023)
by: Mañas, Oscar, et al.
Published: (2023)
Code Comprehension then Auditing for Unsupervised LLM Evaluation
by: Patel, Bhrij, et al.
Published: (2024)
by: Patel, Bhrij, et al.
Published: (2024)
Emergence and Effectiveness of Task Vectors in In-Context Learning: An Encoder Decoder Perspective
by: Han, Seungwook, et al.
Published: (2024)
by: Han, Seungwook, et al.
Published: (2024)
Training Language Models via Neural Cellular Automata
by: Lee, Dan, et al.
Published: (2026)
by: Lee, Dan, et al.
Published: (2026)
Beyond Next-Token Prediction: A Performance Characterization of Diffusion versus Autoregressive Language Models
by: Kim, Minseo, et al.
Published: (2025)
by: Kim, Minseo, et al.
Published: (2025)
HLDC: Hindi Legal Documents Corpus
by: Kapoor, Arnav, et al.
Published: (2022)
by: Kapoor, Arnav, et al.
Published: (2022)
Generation Constraint Scaling Can Mitigate Hallucination
by: Kollias, Georgios, et al.
Published: (2024)
by: Kollias, Georgios, et al.
Published: (2024)
TreeCut: A Synthetic Unanswerable Math Word Problem Dataset for LLM Hallucination Evaluation
by: Ouyang, Jialin
Published: (2025)
by: Ouyang, Jialin
Published: (2025)
Cautious Next Token Prediction
by: Wang, Yizhou, et al.
Published: (2025)
by: Wang, Yizhou, et al.
Published: (2025)
Ladder: A Model-Agnostic Framework Boosting LLM-based Machine Translation to the Next Level
by: Feng, Zhaopeng, et al.
Published: (2024)
by: Feng, Zhaopeng, et al.
Published: (2024)
Value Augmented Sampling for Language Model Alignment and Personalization
by: Han, Seungwook, et al.
Published: (2024)
by: Han, Seungwook, et al.
Published: (2024)
But what is your honest answer? Aiding LLM-judges with honest alternatives using steering vectors
by: Eshuijs, Leon, et al.
Published: (2025)
by: Eshuijs, Leon, et al.
Published: (2025)
Can't Remember Details in Long Documents? You Need Some R&R
by: Agrawal, Devanshu, et al.
Published: (2024)
by: Agrawal, Devanshu, et al.
Published: (2024)
Personas within Parameters: Fine-Tuning Small Language Models with Low-Rank Adapters to Mimic User Behaviors
by: Thakur, Himanshu, et al.
Published: (2025)
by: Thakur, Himanshu, et al.
Published: (2025)
Beyond LLM-as-a-Judge: Deterministic Metrics for Multilingual Generative Text Evaluation
by: Alam, Firoj, et al.
Published: (2026)
by: Alam, Firoj, et al.
Published: (2026)
Easy2Hard-Bench: Standardized Difficulty Labels for Profiling LLM Performance and Generalization
by: Ding, Mucong, et al.
Published: (2024)
by: Ding, Mucong, et al.
Published: (2024)
Translate Smart, not Hard: Cascaded Translation Systems with Quality-Aware Deferral
by: Farinhas, António, et al.
Published: (2025)
by: Farinhas, António, et al.
Published: (2025)
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words
by: Minegishi, Gouki, et al.
Published: (2025)
by: Minegishi, Gouki, et al.
Published: (2025)
From Words to Actions: Unveiling the Theoretical Underpinnings of LLM-Driven Autonomous Systems
by: He, Jianliang, et al.
Published: (2024)
by: He, Jianliang, et al.
Published: (2024)
A Data-Centric Approach To Generate Faithful and High Quality Patient Summaries with Large Language Models
by: Hegselmann, Stefan, et al.
Published: (2024)
by: Hegselmann, Stefan, et al.
Published: (2024)
A Law of Next-Token Prediction in Large Language Models
by: He, Hangfeng, et al.
Published: (2024)
by: He, Hangfeng, et al.
Published: (2024)
Interpretable Next-token Prediction via the Generalized Induction Head
by: Kim, Eunji, et al.
Published: (2024)
by: Kim, Eunji, et al.
Published: (2024)
ProtLLM: An Interleaved Protein-Language LLM with Protein-as-Word Pre-Training
by: Zhuo, Le, et al.
Published: (2024)
by: Zhuo, Le, et al.
Published: (2024)
DeduCE: Deductive Consistency as a Framework to Evaluate LLM Reasoning
by: Pandey, Atharva, et al.
Published: (2025)
by: Pandey, Atharva, et al.
Published: (2025)
Needle in the Haystack for Memory Based Large Language Models
by: Nelson, Elliot, et al.
Published: (2024)
by: Nelson, Elliot, et al.
Published: (2024)
Word-Sequence Entropy: Towards Uncertainty Estimation in Free-Form Medical Question Answering Applications and Beyond
by: Wang, Zhiyuan, et al.
Published: (2024)
by: Wang, Zhiyuan, et al.
Published: (2024)
DeceptionBench: A Comprehensive Benchmark for AI Deception Behaviors in Real-world Scenarios
by: Huang, Yao, et al.
Published: (2025)
by: Huang, Yao, et al.
Published: (2025)
PRECISE: Reducing the Bias of LLM Evaluations Using Prediction-Powered Ranking Estimation
by: Divekar, Abhishek, et al.
Published: (2026)
by: Divekar, Abhishek, et al.
Published: (2026)
Similar Items
-
Addressing LLM Diversity by Infusing Random Concepts
by: Agrawal, Pulin, et al.
Published: (2026) -
Alignment is Localized: A Causal Probe into Preference Layers
by: Chaudhury, Archie
Published: (2025) -
Illuminate: A novel approach for depression detection with explainable analysis and proactive therapy using prompt engineering
by: Agrawal, Aryan
Published: (2024) -
Joint Detection of Fraud and Concept Drift inOnline Conversations with LLM-Assisted Judgment
by: Senol, Ali, et al.
Published: (2025) -
Counting Clues: A Lightweight Probabilistic Baseline Can Match an LLM
by: Jia, Furong, et al.
Published: (2025)