Third-Party Language Model Performance Prediction from Instruction
Fuente:
arXiv
Saved in:
| Main Authors: | Nadkarni, Rahul, Wang, Yizhong, Smith, Noah A. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Rewriting History: A Recipe for Interventional Analyses to Study Data Effects on Model Behavior
by: Nadkarni, Rahul, et al.
Published: (2025)
by: Nadkarni, Rahul, et al.
Published: (2025)
Set the Clock: Temporal Alignment of Pretrained Language Models
by: Zhao, Bowen, et al.
Published: (2024)
by: Zhao, Bowen, et al.
Published: (2024)
Tuning Language Models by Proxy
by: Liu, Alisa, et al.
Published: (2024)
by: Liu, Alisa, et al.
Published: (2024)
Attacks on Third-Party APIs of Large Language Models
by: Zhao, Wanru, et al.
Published: (2024)
by: Zhao, Wanru, et al.
Published: (2024)
HREF: Human Response-Guided Evaluation of Instruction Following in Language Models
by: Lyu, Xinxi, et al.
Published: (2024)
by: Lyu, Xinxi, et al.
Published: (2024)
Can Language Models Act as Knowledge Bases at Scale?
by: He, Qiyuan, et al.
Published: (2024)
by: He, Qiyuan, et al.
Published: (2024)
Online vs Offline: A Comparative Study of First-Party and Third-Party Evaluations of Social Chatbots
by: Svikhnushina, Ekaterina, et al.
Published: (2024)
by: Svikhnushina, Ekaterina, et al.
Published: (2024)
Time is Encoded in the Weights of Finetuned Language Models
by: Nylund, Kai, et al.
Published: (2023)
by: Nylund, Kai, et al.
Published: (2023)
Improving the Accuracy and Efficiency of Legal Document Tagging with Large Language Models and Instruction Prompts
by: Johnson, Emily, et al.
Published: (2025)
by: Johnson, Emily, et al.
Published: (2025)
Sample, Don't Search: Rethinking Test-Time Alignment for Language Models
by: Faria, Gonçalo, et al.
Published: (2025)
by: Faria, Gonçalo, et al.
Published: (2025)
Enhancing Multi-Agent Consensus through Third-Party LLM Integration: Analyzing Uncertainty and Mitigating Hallucinations in Large Language Models
by: Duan, Zhihua, et al.
Published: (2024)
by: Duan, Zhihua, et al.
Published: (2024)
Long Context Alignment with Short Instructions and Synthesized Positions
by: Wu, Wenhao, et al.
Published: (2024)
by: Wu, Wenhao, et al.
Published: (2024)
How Stylistic Similarity Shapes Preferences in Dialogue Dataset with User and Third Party Evaluations
by: Numaya, Ikumi, et al.
Published: (2025)
by: Numaya, Ikumi, et al.
Published: (2025)
Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback
by: Ivison, Hamish, et al.
Published: (2024)
by: Ivison, Hamish, et al.
Published: (2024)
Multi-Party Supervised Fine-tuning of Language Models for Multi-Party Dialogue Generation
by: Wang, Xiaoyu, et al.
Published: (2024)
by: Wang, Xiaoyu, et al.
Published: (2024)
BTR: Binary Token Representations for Efficient Retrieval Augmented Language Models
by: Cao, Qingqing, et al.
Published: (2023)
by: Cao, Qingqing, et al.
Published: (2023)
How Performance Pressure Influences AI-Assisted Decision Making
by: Haduong, Nikita, et al.
Published: (2024)
by: Haduong, Nikita, et al.
Published: (2024)
Sampling from Your Language Model One Byte at a Time
by: Hayase, Jonathan, et al.
Published: (2025)
by: Hayase, Jonathan, et al.
Published: (2025)
Measuring Social Biases in Masked Language Models by Proxy of Prediction Quality
by: Zalkikar, Rahul, et al.
Published: (2024)
by: Zalkikar, Rahul, et al.
Published: (2024)
Large Language Models Naively Recover Ethnicity from Individual Records
by: Dasanaike, Noah
Published: (2026)
by: Dasanaike, Noah
Published: (2026)
Evaluating $n$-Gram Novelty of Language Models Using Rusty-DAWG
by: Merrill, William, et al.
Published: (2024)
by: Merrill, William, et al.
Published: (2024)
Rethinking LLM Watermark Detection in Black-Box Settings: A Non-Intrusive Third-Party Framework
by: Wang, Zhuoshang, et al.
Published: (2026)
by: Wang, Zhuoshang, et al.
Published: (2026)
EvalTree: Profiling Language Model Weaknesses via Hierarchical Capability Trees
by: Zeng, Zhiyuan, et al.
Published: (2025)
by: Zeng, Zhiyuan, et al.
Published: (2025)
Demystifying Prompts in Language Models via Perplexity Estimation
by: Gonen, Hila, et al.
Published: (2022)
by: Gonen, Hila, et al.
Published: (2022)
Fine-grained Hallucination Detection and Editing for Language Models
by: Mishra, Abhika, et al.
Published: (2024)
by: Mishra, Abhika, et al.
Published: (2024)
On Linear Representations and Pretraining Data Frequency in Language Models
by: Merullo, Jack, et al.
Published: (2025)
by: Merullo, Jack, et al.
Published: (2025)
ComPO: Community Preferences for Language Model Personalization
by: Kumar, Sachin, et al.
Published: (2024)
by: Kumar, Sachin, et al.
Published: (2024)
Still Between Us? Evaluating and Improving Voice Assistant Robustness to Third-Party Interruptions
by: Lee, Dongwook, et al.
Published: (2026)
by: Lee, Dongwook, et al.
Published: (2026)
Leveraging Large Language Models for Predictive Analysis of Human Misery
by: Seal, Bishanka, et al.
Published: (2025)
by: Seal, Bishanka, et al.
Published: (2025)
Using Embedding Models to Improve Probabilistic Race Prediction
by: Dasanaike, Noah, et al.
Published: (2026)
by: Dasanaike, Noah, et al.
Published: (2026)
LegiGPT: Party Politics and Transport Policy with Large Language Model
by: Yun, Hyunsoo, et al.
Published: (2025)
by: Yun, Hyunsoo, et al.
Published: (2025)
Does Liking Yellow Imply Driving a School Bus? Semantic Leakage in Language Models
by: Gonen, Hila, et al.
Published: (2024)
by: Gonen, Hila, et al.
Published: (2024)
Limitation Learning: Catching Adverse Dialog with GAIL
by: Kasmanoff, Noah, et al.
Published: (2025)
by: Kasmanoff, Noah, et al.
Published: (2025)
Hybrid Preferences: Learning to Route Instances for Human vs. AI Feedback
by: Miranda, Lester James V., et al.
Published: (2024)
by: Miranda, Lester James V., et al.
Published: (2024)
LIFBench: Evaluating the Instruction Following Performance and Stability of Large Language Models in Long-Context Scenarios
by: Wu, Xiaodong, et al.
Published: (2024)
by: Wu, Xiaodong, et al.
Published: (2024)
Instruction-Following Pruning for Large Language Models
by: Hou, Bairu, et al.
Published: (2025)
by: Hou, Bairu, et al.
Published: (2025)
Summarization-Based Document IDs for Generative Retrieval with Language Models
by: Li, Haoxin, et al.
Published: (2023)
by: Li, Haoxin, et al.
Published: (2023)
Breaking the Curse of Multilinguality with Cross-lingual Expert Language Models
by: Blevins, Terra, et al.
Published: (2024)
by: Blevins, Terra, et al.
Published: (2024)
Resilience of Large Language Models for Noisy Instructions
by: Wang, Bin, et al.
Published: (2024)
by: Wang, Bin, et al.
Published: (2024)
Chain-of-Instructions: Compositional Instruction Tuning on Large Language Models
by: Hayati, Shirley Anugrah, et al.
Published: (2024)
by: Hayati, Shirley Anugrah, et al.
Published: (2024)
Similar Items
-
Rewriting History: A Recipe for Interventional Analyses to Study Data Effects on Model Behavior
by: Nadkarni, Rahul, et al.
Published: (2025) -
Set the Clock: Temporal Alignment of Pretrained Language Models
by: Zhao, Bowen, et al.
Published: (2024) -
Tuning Language Models by Proxy
by: Liu, Alisa, et al.
Published: (2024) -
Attacks on Third-Party APIs of Large Language Models
by: Zhao, Wanru, et al.
Published: (2024) -
HREF: Human Response-Guided Evaluation of Instruction Following in Language Models
by: Lyu, Xinxi, et al.
Published: (2024)