BRIDGE: Predicting Human Task Completion Time From Model Performance
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Fengyuan, Gala, Jay, Nilaksh, Bahdanau, Dzmitry, Reddy, Siva, Larochelle, Hugo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Forecasting Downstream Performance of LLMs With Proxy Metrics
von: Patel, Arkil, et al.
Veröffentlicht: (2026)
von: Patel, Arkil, et al.
Veröffentlicht: (2026)
LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders
von: BehnamGhader, Parishad, et al.
Veröffentlicht: (2024)
von: BehnamGhader, Parishad, et al.
Veröffentlicht: (2024)
Are self-explanations from Large Language Models faithful?
von: Madsen, Andreas, et al.
Veröffentlicht: (2024)
von: Madsen, Andreas, et al.
Veröffentlicht: (2024)
Transformer-Based Multimodal Knowledge Graph Completion with Link-Aware Contexts
von: Ma, Haodi, et al.
Veröffentlicht: (2025)
von: Ma, Haodi, et al.
Veröffentlicht: (2025)
How to Get Your LLM to Generate Challenging Problems for Evaluation
von: Patel, Arkil, et al.
Veröffentlicht: (2025)
von: Patel, Arkil, et al.
Veröffentlicht: (2025)
A density estimation perspective on learning from pairwise human preferences
von: Dumoulin, Vincent, et al.
Veröffentlicht: (2023)
von: Dumoulin, Vincent, et al.
Veröffentlicht: (2023)
ROSA: Random Subspace Adaptation for Efficient Fine-Tuning
von: Hameed, Marawan Gamal Abdel, et al.
Veröffentlicht: (2024)
von: Hameed, Marawan Gamal Abdel, et al.
Veröffentlicht: (2024)
Predicting Task Performance with Context-aware Scaling Laws
von: Montgomery, Kyle, et al.
Veröffentlicht: (2025)
von: Montgomery, Kyle, et al.
Veröffentlicht: (2025)
LLMs can learn self-restraint through iterative self-reflection
von: Piché, Alexandre, et al.
Veröffentlicht: (2024)
von: Piché, Alexandre, et al.
Veröffentlicht: (2024)
Efficient Model Development through Fine-tuning Transfer
von: Lin, Pin-Jie, et al.
Veröffentlicht: (2025)
von: Lin, Pin-Jie, et al.
Veröffentlicht: (2025)
Collaborative Performance Prediction for Large Language Models
von: Zhang, Qiyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Qiyuan, et al.
Veröffentlicht: (2024)
Evaluating In-Context Learning of Libraries for Code Generation
von: Patel, Arkil, et al.
Veröffentlicht: (2023)
von: Patel, Arkil, et al.
Veröffentlicht: (2023)
SBSC: Step-By-Step Coding for Improving Mathematical Olympiad Performance
von: Singh, Kunal, et al.
Veröffentlicht: (2025)
von: Singh, Kunal, et al.
Veröffentlicht: (2025)
Alice in Wonderland: Simple Tasks Showing Complete Reasoning Breakdown in State-Of-the-Art Large Language Models
von: Nezhurina, Marianna, et al.
Veröffentlicht: (2024)
von: Nezhurina, Marianna, et al.
Veröffentlicht: (2024)
The Markovian Thinker: Architecture-Agnostic Linear Scaling of Reasoning
von: Aghajohari, Milad, et al.
Veröffentlicht: (2025)
von: Aghajohari, Milad, et al.
Veröffentlicht: (2025)
Introducing Verification Task of Set Consistency with Set-Consistency Energy Networks
von: Song, Mooho, et al.
Veröffentlicht: (2025)
von: Song, Mooho, et al.
Veröffentlicht: (2025)
Large Language Models in the Task of Automatic Validation of Text Classifier Predictions
von: Tsymbalov, Aleksandr, et al.
Veröffentlicht: (2025)
von: Tsymbalov, Aleksandr, et al.
Veröffentlicht: (2025)
Robustness as an Emergent Property of Task Performance
von: Ashury-Tahan, Shir, et al.
Veröffentlicht: (2026)
von: Ashury-Tahan, Shir, et al.
Veröffentlicht: (2026)
SG-Blend: Learning an Interpolation Between Improved Swish and GELU for Robust Neural Representations
von: Sarkar, Gaurav, et al.
Veröffentlicht: (2025)
von: Sarkar, Gaurav, et al.
Veröffentlicht: (2025)
Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance
von: Ye, Jiasheng, et al.
Veröffentlicht: (2024)
von: Ye, Jiasheng, et al.
Veröffentlicht: (2024)
Diffusion Language Models Can Perform Many Tasks with Scaling and Instruction-Finetuning
von: Ye, Jiasheng, et al.
Veröffentlicht: (2023)
von: Ye, Jiasheng, et al.
Veröffentlicht: (2023)
Observational Scaling Laws and the Predictability of Language Model Performance
von: Ruan, Yangjun, et al.
Veröffentlicht: (2024)
von: Ruan, Yangjun, et al.
Veröffentlicht: (2024)
Test-Time Scaling in Reasoning Models Is Not Effective for Knowledge-Intensive Tasks Yet
von: Zhao, James Xu, et al.
Veröffentlicht: (2025)
von: Zhao, James Xu, et al.
Veröffentlicht: (2025)
From Symbolic Tasks to Code Generation: Diversification Yields Better Task Performers
von: Zhang, Dylan, et al.
Veröffentlicht: (2024)
von: Zhang, Dylan, et al.
Veröffentlicht: (2024)
A Human-Like Reasoning Framework for Multi-Phases Planning Task with Large Language Models
von: Xie, Chengxing, et al.
Veröffentlicht: (2024)
von: Xie, Chengxing, et al.
Veröffentlicht: (2024)
HelpSteer3: Human-Annotated Feedback and Edit Data to Empower Inference-Time Scaling in Open-Ended General-Domain Tasks
von: Wang, Zhilin, et al.
Veröffentlicht: (2025)
von: Wang, Zhilin, et al.
Veröffentlicht: (2025)
Model-based Subsampling for Knowledge Graph Completion
von: Feng, Xincan, et al.
Veröffentlicht: (2023)
von: Feng, Xincan, et al.
Veröffentlicht: (2023)
Modular Multi-Task Learning for Chemical Reaction Prediction
von: Pang, Jiayun, et al.
Veröffentlicht: (2026)
von: Pang, Jiayun, et al.
Veröffentlicht: (2026)
Time-MQA: Time Series Multi-Task Question Answering with Context Enhancement
von: Kong, Yaxuan, et al.
Veröffentlicht: (2025)
von: Kong, Yaxuan, et al.
Veröffentlicht: (2025)
SafeArena: Evaluating the Safety of Autonomous Web Agents
von: Tur, Ada Defne, et al.
Veröffentlicht: (2025)
von: Tur, Ada Defne, et al.
Veröffentlicht: (2025)
Filter-then-Generate: Large Language Models with Structure-Text Adapter for Knowledge Graph Completion
von: Liu, Ben, et al.
Veröffentlicht: (2024)
von: Liu, Ben, et al.
Veröffentlicht: (2024)
Sparse Autoencoder Decomposition of Clinical Sequence Model Representations: Feature Complexity, Task Specialisation, and Mortality Prediction
von: Sainsbury, Chris, et al.
Veröffentlicht: (2026)
von: Sainsbury, Chris, et al.
Veröffentlicht: (2026)
Mind Your Step (by Step): Chain-of-Thought can Reduce Performance on Tasks where Thinking Makes Humans Worse
von: Liu, Ryan, et al.
Veröffentlicht: (2024)
von: Liu, Ryan, et al.
Veröffentlicht: (2024)
Task-Informed Anti-Curriculum by Masking Improves Downstream Performance on Text
von: Jarca, Andrei, et al.
Veröffentlicht: (2025)
von: Jarca, Andrei, et al.
Veröffentlicht: (2025)
Visualizing Uncertainty in Translation Tasks: An Evaluation of LLM Performance and Confidence Metrics
von: Park, Jin Hyun, et al.
Veröffentlicht: (2025)
von: Park, Jin Hyun, et al.
Veröffentlicht: (2025)
AgentRewardBench: Evaluating Automatic Evaluations of Web Agent Trajectories
von: Lù, Xing Han, et al.
Veröffentlicht: (2025)
von: Lù, Xing Han, et al.
Veröffentlicht: (2025)
ClinicRealm: Re-evaluating Large Language Models with Conventional Machine Learning for Non-Generative Clinical Prediction Tasks
von: Zhu, Yinghao, et al.
Veröffentlicht: (2024)
von: Zhu, Yinghao, et al.
Veröffentlicht: (2024)
Language Control Diffusion: Efficiently Scaling through Space, Time, and Tasks
von: Zhang, Edwin, et al.
Veröffentlicht: (2022)
von: Zhang, Edwin, et al.
Veröffentlicht: (2022)
Exploring the Trade-off Between Model Performance and Explanation Plausibility of Text Classifiers Using Human Rationales
von: Resck, Lucas E., et al.
Veröffentlicht: (2024)
von: Resck, Lucas E., et al.
Veröffentlicht: (2024)
Scaling Laws for Predicting Downstream Performance in LLMs
von: Chen, Yangyi, et al.
Veröffentlicht: (2024)
von: Chen, Yangyi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Forecasting Downstream Performance of LLMs With Proxy Metrics
von: Patel, Arkil, et al.
Veröffentlicht: (2026) -
LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders
von: BehnamGhader, Parishad, et al.
Veröffentlicht: (2024) -
Are self-explanations from Large Language Models faithful?
von: Madsen, Andreas, et al.
Veröffentlicht: (2024) -
Transformer-Based Multimodal Knowledge Graph Completion with Link-Aware Contexts
von: Ma, Haodi, et al.
Veröffentlicht: (2025) -
How to Get Your LLM to Generate Challenging Problems for Evaluation
von: Patel, Arkil, et al.
Veröffentlicht: (2025)