BRIDGE: Predicting Human Task Completion Time From Model Performance
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Fengyuan, Gala, Jay, Nilaksh, Bahdanau, Dzmitry, Reddy, Siva, Larochelle, Hugo |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Forecasting Downstream Performance of LLMs With Proxy Metrics
by: Patel, Arkil, et al.
Published: (2026)
by: Patel, Arkil, et al.
Published: (2026)
LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders
by: BehnamGhader, Parishad, et al.
Published: (2024)
by: BehnamGhader, Parishad, et al.
Published: (2024)
Are self-explanations from Large Language Models faithful?
by: Madsen, Andreas, et al.
Published: (2024)
by: Madsen, Andreas, et al.
Published: (2024)
Transformer-Based Multimodal Knowledge Graph Completion with Link-Aware Contexts
by: Ma, Haodi, et al.
Published: (2025)
by: Ma, Haodi, et al.
Published: (2025)
How to Get Your LLM to Generate Challenging Problems for Evaluation
by: Patel, Arkil, et al.
Published: (2025)
by: Patel, Arkil, et al.
Published: (2025)
A density estimation perspective on learning from pairwise human preferences
by: Dumoulin, Vincent, et al.
Published: (2023)
by: Dumoulin, Vincent, et al.
Published: (2023)
ROSA: Random Subspace Adaptation for Efficient Fine-Tuning
by: Hameed, Marawan Gamal Abdel, et al.
Published: (2024)
by: Hameed, Marawan Gamal Abdel, et al.
Published: (2024)
Predicting Task Performance with Context-aware Scaling Laws
by: Montgomery, Kyle, et al.
Published: (2025)
by: Montgomery, Kyle, et al.
Published: (2025)
LLMs can learn self-restraint through iterative self-reflection
by: Piché, Alexandre, et al.
Published: (2024)
by: Piché, Alexandre, et al.
Published: (2024)
Efficient Model Development through Fine-tuning Transfer
by: Lin, Pin-Jie, et al.
Published: (2025)
by: Lin, Pin-Jie, et al.
Published: (2025)
Collaborative Performance Prediction for Large Language Models
by: Zhang, Qiyuan, et al.
Published: (2024)
by: Zhang, Qiyuan, et al.
Published: (2024)
Evaluating In-Context Learning of Libraries for Code Generation
by: Patel, Arkil, et al.
Published: (2023)
by: Patel, Arkil, et al.
Published: (2023)
SBSC: Step-By-Step Coding for Improving Mathematical Olympiad Performance
by: Singh, Kunal, et al.
Published: (2025)
by: Singh, Kunal, et al.
Published: (2025)
Alice in Wonderland: Simple Tasks Showing Complete Reasoning Breakdown in State-Of-the-Art Large Language Models
by: Nezhurina, Marianna, et al.
Published: (2024)
by: Nezhurina, Marianna, et al.
Published: (2024)
The Markovian Thinker: Architecture-Agnostic Linear Scaling of Reasoning
by: Aghajohari, Milad, et al.
Published: (2025)
by: Aghajohari, Milad, et al.
Published: (2025)
Introducing Verification Task of Set Consistency with Set-Consistency Energy Networks
by: Song, Mooho, et al.
Published: (2025)
by: Song, Mooho, et al.
Published: (2025)
Large Language Models in the Task of Automatic Validation of Text Classifier Predictions
by: Tsymbalov, Aleksandr, et al.
Published: (2025)
by: Tsymbalov, Aleksandr, et al.
Published: (2025)
Robustness as an Emergent Property of Task Performance
by: Ashury-Tahan, Shir, et al.
Published: (2026)
by: Ashury-Tahan, Shir, et al.
Published: (2026)
SG-Blend: Learning an Interpolation Between Improved Swish and GELU for Robust Neural Representations
by: Sarkar, Gaurav, et al.
Published: (2025)
by: Sarkar, Gaurav, et al.
Published: (2025)
Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance
by: Ye, Jiasheng, et al.
Published: (2024)
by: Ye, Jiasheng, et al.
Published: (2024)
Diffusion Language Models Can Perform Many Tasks with Scaling and Instruction-Finetuning
by: Ye, Jiasheng, et al.
Published: (2023)
by: Ye, Jiasheng, et al.
Published: (2023)
Observational Scaling Laws and the Predictability of Language Model Performance
by: Ruan, Yangjun, et al.
Published: (2024)
by: Ruan, Yangjun, et al.
Published: (2024)
Test-Time Scaling in Reasoning Models Is Not Effective for Knowledge-Intensive Tasks Yet
by: Zhao, James Xu, et al.
Published: (2025)
by: Zhao, James Xu, et al.
Published: (2025)
From Symbolic Tasks to Code Generation: Diversification Yields Better Task Performers
by: Zhang, Dylan, et al.
Published: (2024)
by: Zhang, Dylan, et al.
Published: (2024)
A Human-Like Reasoning Framework for Multi-Phases Planning Task with Large Language Models
by: Xie, Chengxing, et al.
Published: (2024)
by: Xie, Chengxing, et al.
Published: (2024)
HelpSteer3: Human-Annotated Feedback and Edit Data to Empower Inference-Time Scaling in Open-Ended General-Domain Tasks
by: Wang, Zhilin, et al.
Published: (2025)
by: Wang, Zhilin, et al.
Published: (2025)
Model-based Subsampling for Knowledge Graph Completion
by: Feng, Xincan, et al.
Published: (2023)
by: Feng, Xincan, et al.
Published: (2023)
Modular Multi-Task Learning for Chemical Reaction Prediction
by: Pang, Jiayun, et al.
Published: (2026)
by: Pang, Jiayun, et al.
Published: (2026)
Time-MQA: Time Series Multi-Task Question Answering with Context Enhancement
by: Kong, Yaxuan, et al.
Published: (2025)
by: Kong, Yaxuan, et al.
Published: (2025)
SafeArena: Evaluating the Safety of Autonomous Web Agents
by: Tur, Ada Defne, et al.
Published: (2025)
by: Tur, Ada Defne, et al.
Published: (2025)
Filter-then-Generate: Large Language Models with Structure-Text Adapter for Knowledge Graph Completion
by: Liu, Ben, et al.
Published: (2024)
by: Liu, Ben, et al.
Published: (2024)
Sparse Autoencoder Decomposition of Clinical Sequence Model Representations: Feature Complexity, Task Specialisation, and Mortality Prediction
by: Sainsbury, Chris, et al.
Published: (2026)
by: Sainsbury, Chris, et al.
Published: (2026)
Mind Your Step (by Step): Chain-of-Thought can Reduce Performance on Tasks where Thinking Makes Humans Worse
by: Liu, Ryan, et al.
Published: (2024)
by: Liu, Ryan, et al.
Published: (2024)
Task-Informed Anti-Curriculum by Masking Improves Downstream Performance on Text
by: Jarca, Andrei, et al.
Published: (2025)
by: Jarca, Andrei, et al.
Published: (2025)
Visualizing Uncertainty in Translation Tasks: An Evaluation of LLM Performance and Confidence Metrics
by: Park, Jin Hyun, et al.
Published: (2025)
by: Park, Jin Hyun, et al.
Published: (2025)
AgentRewardBench: Evaluating Automatic Evaluations of Web Agent Trajectories
by: Lù, Xing Han, et al.
Published: (2025)
by: Lù, Xing Han, et al.
Published: (2025)
ClinicRealm: Re-evaluating Large Language Models with Conventional Machine Learning for Non-Generative Clinical Prediction Tasks
by: Zhu, Yinghao, et al.
Published: (2024)
by: Zhu, Yinghao, et al.
Published: (2024)
Language Control Diffusion: Efficiently Scaling through Space, Time, and Tasks
by: Zhang, Edwin, et al.
Published: (2022)
by: Zhang, Edwin, et al.
Published: (2022)
Exploring the Trade-off Between Model Performance and Explanation Plausibility of Text Classifiers Using Human Rationales
by: Resck, Lucas E., et al.
Published: (2024)
by: Resck, Lucas E., et al.
Published: (2024)
Scaling Laws for Predicting Downstream Performance in LLMs
by: Chen, Yangyi, et al.
Published: (2024)
by: Chen, Yangyi, et al.
Published: (2024)
Similar Items
-
Forecasting Downstream Performance of LLMs With Proxy Metrics
by: Patel, Arkil, et al.
Published: (2026) -
LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders
by: BehnamGhader, Parishad, et al.
Published: (2024) -
Are self-explanations from Large Language Models faithful?
by: Madsen, Andreas, et al.
Published: (2024) -
Transformer-Based Multimodal Knowledge Graph Completion with Link-Aware Contexts
by: Ma, Haodi, et al.
Published: (2025) -
How to Get Your LLM to Generate Challenging Problems for Evaluation
by: Patel, Arkil, et al.
Published: (2025)