UQ: Assessing Language Models on Unsolved Questions
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Nie, Fan, Liu, Ken Ziyu, Wang, Zihao, Sun, Rui, Liu, Wei, Shi, Weijia, Yao, Huaxiu, Zhang, Linjun, Ng, Andrew Y., Zou, James, Koyejo, Sanmi, Choi, Yejin, Liang, Percy, Muennighoff, Niklas |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Understanding Adversarial Transfer: Why Representation-Space Attacks Fail Where Data-Space Attacks Succeed
par: Gupta, Isha, et autres
Publié: (2025)
par: Gupta, Isha, et autres
Publié: (2025)
On Fairness of Low-Rank Adaptation of Large Models
par: Ding, Zhoujie, et autres
Publié: (2024)
par: Ding, Zhoujie, et autres
Publié: (2024)
FactTest: Factuality Testing in Large Language Models with Finite-Sample and Distribution-Free Guarantees
par: Nie, Fan, et autres
Publié: (2024)
par: Nie, Fan, et autres
Publié: (2024)
Language Models May Verbatim Complete Text They Were Not Explicitly Trained On
par: Liu, Ken Ziyu, et autres
Publié: (2025)
par: Liu, Ken Ziyu, et autres
Publié: (2025)
When Visualizing is the First Step to Reasoning: MIRA, a Benchmark for Visual Chain-of-Thought
par: Zhou, Yiyang, et autres
Publié: (2025)
par: Zhou, Yiyang, et autres
Publié: (2025)
Extracting books from production language models
par: Ahmed, Ahmed, et autres
Publié: (2026)
par: Ahmed, Ahmed, et autres
Publié: (2026)
Causally Inspired Regularization Enables Domain General Representations
par: Salaudeen, Olawale, et autres
Publié: (2024)
par: Salaudeen, Olawale, et autres
Publié: (2024)
Let's Measure Information Step-by-Step: AI-Based Evaluation Beyond Vibes
par: Robertson, Zachary, et autres
Publié: (2025)
par: Robertson, Zachary, et autres
Publié: (2025)
CURE: Cultural Understanding and Reasoning Evaluation - A Framework for "Thick" Culture Alignment Evaluation in LLMs
par: Vo, Truong, et autres
Publié: (2025)
par: Vo, Truong, et autres
Publié: (2025)
Reliable and Efficient Amortized Model-based Evaluation
par: Truong, Sang, et autres
Publié: (2025)
par: Truong, Sang, et autres
Publié: (2025)
SpecEval: Evaluating Model Adherence to Behavior Specifications
par: Ahmed, Ahmed, et autres
Publié: (2025)
par: Ahmed, Ahmed, et autres
Publié: (2025)
Investigating Data Contamination for Pre-training Language Models
par: Jiang, Minhao, et autres
Publié: (2024)
par: Jiang, Minhao, et autres
Publié: (2024)
A Framework for Objective-Driven Dynamical Stochastic Fields
par: Zhang, Yibo Jacky, et autres
Publié: (2025)
par: Zhang, Yibo Jacky, et autres
Publié: (2025)
From Passive to Active Reasoning: Can Large Language Models Ask the Right Questions under Incomplete Information?
par: Zhou, Zhanke, et autres
Publié: (2025)
par: Zhou, Zhanke, et autres
Publié: (2025)
Humanline: Online Alignment as Perceptual Loss
par: Liu, Sijia, et autres
Publié: (2025)
par: Liu, Sijia, et autres
Publié: (2025)
MMed-RAG: Versatile Multimodal RAG System for Medical Vision Language Models
par: Xia, Peng, et autres
Publié: (2024)
par: Xia, Peng, et autres
Publié: (2024)
In-Context Learning of Energy Functions
par: Schaeffer, Rylan, et autres
Publié: (2024)
par: Schaeffer, Rylan, et autres
Publié: (2024)
Discovering Implicit Large Language Model Alignment Objectives
par: Chen, Edward, et autres
Publié: (2026)
par: Chen, Edward, et autres
Publié: (2026)
High-Dimensional Markov-switching Ordinary Differential Processes
par: Tsai, Katherine, et autres
Publié: (2024)
par: Tsai, Katherine, et autres
Publié: (2024)
Distributional Machine Unlearning via Selective Data Removal
par: Allouah, Youssef, et autres
Publié: (2025)
par: Allouah, Youssef, et autres
Publié: (2025)
SCENEBench: An Audio Understanding Benchmark Grounded in Assistive and Industrial Use Cases
par: Iyer, Laya, et autres
Publié: (2026)
par: Iyer, Laya, et autres
Publié: (2026)
HiFA: High-fidelity Text-to-3D Generation with Advanced Diffusion Guidance
par: Zhu, Junzhe, et autres
Publié: (2023)
par: Zhu, Junzhe, et autres
Publié: (2023)
Unsolved Problems in Spectral Graph Theory
par: Liu, Lele, et autres
Publié: (2023)
par: Liu, Lele, et autres
Publié: (2023)
Reasoning Models Don't Just Think Longer, They Move Differently
par: Gjølbye, Anders, et autres
Publié: (2026)
par: Gjølbye, Anders, et autres
Publié: (2026)
Is Backpropagation Optimal? When Synthetic Gradients Improve Sample Efficiency
par: Zhang, Yibo Jacky, et autres
Publié: (2026)
par: Zhang, Yibo Jacky, et autres
Publié: (2026)
The Inadequacy of Offline LLM Evaluations: A Need to Account for Personalization in Model Behavior
par: Wang, Angelina, et autres
Publié: (2025)
par: Wang, Angelina, et autres
Publié: (2025)
Principled Federated Domain Adaptation: Gradient Projection and Auto-Weighting
par: Jiang, Enyi, et autres
Publié: (2023)
par: Jiang, Enyi, et autres
Publié: (2023)
Distribution-Free Fair Federated Learning with Small Samples
par: Yin, Qichuan, et autres
Publié: (2024)
par: Yin, Qichuan, et autres
Publié: (2024)
HEART: A Unified Benchmark for Assessing Humans and LLMs in Emotional Support Dialogue
par: Iyer, Laya, et autres
Publié: (2026)
par: Iyer, Laya, et autres
Publié: (2026)
Why Do Safety Guardrails Degrade Across Languages?
par: Zhang, Max, et autres
Publié: (2026)
par: Zhang, Max, et autres
Publié: (2026)
Are Domain Generalization Benchmarks with Accuracy on the Line Misspecified?
par: Salaudeen, Olawale, et autres
Publié: (2025)
par: Salaudeen, Olawale, et autres
Publié: (2025)
Pretraining Scaling Laws for Generative Evaluations of Language Models
par: Schaeffer, Rylan, et autres
Publié: (2025)
par: Schaeffer, Rylan, et autres
Publié: (2025)
The Utility and Complexity of in- and out-of-Distribution Machine Unlearning
par: Allouah, Youssef, et autres
Publié: (2024)
par: Allouah, Youssef, et autres
Publié: (2024)
Logits are All We Need to Adapt Closed Models
par: Hiranandani, Gaurush, et autres
Publié: (2025)
par: Hiranandani, Gaurush, et autres
Publié: (2025)
Invariant Aggregator for Defending against Federated Backdoor Attacks
par: Wang, Xiaoyang, et autres
Publié: (2022)
par: Wang, Xiaoyang, et autres
Publié: (2022)
Stop Automating Peer Review Without Rigorous Evaluation
par: Baumann, Joachim, et autres
Publié: (2026)
par: Baumann, Joachim, et autres
Publié: (2026)
C-Pack: Packed Resources For General Chinese Embeddings
par: Xiao, Shitao, et autres
Publié: (2023)
par: Xiao, Shitao, et autres
Publié: (2023)
Mixture of Rationale: Multi-Modal Reasoning Mixture for Visual Question Answering
par: Li, Tao, et autres
Publié: (2024)
par: Li, Tao, et autres
Publié: (2024)
Conformal Prediction for Deep Classifier via Label Ranking
par: Huang, Jianguo, et autres
Publié: (2023)
par: Huang, Jianguo, et autres
Publié: (2023)
Fairness through Difference Awareness: Measuring Desired Group Discrimination in LLMs
par: Wang, Angelina, et autres
Publié: (2025)
par: Wang, Angelina, et autres
Publié: (2025)
Documents similaires
-
Understanding Adversarial Transfer: Why Representation-Space Attacks Fail Where Data-Space Attacks Succeed
par: Gupta, Isha, et autres
Publié: (2025) -
On Fairness of Low-Rank Adaptation of Large Models
par: Ding, Zhoujie, et autres
Publié: (2024) -
FactTest: Factuality Testing in Large Language Models with Finite-Sample and Distribution-Free Guarantees
par: Nie, Fan, et autres
Publié: (2024) -
Language Models May Verbatim Complete Text They Were Not Explicitly Trained On
par: Liu, Ken Ziyu, et autres
Publié: (2025) -
When Visualizing is the First Step to Reasoning: MIRA, a Benchmark for Visual Chain-of-Thought
par: Zhou, Yiyang, et autres
Publié: (2025)