Pearls from Pebbles: Improved Confidence Functions for Auto-labeling
Fuente:
arXiv
Guardado en:
| Autores principales: | Vishwakarma, Harit, Reid, Chen, Tay, Sui Jiet, Namburi, Satya Sai Srinath, Sala, Frederic, Vinayak, Ramya Korlakai |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Promises and Pitfalls of Threshold-based Auto-labeling
por: Vishwakarma, Harit, et al.
Publicado: (2022)
por: Vishwakarma, Harit, et al.
Publicado: (2022)
Taming False Positives in Out-of-Distribution Detection with Human Feedback
por: Vishwakarma, Harit, et al.
Publicado: (2024)
por: Vishwakarma, Harit, et al.
Publicado: (2024)
Adaptive Scoring and Thresholding with Human Feedback for Robust Out-of-Distribution Detection
por: Yamada, Daisuke, et al.
Publicado: (2025)
por: Yamada, Daisuke, et al.
Publicado: (2025)
CARE: Confounder-Aware Aggregation for Reliable LLM Evaluation
por: Zhao, Jitian, et al.
Publicado: (2026)
por: Zhao, Jitian, et al.
Publicado: (2026)
Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation
por: Huang, Tzu-Heng, et al.
Publicado: (2025)
por: Huang, Tzu-Heng, et al.
Publicado: (2025)
GRASP: Graph Agentic Search over Propositions for Multi-hop Question Answering
por: Jenkins, Stockton, et al.
Publicado: (2026)
por: Jenkins, Stockton, et al.
Publicado: (2026)
OTTER: Effortless Label Distribution Adaptation of Zero-shot Models
por: Shin, Changho, et al.
Publicado: (2024)
por: Shin, Changho, et al.
Publicado: (2024)
Pretrained Hybrids with MAD Skills
por: Roberts, Nicholas, et al.
Publicado: (2024)
por: Roberts, Nicholas, et al.
Publicado: (2024)
Metric Learning in an RKHS
por: Tatli, Gokcan, et al.
Publicado: (2025)
por: Tatli, Gokcan, et al.
Publicado: (2025)
LETS Forecast: Learning Embedology for Time Series Forecasting
por: Majeedi, Abrar, et al.
Publicado: (2025)
por: Majeedi, Abrar, et al.
Publicado: (2025)
Learning from Less: Measuring the Effectiveness of RLVR in Low Data and Compute Regimes
por: Bauer, Justin, et al.
Publicado: (2026)
por: Bauer, Justin, et al.
Publicado: (2026)
Is Conformal Factuality for RAG-based LLMs Robust? Novel Metrics and Systematic Insights
por: Chen, Yi, et al.
Publicado: (2026)
por: Chen, Yi, et al.
Publicado: (2026)
R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training
por: Ge, Albert, et al.
Publicado: (2025)
por: Ge, Albert, et al.
Publicado: (2025)
Brain alignment of reasoning and action representations from vision-language and action models during naturalistic gameplay
por: Oota, Subba Reddy, et al.
Publicado: (2026)
por: Oota, Subba Reddy, et al.
Publicado: (2026)
Linguistic properties and model scale in brain encoding: from small to compressed language models
por: Oota, Subba Reddy, et al.
Publicado: (2026)
por: Oota, Subba Reddy, et al.
Publicado: (2026)
Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain)
por: Oota, Subba Reddy, et al.
Publicado: (2025)
por: Oota, Subba Reddy, et al.
Publicado: (2025)
Task-conditioned probing of instruction-tuned multimodal LLMs: Region-specific brain alignment patterns under naturalistic stimuli
por: Oota, Subba Reddy, et al.
Publicado: (2025)
por: Oota, Subba Reddy, et al.
Publicado: (2025)
Multimodal Data Curation via Object Detection and Filter Ensembles
por: Huang, Tzu-Heng, et al.
Publicado: (2024)
por: Huang, Tzu-Heng, et al.
Publicado: (2024)
Prune 'n Predict: Optimizing LLM Decision-making with Conformal Prediction
por: Vishwakarma, Harit, et al.
Publicado: (2024)
por: Vishwakarma, Harit, et al.
Publicado: (2024)
Bridging Lifelong and Multi-Task Representation Learning via Algorithm and Complexity Measure
por: Wang, Zhi, et al.
Publicado: (2025)
por: Wang, Zhi, et al.
Publicado: (2025)
Metric Learning from Limited Pairwise Preference Comparisons
por: Wang, Zhi, et al.
Publicado: (2024)
por: Wang, Zhi, et al.
Publicado: (2024)
Tabby: A Language Model Architecture for Tabular and Structured Data Synthesis
por: Cromp, Sonia, et al.
Publicado: (2025)
por: Cromp, Sonia, et al.
Publicado: (2025)
Causal Spherical Hypergraph Networks for Modelling Social Uncertainty
por: Harit, Anoushka, et al.
Publicado: (2025)
por: Harit, Anoushka, et al.
Publicado: (2025)
PAL: Pluralistic Alignment Framework for Learning from Heterogeneous Preferences
por: Chen, Daiwei, et al.
Publicado: (2024)
por: Chen, Daiwei, et al.
Publicado: (2024)
RicciFlowRec: A Geometric Root Cause Recommender Using Ricci Curvature on Financial Graphs
por: Sun, Zhongtian, et al.
Publicado: (2025)
por: Sun, Zhongtian, et al.
Publicado: (2025)
Actionable Interpretability via Causal Hypergraphs: Unravelling Batch Size Effects in Deep Learning
por: Sun, Zhongtian, et al.
Publicado: (2025)
por: Sun, Zhongtian, et al.
Publicado: (2025)
RICA2: Rubric-Informed, Calibrated Assessment of Actions
por: Majeedi, Abrar, et al.
Publicado: (2024)
por: Majeedi, Abrar, et al.
Publicado: (2024)
Maximizing Confidence Alone Improves Reasoning
por: Prabhudesai, Mihir, et al.
Publicado: (2025)
por: Prabhudesai, Mihir, et al.
Publicado: (2025)
Auto FAQ Generation
por: Kalvakolanu, Anjaneya Teja, et al.
Publicado: (2024)
por: Kalvakolanu, Anjaneya Teja, et al.
Publicado: (2024)
Grounded Token Initialization for New Vocabulary in LMs for Generative Recommendation
por: Chen, Daiwei, et al.
Publicado: (2026)
por: Chen, Daiwei, et al.
Publicado: (2026)
ManifoldMind: Dynamic Hyperbolic Reasoning for Trustworthy Recommendations
por: Harit, Anoushka, et al.
Publicado: (2025)
por: Harit, Anoushka, et al.
Publicado: (2025)
Dual-Signal Adaptive KV-Cache Optimization for Long-Form Video Understanding in Vision-Language Models
por: Sai, Vishnu, et al.
Publicado: (2026)
por: Sai, Vishnu, et al.
Publicado: (2026)
COSMOS: Predictable and Cost-Effective Adaptation of LLMs
por: Wang, Jiayu, et al.
Publicado: (2025)
por: Wang, Jiayu, et al.
Publicado: (2025)
Weak-to-Strong Generalization Through the Data-Centric Lens
por: Shin, Changho, et al.
Publicado: (2024)
por: Shin, Changho, et al.
Publicado: (2024)
Beyond Confidence: The Rhythms of Reasoning in Generative Models
por: Liu, Deyuan, et al.
Publicado: (2026)
por: Liu, Deyuan, et al.
Publicado: (2026)
Confidence Interval Estimation of Predictive Performance in the Context of AutoML
por: Paraschakis, Konstantinos, et al.
Publicado: (2024)
por: Paraschakis, Konstantinos, et al.
Publicado: (2024)
Confidence Improves Self-Consistency in LLMs
por: Taubenfeld, Amir, et al.
Publicado: (2025)
por: Taubenfeld, Amir, et al.
Publicado: (2025)
Hyperparameter Importance Analysis for Multi-Objective AutoML
por: Theodorakopoulos, Daphne, et al.
Publicado: (2024)
por: Theodorakopoulos, Daphne, et al.
Publicado: (2024)
Breaking Down Financial News Impact: A Novel AI Approach with Geometric Hypergraphs
por: Harit, Anoushka, et al.
Publicado: (2024)
por: Harit, Anoushka, et al.
Publicado: (2024)
Auto311: A Confidence-guided Automated System for Non-emergency Calls
por: Chen, Zirong, et al.
Publicado: (2023)
por: Chen, Zirong, et al.
Publicado: (2023)
Ejemplares similares
-
Promises and Pitfalls of Threshold-based Auto-labeling
por: Vishwakarma, Harit, et al.
Publicado: (2022) -
Taming False Positives in Out-of-Distribution Detection with Human Feedback
por: Vishwakarma, Harit, et al.
Publicado: (2024) -
Adaptive Scoring and Thresholding with Human Feedback for Robust Out-of-Distribution Detection
por: Yamada, Daisuke, et al.
Publicado: (2025) -
CARE: Confounder-Aware Aggregation for Reliable LLM Evaluation
por: Zhao, Jitian, et al.
Publicado: (2026) -
Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation
por: Huang, Tzu-Heng, et al.
Publicado: (2025)