Pearls from Pebbles: Improved Confidence Functions for Auto-labeling
Fuente:
arXiv
Salvato in:
| Autori principali: | Vishwakarma, Harit, Reid, Chen, Tay, Sui Jiet, Namburi, Satya Sai Srinath, Sala, Frederic, Vinayak, Ramya Korlakai |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Promises and Pitfalls of Threshold-based Auto-labeling
di: Vishwakarma, Harit, et al.
Pubblicazione: (2022)
di: Vishwakarma, Harit, et al.
Pubblicazione: (2022)
Taming False Positives in Out-of-Distribution Detection with Human Feedback
di: Vishwakarma, Harit, et al.
Pubblicazione: (2024)
di: Vishwakarma, Harit, et al.
Pubblicazione: (2024)
Adaptive Scoring and Thresholding with Human Feedback for Robust Out-of-Distribution Detection
di: Yamada, Daisuke, et al.
Pubblicazione: (2025)
di: Yamada, Daisuke, et al.
Pubblicazione: (2025)
CARE: Confounder-Aware Aggregation for Reliable LLM Evaluation
di: Zhao, Jitian, et al.
Pubblicazione: (2026)
di: Zhao, Jitian, et al.
Pubblicazione: (2026)
Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation
di: Huang, Tzu-Heng, et al.
Pubblicazione: (2025)
di: Huang, Tzu-Heng, et al.
Pubblicazione: (2025)
GRASP: Graph Agentic Search over Propositions for Multi-hop Question Answering
di: Jenkins, Stockton, et al.
Pubblicazione: (2026)
di: Jenkins, Stockton, et al.
Pubblicazione: (2026)
OTTER: Effortless Label Distribution Adaptation of Zero-shot Models
di: Shin, Changho, et al.
Pubblicazione: (2024)
di: Shin, Changho, et al.
Pubblicazione: (2024)
Pretrained Hybrids with MAD Skills
di: Roberts, Nicholas, et al.
Pubblicazione: (2024)
di: Roberts, Nicholas, et al.
Pubblicazione: (2024)
Metric Learning in an RKHS
di: Tatli, Gokcan, et al.
Pubblicazione: (2025)
di: Tatli, Gokcan, et al.
Pubblicazione: (2025)
LETS Forecast: Learning Embedology for Time Series Forecasting
di: Majeedi, Abrar, et al.
Pubblicazione: (2025)
di: Majeedi, Abrar, et al.
Pubblicazione: (2025)
Learning from Less: Measuring the Effectiveness of RLVR in Low Data and Compute Regimes
di: Bauer, Justin, et al.
Pubblicazione: (2026)
di: Bauer, Justin, et al.
Pubblicazione: (2026)
Is Conformal Factuality for RAG-based LLMs Robust? Novel Metrics and Systematic Insights
di: Chen, Yi, et al.
Pubblicazione: (2026)
di: Chen, Yi, et al.
Pubblicazione: (2026)
R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training
di: Ge, Albert, et al.
Pubblicazione: (2025)
di: Ge, Albert, et al.
Pubblicazione: (2025)
Brain alignment of reasoning and action representations from vision-language and action models during naturalistic gameplay
di: Oota, Subba Reddy, et al.
Pubblicazione: (2026)
di: Oota, Subba Reddy, et al.
Pubblicazione: (2026)
Linguistic properties and model scale in brain encoding: from small to compressed language models
di: Oota, Subba Reddy, et al.
Pubblicazione: (2026)
di: Oota, Subba Reddy, et al.
Pubblicazione: (2026)
Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain)
di: Oota, Subba Reddy, et al.
Pubblicazione: (2025)
di: Oota, Subba Reddy, et al.
Pubblicazione: (2025)
Task-conditioned probing of instruction-tuned multimodal LLMs: Region-specific brain alignment patterns under naturalistic stimuli
di: Oota, Subba Reddy, et al.
Pubblicazione: (2025)
di: Oota, Subba Reddy, et al.
Pubblicazione: (2025)
Multimodal Data Curation via Object Detection and Filter Ensembles
di: Huang, Tzu-Heng, et al.
Pubblicazione: (2024)
di: Huang, Tzu-Heng, et al.
Pubblicazione: (2024)
Prune 'n Predict: Optimizing LLM Decision-making with Conformal Prediction
di: Vishwakarma, Harit, et al.
Pubblicazione: (2024)
di: Vishwakarma, Harit, et al.
Pubblicazione: (2024)
Bridging Lifelong and Multi-Task Representation Learning via Algorithm and Complexity Measure
di: Wang, Zhi, et al.
Pubblicazione: (2025)
di: Wang, Zhi, et al.
Pubblicazione: (2025)
Metric Learning from Limited Pairwise Preference Comparisons
di: Wang, Zhi, et al.
Pubblicazione: (2024)
di: Wang, Zhi, et al.
Pubblicazione: (2024)
Tabby: A Language Model Architecture for Tabular and Structured Data Synthesis
di: Cromp, Sonia, et al.
Pubblicazione: (2025)
di: Cromp, Sonia, et al.
Pubblicazione: (2025)
Causal Spherical Hypergraph Networks for Modelling Social Uncertainty
di: Harit, Anoushka, et al.
Pubblicazione: (2025)
di: Harit, Anoushka, et al.
Pubblicazione: (2025)
PAL: Pluralistic Alignment Framework for Learning from Heterogeneous Preferences
di: Chen, Daiwei, et al.
Pubblicazione: (2024)
di: Chen, Daiwei, et al.
Pubblicazione: (2024)
RicciFlowRec: A Geometric Root Cause Recommender Using Ricci Curvature on Financial Graphs
di: Sun, Zhongtian, et al.
Pubblicazione: (2025)
di: Sun, Zhongtian, et al.
Pubblicazione: (2025)
Actionable Interpretability via Causal Hypergraphs: Unravelling Batch Size Effects in Deep Learning
di: Sun, Zhongtian, et al.
Pubblicazione: (2025)
di: Sun, Zhongtian, et al.
Pubblicazione: (2025)
RICA2: Rubric-Informed, Calibrated Assessment of Actions
di: Majeedi, Abrar, et al.
Pubblicazione: (2024)
di: Majeedi, Abrar, et al.
Pubblicazione: (2024)
Maximizing Confidence Alone Improves Reasoning
di: Prabhudesai, Mihir, et al.
Pubblicazione: (2025)
di: Prabhudesai, Mihir, et al.
Pubblicazione: (2025)
Auto FAQ Generation
di: Kalvakolanu, Anjaneya Teja, et al.
Pubblicazione: (2024)
di: Kalvakolanu, Anjaneya Teja, et al.
Pubblicazione: (2024)
Grounded Token Initialization for New Vocabulary in LMs for Generative Recommendation
di: Chen, Daiwei, et al.
Pubblicazione: (2026)
di: Chen, Daiwei, et al.
Pubblicazione: (2026)
ManifoldMind: Dynamic Hyperbolic Reasoning for Trustworthy Recommendations
di: Harit, Anoushka, et al.
Pubblicazione: (2025)
di: Harit, Anoushka, et al.
Pubblicazione: (2025)
Dual-Signal Adaptive KV-Cache Optimization for Long-Form Video Understanding in Vision-Language Models
di: Sai, Vishnu, et al.
Pubblicazione: (2026)
di: Sai, Vishnu, et al.
Pubblicazione: (2026)
COSMOS: Predictable and Cost-Effective Adaptation of LLMs
di: Wang, Jiayu, et al.
Pubblicazione: (2025)
di: Wang, Jiayu, et al.
Pubblicazione: (2025)
Weak-to-Strong Generalization Through the Data-Centric Lens
di: Shin, Changho, et al.
Pubblicazione: (2024)
di: Shin, Changho, et al.
Pubblicazione: (2024)
Beyond Confidence: The Rhythms of Reasoning in Generative Models
di: Liu, Deyuan, et al.
Pubblicazione: (2026)
di: Liu, Deyuan, et al.
Pubblicazione: (2026)
Confidence Interval Estimation of Predictive Performance in the Context of AutoML
di: Paraschakis, Konstantinos, et al.
Pubblicazione: (2024)
di: Paraschakis, Konstantinos, et al.
Pubblicazione: (2024)
Confidence Improves Self-Consistency in LLMs
di: Taubenfeld, Amir, et al.
Pubblicazione: (2025)
di: Taubenfeld, Amir, et al.
Pubblicazione: (2025)
Hyperparameter Importance Analysis for Multi-Objective AutoML
di: Theodorakopoulos, Daphne, et al.
Pubblicazione: (2024)
di: Theodorakopoulos, Daphne, et al.
Pubblicazione: (2024)
Breaking Down Financial News Impact: A Novel AI Approach with Geometric Hypergraphs
di: Harit, Anoushka, et al.
Pubblicazione: (2024)
di: Harit, Anoushka, et al.
Pubblicazione: (2024)
Auto311: A Confidence-guided Automated System for Non-emergency Calls
di: Chen, Zirong, et al.
Pubblicazione: (2023)
di: Chen, Zirong, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Promises and Pitfalls of Threshold-based Auto-labeling
di: Vishwakarma, Harit, et al.
Pubblicazione: (2022) -
Taming False Positives in Out-of-Distribution Detection with Human Feedback
di: Vishwakarma, Harit, et al.
Pubblicazione: (2024) -
Adaptive Scoring and Thresholding with Human Feedback for Robust Out-of-Distribution Detection
di: Yamada, Daisuke, et al.
Pubblicazione: (2025) -
CARE: Confounder-Aware Aggregation for Reliable LLM Evaluation
di: Zhao, Jitian, et al.
Pubblicazione: (2026) -
Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation
di: Huang, Tzu-Heng, et al.
Pubblicazione: (2025)