LLM Confidence Evaluation Measures in Zero-Shot CSS Classification
Fuente:
arXiv
Saved in:
| Main Authors: | Farr, David, Cruickshank, Iain, Manzonelli, Nico, Clark, Nicholas, Starbird, Kate, West, Jevin |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLM Chain Ensembles for Scalable and Accurate Data Annotation
by: Farr, David, et al.
Published: (2024)
by: Farr, David, et al.
Published: (2024)
RED-CT: A Systems Design Methodology for Using LLM-labeled Data to Train and Deploy Edge Classifiers for Computational Social Science
by: Farr, David, et al.
Published: (2024)
by: Farr, David, et al.
Published: (2024)
The Cost of Consensus: Malignant Epistemic Herding and Adaptive Gating in Distributed Multi-Agent Search
by: Farr, David, et al.
Published: (2026)
by: Farr, David, et al.
Published: (2026)
Reliable Annotations with Less Effort: Evaluating LLM-Human Collaboration in Search Clarifications
by: Tavakoli, Leila, et al.
Published: (2025)
by: Tavakoli, Leila, et al.
Published: (2025)
BloomIntent: Automating Search Evaluation with LLM-Generated Fine-Grained User Intents
by: Choi, Yoonseo, et al.
Published: (2025)
by: Choi, Yoonseo, et al.
Published: (2025)
Seek and You Shall Find: Design & Evaluation of a Context-Aware Interactive Search Companion
by: Bink, Markus, et al.
Published: (2026)
by: Bink, Markus, et al.
Published: (2026)
The Fault in Our Recommendations: On the Perils of Optimizing the Measurable
by: Besbes, Omar, et al.
Published: (2024)
by: Besbes, Omar, et al.
Published: (2024)
Comparing Traditional and LLM-based Search for Image Geolocation
by: Wazzan, Albatool, et al.
Published: (2024)
by: Wazzan, Albatool, et al.
Published: (2024)
Leveraging Multimodal LLM for Inspirational User Interface Search
by: Park, Seokhyeon, et al.
Published: (2025)
by: Park, Seokhyeon, et al.
Published: (2025)
OwlerLite: Scope- and Freshness-Aware Web Retrieval for LLM Assistants
by: Zerhoudi, Saber, et al.
Published: (2026)
by: Zerhoudi, Saber, et al.
Published: (2026)
Towards a Signal Detection Based Measure for Assessing Information Quality of Explainable Recommender Systems
by: Son, Yeonbin, et al.
Published: (2025)
by: Son, Yeonbin, et al.
Published: (2025)
SymbioticRAG: Enhancing Document Intelligence Through Human-LLM Symbiotic Collaboration
by: Sun, Qiang, et al.
Published: (2025)
by: Sun, Qiang, et al.
Published: (2025)
DataScout: Automatic Data Fact Retrieval for Statement Augmentation with an LLM-Based Agent
by: Chen, Chuer, et al.
Published: (2025)
by: Chen, Chuer, et al.
Published: (2025)
ChartifyText: Automated Chart Generation from Data-Involved Texts via LLM
by: Zhang, Songheng, et al.
Published: (2024)
by: Zhang, Songheng, et al.
Published: (2024)
Echoes in the Loop: Diagnosing Risks in LLM-Powered Recommender Systems under Feedback Loops
by: Park, Donguk, et al.
Published: (2026)
by: Park, Donguk, et al.
Published: (2026)
Orbit: A Framework for Designing and Evaluating Multi-objective Rankers
by: Yang, Chenyang, et al.
Published: (2024)
by: Yang, Chenyang, et al.
Published: (2024)
Learning Outcomes, Assessment, and Evaluation in Educational Recommender Systems: A Systematic Review
by: Askarbekuly, Nursultan, et al.
Published: (2024)
by: Askarbekuly, Nursultan, et al.
Published: (2024)
Evolving Paradigms in Task-Based Search and Learning: A Comparative Analysis of Traditional Search Engine with LLM-Enhanced Conversational Search System
by: Guan, Zhitong, et al.
Published: (2025)
by: Guan, Zhitong, et al.
Published: (2025)
From Tool to Teacher: Rethinking Search Systems as Instructive Interfaces
by: Elsweiler, David
Published: (2026)
by: Elsweiler, David
Published: (2026)
Display Content, Display Methods and Evaluation Methods of the HCI in Explainable Recommender Systems: A Survey
by: Li, Weiqing, et al.
Published: (2025)
by: Li, Weiqing, et al.
Published: (2025)
Blending Queries and Conversations: Understanding Tactics, Trust, Verification, and System Choice in Web Search and Chat Interactions
by: Mayerhofer, Kerstin, et al.
Published: (2025)
by: Mayerhofer, Kerstin, et al.
Published: (2025)
What's in People's Digital File Collections?
by: Dinneen, Jesse David, et al.
Published: (2024)
by: Dinneen, Jesse David, et al.
Published: (2024)
FairEval: Evaluating Fairness in LLM-Based Recommendations with Personality Awareness
by: Sah, Chandan Kumar, et al.
Published: (2025)
by: Sah, Chandan Kumar, et al.
Published: (2025)
"Can You Tell Me?": Designing Copilots to Support Human Judgement in Online Information Seeking
by: Bink, Markus, et al.
Published: (2026)
by: Bink, Markus, et al.
Published: (2026)
Efficiency and Effectiveness of LLM-Based Summarization of Evidence in Crowdsourced Fact-Checking
by: Roitero, Kevin, et al.
Published: (2025)
by: Roitero, Kevin, et al.
Published: (2025)
Verify as You Go: An LLM-Powered Browser Extension for Fake News Detection
by: Sallami, Dorsaf, et al.
Published: (2026)
by: Sallami, Dorsaf, et al.
Published: (2026)
iSee: Advancing Multi-Shot Explainable AI Using Case-based Recommendations
by: Wijekoon, Anjana, et al.
Published: (2024)
by: Wijekoon, Anjana, et al.
Published: (2024)
Tip of the Tongue Query Elicitation for Simulated Evaluation
by: He, Yifan, et al.
Published: (2025)
by: He, Yifan, et al.
Published: (2025)
Balancing Domestic and Global Perspectives: Evaluating Dual-Calibration and LLM-Generated Nudges for Diverse News Recommendation
by: Sun, Ruixuan, et al.
Published: (2026)
by: Sun, Ruixuan, et al.
Published: (2026)
Argumentative Experience: Reducing Confirmation Bias on Controversial Issues through LLM-Generated Multi-Persona Debates
by: Shi, Li, et al.
Published: (2024)
by: Shi, Li, et al.
Published: (2024)
Designing and Evaluating an Educational Recommender System with Different Levels of User Control
by: Ain, Qurat Ul, et al.
Published: (2025)
by: Ain, Qurat Ul, et al.
Published: (2025)
Context Does Matter: Implications for Crowdsourced Evaluation Labels in Task-Oriented Dialogue Systems
by: Siro, Clemencia, et al.
Published: (2024)
by: Siro, Clemencia, et al.
Published: (2024)
Dagstuhl Perspectives Workshop 24352 -- Conversational Agents: A Framework for Evaluation (CAFE): Manifesto
by: Bauer, Christine, et al.
Published: (2025)
by: Bauer, Christine, et al.
Published: (2025)
Colour Contrast on the Web: A WCAG 2.1 Level AA Compliance Audit of Common Crawl's Top 500 Domains
by: Vaughan, Thom, et al.
Published: (2026)
by: Vaughan, Thom, et al.
Published: (2026)
RecGaze: The First Eye Tracking and User Interaction Dataset for Carousel Interfaces
by: de Leon-Martinez, Santiago, et al.
Published: (2025)
by: de Leon-Martinez, Santiago, et al.
Published: (2025)
How to Make Museums More Interactive? Case Study of Artistic Chatbot
by: Kucia, Filip J., et al.
Published: (2025)
by: Kucia, Filip J., et al.
Published: (2025)
From SERPs to Agents: A Platform for Comparative Studies of Information Interaction
by: Zerhoudi, Saber, et al.
Published: (2026)
by: Zerhoudi, Saber, et al.
Published: (2026)
Task Supportive and Personalized Human-Large Language Model Interaction: A User Study
by: Wang, Ben, et al.
Published: (2024)
by: Wang, Ben, et al.
Published: (2024)
Enhancing EmoBot: An In-Depth Analysis of User Satisfaction and Faults in an Emotion-Aware Chatbot
by: Mubassira, Taseen, et al.
Published: (2024)
by: Mubassira, Taseen, et al.
Published: (2024)
Beyond Centralization: User-Controlled Federated Recommendations in Practice
by: Slokom, Manel, et al.
Published: (2026)
by: Slokom, Manel, et al.
Published: (2026)
Similar Items
-
LLM Chain Ensembles for Scalable and Accurate Data Annotation
by: Farr, David, et al.
Published: (2024) -
RED-CT: A Systems Design Methodology for Using LLM-labeled Data to Train and Deploy Edge Classifiers for Computational Social Science
by: Farr, David, et al.
Published: (2024) -
The Cost of Consensus: Malignant Epistemic Herding and Adaptive Gating in Distributed Multi-Agent Search
by: Farr, David, et al.
Published: (2026) -
Reliable Annotations with Less Effort: Evaluating LLM-Human Collaboration in Search Clarifications
by: Tavakoli, Leila, et al.
Published: (2025) -
BloomIntent: Automating Search Evaluation with LLM-Generated Fine-Grained User Intents
by: Choi, Yoonseo, et al.
Published: (2025)