Self-rationalization improves LLM as a fine-grained judge
Fuente:
arXiv
Saved in:
| Main Authors: | Trivedi, Prapti, Gulati, Aditya, Molenschot, Oliver, Rajeev, Meghana Arakkal, Ramamurthy, Rajkumar, Stevens, Keith, Chaudhery, Tanveesh Singh, Jambholkar, Jahnavi, Zou, James, Rajani, Nazneen |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VERITAS: A Unified Approach to Reliability Evaluation
by: Ramamurthy, Rajkumar, et al.
Published: (2024)
by: Ramamurthy, Rajkumar, et al.
Published: (2024)
Cats Confuse Reasoning LLM: Query Agnostic Adversarial Triggers for Reasoning Models
by: Rajeev, Meghana, et al.
Published: (2025)
by: Rajeev, Meghana, et al.
Published: (2025)
Impatient Users Confuse AI Agents: High-fidelity Simulations of Human Traits for Testing Agents
by: He, Muyu, et al.
Published: (2025)
by: He, Muyu, et al.
Published: (2025)
Why do we Trust Chatbots? From Normative Principles to Behavioral Drivers
by: Gulati, Aditya, et al.
Published: (2026)
by: Gulati, Aditya, et al.
Published: (2026)
Stage-wise Fine-tuning for Graph-to-Text Generation
by: Wang, Qingyun, et al.
Published: (2021)
by: Wang, Qingyun, et al.
Published: (2021)
The Valley of Code Reasoning: Scaling Knowledge Distillation of Large Language Models
by: He, Muyu, et al.
Published: (2025)
by: He, Muyu, et al.
Published: (2025)
Lookism: The overlooked bias in computer vision
by: Gulati, Aditya, et al.
Published: (2024)
by: Gulati, Aditya, et al.
Published: (2024)
Reverse
by: Rahman, Prapti
Published: (2024)
by: Rahman, Prapti
Published: (2024)
bayesics: Core Statistical Methods via Bayesian Inference in R
by: Sewell, Daniel K., et al.
Published: (2026)
by: Sewell, Daniel K., et al.
Published: (2026)
ReviewRobot: Explainable Paper Review Generation based on Knowledge Synthesis
by: Wang, Qingyun, et al.
Published: (2020)
by: Wang, Qingyun, et al.
Published: (2020)
What's documented in AI? Systematic Analysis of 32K AI Model Cards
by: Liang, Weixin, et al.
Published: (2024)
by: Liang, Weixin, et al.
Published: (2024)
$\texttt{YC-Bench}$: Benchmarking AI Agents for Long-Term Planning and Consistent Execution
by: He, Muyu, et al.
Published: (2026)
by: He, Muyu, et al.
Published: (2026)
When Algorithms Play Favorites: Lookism in the Generation and Perception of Faces
by: Doh, Miriam, et al.
Published: (2025)
by: Doh, Miriam, et al.
Published: (2025)
Aesthetics as Structural Harm: Algorithmic Lookism Across Text-to-Image Generation and Classification
by: Doh, Miriam, et al.
Published: (2026)
by: Doh, Miriam, et al.
Published: (2026)
Happy Young Women, Grumpy Old Men? Emotion-Driven Demographic Biases in Synthetic Face Generation
by: Wei, Mengting, et al.
Published: (2026)
by: Wei, Mengting, et al.
Published: (2026)
Media representation of women migrant workers: a critical look
by: Nazneen Ahmed (Author)
Published: (2023)
by: Nazneen Ahmed (Author)
Published: (2023)
Marine fisheries development in India, 2000 AD: a perspective
by: Ramamurthy, S.
Published: (1988)
by: Ramamurthy, S.
Published: (1988)
Emotion Recognition from the perspective of Activity Recognition
by: Nagendra, Savinay, et al.
Published: (2024)
by: Nagendra, Savinay, et al.
Published: (2024)
An Analytical Study of Fear of Missing Out (Fomo) and Its Impact on Online Purchase Behaviour
by: siri, Kilarapu Jahnavi
Published: (2025)
by: siri, Kilarapu Jahnavi
Published: (2025)
Composing Copyless Streaming String Transducers
by: Alur, Rajeev, et al.
Published: (2022)
by: Alur, Rajeev, et al.
Published: (2022)
Ram-Maha/Visual-Tasks-as-Language-agnostic-window-into-early-reading-heterogeneity-: Release for Current Biology submission
by: Maha Ramamurthy
Published: (2026)
by: Maha Ramamurthy
Published: (2026)
Thermal‐Spin Conversion: Mechanism, Materials, and Determinants of the Spin Seebeck Effect
by: Meghana Mishra, et al.
Published: (2025)
by: Meghana Mishra, et al.
Published: (2025)
Mind the Style: Impact of Communication Style on Human-Chatbot Interaction
by: Derner, Erik, et al.
Published: (2026)
by: Derner, Erik, et al.
Published: (2026)
Normalized Space Alignment: A Versatile Metric for Representation Analysis
by: Ebadulla, Danish, et al.
Published: (2024)
by: Ebadulla, Danish, et al.
Published: (2024)
Vigilance Beyondthe Scalpel: The Post-op Story
by: Rajani Nawal,
Published: (2025)
by: Rajani Nawal,
Published: (2025)
Abnormal origin of posterior circumflex humeral artery and subscapular artery: case report and review of the literature
by: Rajani Singh
Published: (2017)
by: Rajani Singh
Published: (2017)
Relating Suicide: A Personal and Critical Perspective. By A.Whitehead, Bloomsbury Publishing, 2023. 128 pp. £14.99 (paperback); £45.00 (hardback); £13.49 (e‐book). ISBN: 135019218X, 9781350192188
by: Laila Rajani
Published: (2025)
by: Laila Rajani
Published: (2025)
Anomalies of radial and ulnar arteries
by: Rajani Singh
Published: (2017)
by: Rajani Singh
Published: (2017)
Code Review Automation Via Multi-task Federated LLM -- An Empirical Study
by: Kumar, Jahnavi, et al.
Published: (2024)
by: Kumar, Jahnavi, et al.
Published: (2024)
Anterior Commissure Position Relative to the Thyroid Cartilage for Safe Chondrolaryngoplasty
by: Jahnavi, et al.
Published: (2025)
by: Jahnavi, et al.
Published: (2025)
Idiosyncrasies Unveiled: Examining the Pace, Patterns and Predictors of Biotic Diversification in Peninsular India
by: Pragyadeep Roy, et al.
Published: (2025)
by: Pragyadeep Roy, et al.
Published: (2025)
PRISM: Prompt Reliability via Iterative Simulation and Monitoring for Enterprise Conversational AI
by: Chaitanya, Keshava, et al.
Published: (2026)
by: Chaitanya, Keshava, et al.
Published: (2026)
NordFKB: a fine-grained benchmark dataset for geospatial AI in Norway
by: Jyhne, Sander Riisøen, et al.
Published: (2025)
by: Jyhne, Sander Riisøen, et al.
Published: (2025)
Beauty and the Bias: Exploring the Impact of Attractiveness on Multimodal Large Language Models
by: Gulati, Aditya, et al.
Published: (2025)
by: Gulati, Aditya, et al.
Published: (2025)
Multidimensional clustering in judge designs
by: Ligtenberg, Johannes W., et al.
Published: (2024)
by: Ligtenberg, Johannes W., et al.
Published: (2024)
Review on Probiotics as A Health Supplement
by: Subhajit Patra, Priyanka Ray, Prapti Chakraborty*
Published: (2025)
by: Subhajit Patra, Priyanka Ray, Prapti Chakraborty*
Published: (2025)
Enhancing Electrocardiogram Signal Analysis Using NLP-Inspired Techniques: A Novel Approach with Embedding and Self-Attention
by: Ganguly, Prapti, et al.
Published: (2024)
by: Ganguly, Prapti, et al.
Published: (2024)
Seed Kernel Counting using Domain Randomization and Object Tracking Neural Networks
by: Margapuri, Venkat, et al.
Published: (2023)
by: Margapuri, Venkat, et al.
Published: (2023)
Leaf Angle Estimation using Mask R-CNN and LETR Vision Transformer
by: Margapuri, Venkat, et al.
Published: (2024)
by: Margapuri, Venkat, et al.
Published: (2024)
Cryptography in the Common Haar State Model: Feasibility Results and Separations
by: Ananth, Prabhanjan, et al.
Published: (2024)
by: Ananth, Prabhanjan, et al.
Published: (2024)
Similar Items
-
VERITAS: A Unified Approach to Reliability Evaluation
by: Ramamurthy, Rajkumar, et al.
Published: (2024) -
Cats Confuse Reasoning LLM: Query Agnostic Adversarial Triggers for Reasoning Models
by: Rajeev, Meghana, et al.
Published: (2025) -
Impatient Users Confuse AI Agents: High-fidelity Simulations of Human Traits for Testing Agents
by: He, Muyu, et al.
Published: (2025) -
Why do we Trust Chatbots? From Normative Principles to Behavioral Drivers
by: Gulati, Aditya, et al.
Published: (2026) -
Stage-wise Fine-tuning for Graph-to-Text Generation
by: Wang, Qingyun, et al.
Published: (2021)