CARE: Confounder-Aware Aggregation for Reliable LLM Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Jitian, Shin, Changho, Huang, Tzu-Heng, GNVV, Satya Sai Srinath Namburi, Sala, Frederic |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training
by: Ge, Albert, et al.
Published: (2025)
by: Ge, Albert, et al.
Published: (2025)
OTTER: Effortless Label Distribution Adaptation of Zero-shot Models
by: Shin, Changho, et al.
Published: (2024)
by: Shin, Changho, et al.
Published: (2024)
Pretrained Hybrids with MAD Skills
by: Roberts, Nicholas, et al.
Published: (2024)
by: Roberts, Nicholas, et al.
Published: (2024)
LETS Forecast: Learning Embedology for Time Series Forecasting
by: Majeedi, Abrar, et al.
Published: (2025)
by: Majeedi, Abrar, et al.
Published: (2025)
Pearls from Pebbles: Improved Confidence Functions for Auto-labeling
by: Vishwakarma, Harit, et al.
Published: (2024)
by: Vishwakarma, Harit, et al.
Published: (2024)
Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation
by: Huang, Tzu-Heng, et al.
Published: (2025)
by: Huang, Tzu-Heng, et al.
Published: (2025)
Weak-to-Strong Generalization Through the Data-Centric Lens
by: Shin, Changho, et al.
Published: (2024)
by: Shin, Changho, et al.
Published: (2024)
Quantifying Structure in CLIP Embeddings: A Statistical Framework for Concept Interpretation
by: Zhao, Jitian, et al.
Published: (2025)
by: Zhao, Jitian, et al.
Published: (2025)
MoRe Fine-Tuning with 10x Fewer Parameters
by: Tan, Wenxuan, et al.
Published: (2024)
by: Tan, Wenxuan, et al.
Published: (2024)
Zero-Shot Robustification of Zero-Shot Models
by: Adila, Dyah, et al.
Published: (2023)
by: Adila, Dyah, et al.
Published: (2023)
Tabby: A Language Model Architecture for Tabular and Structured Data Synthesis
by: Cromp, Sonia, et al.
Published: (2025)
by: Cromp, Sonia, et al.
Published: (2025)
The ALCHEmist: Automated Labeling 500x CHEaper Than LLM Data Annotators
by: Huang, Tzu-Heng, et al.
Published: (2024)
by: Huang, Tzu-Heng, et al.
Published: (2024)
Evaluating Sample Utility for Efficient Data Selection by Mimicking Model Weights
by: Huang, Tzu-Heng, et al.
Published: (2025)
by: Huang, Tzu-Heng, et al.
Published: (2025)
Brain alignment of reasoning and action representations from vision-language and action models during naturalistic gameplay
by: Oota, Subba Reddy, et al.
Published: (2026)
by: Oota, Subba Reddy, et al.
Published: (2026)
Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain)
by: Oota, Subba Reddy, et al.
Published: (2025)
by: Oota, Subba Reddy, et al.
Published: (2025)
Linguistic properties and model scale in brain encoding: from small to compressed language models
by: Oota, Subba Reddy, et al.
Published: (2026)
by: Oota, Subba Reddy, et al.
Published: (2026)
Meaningful Causal Aggregation and Paradoxical Confounding
by: Zhu, Yuchen, et al.
Published: (2023)
by: Zhu, Yuchen, et al.
Published: (2023)
RubiCap: Rubric-Guided Reinforcement Learning for Dense Image Captioning
by: Huang, Tzu-Heng, et al.
Published: (2026)
by: Huang, Tzu-Heng, et al.
Published: (2026)
Personalize Your LLM: Fake it then Align it
by: Zhang, Yijing, et al.
Published: (2025)
by: Zhang, Yijing, et al.
Published: (2025)
Multimodal Data Curation via Object Detection and Filter Ensembles
by: Huang, Tzu-Heng, et al.
Published: (2024)
by: Huang, Tzu-Heng, et al.
Published: (2024)
Task-conditioned probing of instruction-tuned multimodal LLMs: Region-specific brain alignment patterns under naturalistic stimuli
by: Oota, Subba Reddy, et al.
Published: (2025)
by: Oota, Subba Reddy, et al.
Published: (2025)
AI-CARE: Carbon-Aware Reporting Evaluation Metric for AI Models
by: Santosh, KC, et al.
Published: (2026)
by: Santosh, KC, et al.
Published: (2026)
Evaluating Language Model Context Windows: A "Working Memory" Test and Inference-time Correction
by: Dsouza, Amanda, et al.
Published: (2024)
by: Dsouza, Amanda, et al.
Published: (2024)
ReTAMamba: Reliability-Aware Temporal Aggregation with Mamba for Irregular Clinical Time Series Prediction
by: Kim, Jinwoong, et al.
Published: (2026)
by: Kim, Jinwoong, et al.
Published: (2026)
FedUAF: Uncertainty-Aware Fusion with Reliability-Guided Aggregation for Multimodal Federated Sentiment Analysis
by: Zhu, Xianxun, et al.
Published: (2026)
by: Zhu, Xianxun, et al.
Published: (2026)
COSMOS: Predictable and Cost-Effective Adaptation of LLMs
by: Wang, Jiayu, et al.
Published: (2025)
by: Wang, Jiayu, et al.
Published: (2025)
RICA2: Rubric-Informed, Calibrated Assessment of Actions
by: Majeedi, Abrar, et al.
Published: (2024)
by: Majeedi, Abrar, et al.
Published: (2024)
Confounded Causal Imitation Learning with Instrumental Variables
by: Zeng, Yan, et al.
Published: (2025)
by: Zeng, Yan, et al.
Published: (2025)
FedCARE: Federated Unlearning with Conflict-Aware Projection and Relearning-Resistant Recovery
by: Li, Yue, et al.
Published: (2026)
by: Li, Yue, et al.
Published: (2026)
CARE-RL: Capability-Aware Reinforcement Learning for Mitigating Cross-Domain Conflicts
by: Zhang, Rui, et al.
Published: (2026)
by: Zhang, Rui, et al.
Published: (2026)
Training on the Test Task Confounds Evaluation and Emergence
by: Dominguez-Olmedo, Ricardo, et al.
Published: (2024)
by: Dominguez-Olmedo, Ricardo, et al.
Published: (2024)
Margin-Adaptive Confidence Ranking for Reliable LLM Judgement
by: Jin, Gaojie, et al.
Published: (2026)
by: Jin, Gaojie, et al.
Published: (2026)
CARE: Covariance-Aware and Rank-Enhanced Decomposition for Enabling Multi-Head Latent Attention
by: Zhou, Zhongzhu, et al.
Published: (2026)
by: Zhou, Zhongzhu, et al.
Published: (2026)
Mobility-Aware Cache Framework for Scalable LLM-Based Human Mobility Simulation
by: Yan, Hua, et al.
Published: (2026)
by: Yan, Hua, et al.
Published: (2026)
Is Free Self-Alignment Possible?
by: Adila, Dyah, et al.
Published: (2024)
by: Adila, Dyah, et al.
Published: (2024)
Relational Causal Discovery with Latent Confounders
by: Negro, Matteo, et al.
Published: (2025)
by: Negro, Matteo, et al.
Published: (2025)
Promises and Pitfalls of Threshold-based Auto-labeling
by: Vishwakarma, Harit, et al.
Published: (2022)
by: Vishwakarma, Harit, et al.
Published: (2022)
CARE-RFT: Confidence-Anchored Reinforcement Finetuning for Reliable Reasoning in Large Language Models
by: Li, Shuozhe, et al.
Published: (2026)
by: Li, Shuozhe, et al.
Published: (2026)
Dual-Signal Adaptive KV-Cache Optimization for Long-Form Video Understanding in Vision-Language Models
by: Sai, Vishnu, et al.
Published: (2026)
by: Sai, Vishnu, et al.
Published: (2026)
Towards Reliable, Uncertainty-Aware Alignment
by: Banerjee, Debangshu, et al.
Published: (2025)
by: Banerjee, Debangshu, et al.
Published: (2025)
Similar Items
-
R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training
by: Ge, Albert, et al.
Published: (2025) -
OTTER: Effortless Label Distribution Adaptation of Zero-shot Models
by: Shin, Changho, et al.
Published: (2024) -
Pretrained Hybrids with MAD Skills
by: Roberts, Nicholas, et al.
Published: (2024) -
LETS Forecast: Learning Embedology for Time Series Forecasting
by: Majeedi, Abrar, et al.
Published: (2025) -
Pearls from Pebbles: Improved Confidence Functions for Auto-labeling
by: Vishwakarma, Harit, et al.
Published: (2024)