FESTA: Functionally Equivalent Sampling for Trust Assessment of Multimodal LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Bhattacharya, Debarpan, Kulkarni, Apoorva, Ganapathy, Sriram |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Benchmarking and Confidence Evaluation of LALMs For Temporal Reasoning
by: Bhattacharya, Debarpan, et al.
Published: (2025)
by: Bhattacharya, Debarpan, et al.
Published: (2025)
Gradient-free Post-hoc Explainability Using Distillation Aided Learnable Approach
by: Bhattacharya, Debarpan, et al.
Published: (2024)
by: Bhattacharya, Debarpan, et al.
Published: (2024)
Towards the Next Frontier in Speech Representation Learning Using Disentanglement
by: Krishna, Varun, et al.
Published: (2024)
by: Krishna, Varun, et al.
Published: (2024)
LLM Augmented LLMs: Expanding Capabilities through Composition
by: Bansal, Rachit, et al.
Published: (2024)
by: Bansal, Rachit, et al.
Published: (2024)
PromptMind Team at EHRSQL-2024: Improving Reliability of SQL Generation using Ensemble LLMs
by: Gundabathula, Satya K, et al.
Published: (2024)
by: Gundabathula, Satya K, et al.
Published: (2024)
Benchmarking Multi-Agent LLM Architectures for Financial Document Processing: A Comparative Study of Orchestration Patterns, Cost-Accuracy Tradeoffs and Production Scaling Strategies
by: Kulkarni, Siddhant, et al.
Published: (2026)
by: Kulkarni, Siddhant, et al.
Published: (2026)
Sample-Efficient Alignment for LLMs
by: Liu, Zichen, et al.
Published: (2024)
by: Liu, Zichen, et al.
Published: (2024)
An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
by: Hao, Yuren, et al.
Published: (2025)
by: Hao, Yuren, et al.
Published: (2025)
Improving Self-supervised Pre-training using Accent-Specific Codebooks
by: Prabhu, Darshan, et al.
Published: (2024)
by: Prabhu, Darshan, et al.
Published: (2024)
Sample then Identify: A General Framework for Risk Control and Assessment in Multimodal Large Language Models
by: Wang, Qingni, et al.
Published: (2024)
by: Wang, Qingni, et al.
Published: (2024)
Multimodal Fusion with LLMs for Engagement Prediction in Natural Conversation
by: Ma, Cheng Charles, et al.
Published: (2024)
by: Ma, Cheng Charles, et al.
Published: (2024)
A Surprising Failure? Multimodal LLMs and the NLVR Challenge
by: Wu, Anne, et al.
Published: (2024)
by: Wu, Anne, et al.
Published: (2024)
On the Emergence of Thinking in LLMs I: Searching for the Right Intuition
by: Ye, Guanghao, et al.
Published: (2025)
by: Ye, Guanghao, et al.
Published: (2025)
Rethinking GSPO: The Perplexity-Entropy Equivalence
by: Liu, Chi
Published: (2025)
by: Liu, Chi
Published: (2025)
Wings: Learning Multimodal LLMs without Text-only Forgetting
by: Zhang, Yi-Kai, et al.
Published: (2024)
by: Zhang, Yi-Kai, et al.
Published: (2024)
Towards Robust Evaluation of Unlearning in LLMs via Data Transformations
by: Joshi, Abhinav, et al.
Published: (2024)
by: Joshi, Abhinav, et al.
Published: (2024)
Beyond Confidence: Rethinking Self-Assessments for Performance Prediction in LLMs
by: Bhattacharyya, Sree, et al.
Published: (2026)
by: Bhattacharyya, Sree, et al.
Published: (2026)
PromptMind Team at MEDIQA-CORR 2024: Improving Clinical Text Correction with Error Categorization and LLM Ensembles
by: Gundabathula, Satya Kesav, et al.
Published: (2024)
by: Gundabathula, Satya Kesav, et al.
Published: (2024)
RCT Rejection Sampling for Causal Estimation Evaluation
by: Keith, Katherine A., et al.
Published: (2023)
by: Keith, Katherine A., et al.
Published: (2023)
UnStar: Unlearning with Self-Taught Anti-Sample Reasoning for LLMs
by: Sinha, Yash, et al.
Published: (2024)
by: Sinha, Yash, et al.
Published: (2024)
Understanding Multimodal LLMs Under Distribution Shifts: An Information-Theoretic Approach
by: Oh, Changdae, et al.
Published: (2025)
by: Oh, Changdae, et al.
Published: (2025)
Modality Collapse as Mismatched Decoding: Information-Theoretic Limits of Multimodal LLMs
by: Billa, Jayadev
Published: (2026)
by: Billa, Jayadev
Published: (2026)
Equivalent Linear Mappings of Large Language Models
by: Golden, James R.
Published: (2025)
by: Golden, James R.
Published: (2025)
Do LLMs Encode Functional Importance of Reasoning Tokens?
by: Singh, Janvijay, et al.
Published: (2026)
by: Singh, Janvijay, et al.
Published: (2026)
STEM: Efficient Relative Capability Evaluation of LLMs through Structured Transition Samples
by: Hu, Haiquan, et al.
Published: (2025)
by: Hu, Haiquan, et al.
Published: (2025)
MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples
by: Xie, Shuo, et al.
Published: (2024)
by: Xie, Shuo, et al.
Published: (2024)
Reparameterized LLM Training via Orthogonal Equivalence Transformation
by: Qiu, Zeju, et al.
Published: (2025)
by: Qiu, Zeju, et al.
Published: (2025)
CorrSynth -- A Correlated Sampling Method for Diverse Dataset Generation from LLMs
by: Kowshik, Suhas S, et al.
Published: (2024)
by: Kowshik, Suhas S, et al.
Published: (2024)
Robust Checkpoint Selection for Multimodal LLMs via Agentic Evaluation and Stability-Aware Ranking
by: Xu, Qinwu, et al.
Published: (2026)
by: Xu, Qinwu, et al.
Published: (2026)
Contrastive Learning to Improve Retrieval for Real-world Fact Checking
by: Sriram, Aniruddh, et al.
Published: (2024)
by: Sriram, Aniruddh, et al.
Published: (2024)
PARAMANU-GANITA: Can Small Math Language Models Rival with Large Language Models on Mathematical Reasoning?
by: Niyogi, Mitodru, et al.
Published: (2024)
by: Niyogi, Mitodru, et al.
Published: (2024)
BhashaSetu: Cross-Lingual Knowledge Transfer from High-Resource to Extreme Low-Resource Languages
by: Maji, Subhadip, et al.
Published: (2026)
by: Maji, Subhadip, et al.
Published: (2026)
What can Large Language Models Capture about Code Functional Equivalence?
by: Maveli, Nickil, et al.
Published: (2024)
by: Maveli, Nickil, et al.
Published: (2024)
Tool Preferences in Agentic LLMs are Unreliable
by: Faghih, Kazem, et al.
Published: (2025)
by: Faghih, Kazem, et al.
Published: (2025)
Unified Tool Integration for LLMs: A Protocol-Agnostic Approach to Function Calling
by: Ding, Peng, et al.
Published: (2025)
by: Ding, Peng, et al.
Published: (2025)
Concurrency without Model Changes: Future-based Asynchronous Function Calling for LLMs
by: Feng, Guangyu, et al.
Published: (2026)
by: Feng, Guangyu, et al.
Published: (2026)
Less is More for Improving Automatic Evaluation of Factual Consistency
by: Wang, Tong, et al.
Published: (2024)
by: Wang, Tong, et al.
Published: (2024)
TrustLDM: Benchmarking Trustworthiness in Language Diffusion Models
by: Mo, Yichuan, et al.
Published: (2026)
by: Mo, Yichuan, et al.
Published: (2026)
Ayn: A Tiny yet Competitive Indian Legal Language Model Pretrained from Scratch
by: Niyogi, Mitodru, et al.
Published: (2024)
by: Niyogi, Mitodru, et al.
Published: (2024)
Don't Trust: Verify -- Grounding LLM Quantitative Reasoning with Autoformalization
by: Zhou, Jin Peng, et al.
Published: (2024)
by: Zhou, Jin Peng, et al.
Published: (2024)
Similar Items
-
Benchmarking and Confidence Evaluation of LALMs For Temporal Reasoning
by: Bhattacharya, Debarpan, et al.
Published: (2025) -
Gradient-free Post-hoc Explainability Using Distillation Aided Learnable Approach
by: Bhattacharya, Debarpan, et al.
Published: (2024) -
Towards the Next Frontier in Speech Representation Learning Using Disentanglement
by: Krishna, Varun, et al.
Published: (2024) -
LLM Augmented LLMs: Expanding Capabilities through Composition
by: Bansal, Rachit, et al.
Published: (2024) -
PromptMind Team at EHRSQL-2024: Improving Reliability of SQL Generation using Ensemble LLMs
by: Gundabathula, Satya K, et al.
Published: (2024)