GAICo: A Deployed and Extensible Framework for Evaluating Diverse and Multimodal Generative AI Outputs
Fuente:
arXiv
Saved in:
| Main Authors: | Gupta, Nitin, Koppisetti, Pallav, Lakkaraju, Kausik, Srivastava, Biplav |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SafeChat: A Framework for Building Trustworthy Collaborative Assistants and a Case Study of its Usefulness
by: Srivastava, Biplav, et al.
Published: (2025)
by: Srivastava, Biplav, et al.
Published: (2025)
BEACON: Balancing Convenience and Nutrition in Meals With Long-Term Group Recommendations and Reasoning on Multimodal Recipes
by: Nagpal, Vansh, et al.
Published: (2024)
by: Nagpal, Vansh, et al.
Published: (2024)
The Effect of Human v/s Synthetic Test Data and Round-tripping on Assessment of Sentiment Analysis Systems for Bias
by: Lakkaraju, Kausik, et al.
Published: (2024)
by: Lakkaraju, Kausik, et al.
Published: (2024)
PLANTS: A Novel Problem and Dataset for Summarization of Planning-Like (PL) Tasks
by: Pallagani, Vishal, et al.
Published: (2024)
by: Pallagani, Vishal, et al.
Published: (2024)
Holistic Explainable AI (H-XAI): Extending Transparency Beyond Developers in AI-Driven Decision Making
by: Lakkaraju, Kausik, et al.
Published: (2025)
by: Lakkaraju, Kausik, et al.
Published: (2025)
A Novel Approach to Balance Convenience and Nutrition in Meals With Long-Term Group Recommendations and Reasoning on Multimodal Recipes and its Implementation in BEACON
by: Nagpal, Vansh, et al.
Published: (2024)
by: Nagpal, Vansh, et al.
Published: (2024)
On Identifying Why and When Foundation Models Perform Well on Time-Series Forecasting Using Automated Explanations and Rating
by: Widener, Michael, et al.
Published: (2025)
by: Widener, Michael, et al.
Published: (2025)
A Neurosymbolic Fast and Slow Architecture for Graph Coloring
by: Khandelwal, Vedant, et al.
Published: (2024)
by: Khandelwal, Vedant, et al.
Published: (2024)
Trust and ethical considerations in a multi-modal, explainable AI-driven chatbot tutoring system: The case of collaboratively solving Rubik's Cube
by: Lakkaraju, Kausik, et al.
Published: (2024)
by: Lakkaraju, Kausik, et al.
Published: (2024)
FABLE: A Novel Data-Flow Analysis Benchmark on Procedural Text for Large Language Model Evaluation
by: Pallagani, Vishal, et al.
Published: (2025)
by: Pallagani, Vishal, et al.
Published: (2025)
Development and Evaluation of a Retrieval-Augmented Generation Tool for Creating SAPPhIRE Models of Artificial Systems
by: Majumder, Anubhab, et al.
Published: (2024)
by: Majumder, Anubhab, et al.
Published: (2024)
Assess and Prompt: A Generative RL Framework for Improving Engagement in Online Mental Health Communities
by: Gaur, Bhagesh, et al.
Published: (2025)
by: Gaur, Bhagesh, et al.
Published: (2025)
SocioEval: A Template-Based Framework for Evaluating Socioeconomic Status Bias in Foundation Models
by: Kumar, Divyanshu, et al.
Published: (2026)
by: Kumar, Divyanshu, et al.
Published: (2026)
G2: Guided Generation for Enhanced Output Diversity in LLMs
by: Ruan, Zhiwen, et al.
Published: (2025)
by: Ruan, Zhiwen, et al.
Published: (2025)
Scaling Efficient LLMs
by: Kausik, B. N.
Published: (2024)
by: Kausik, B. N.
Published: (2024)
Rating Multi-Modal Time-Series Forecasting Models (MM-TSFM) for Robustness Through a Causal Lens
by: Lakkaraju, Kausik, et al.
Published: (2024)
by: Lakkaraju, Kausik, et al.
Published: (2024)
On Sample-Efficient Generalized Planning via Learned Transition Models
by: Gupta, Nitin, et al.
Published: (2026)
by: Gupta, Nitin, et al.
Published: (2026)
ZeroSumEval: An Extensible Framework For Scaling LLM Evaluation with Inter-Model Competition
by: Alyahya, Hisham A., et al.
Published: (2025)
by: Alyahya, Hisham A., et al.
Published: (2025)
Do Voters Get the Information They Want? Understanding Authentic Voter FAQs in the US and How to Improve for Informed Electoral Participation
by: Rawte, Vipula, et al.
Published: (2024)
by: Rawte, Vipula, et al.
Published: (2024)
A Study on Effect of Reference Knowledge Choice in Generating Technical Content Relevant to SAPPhIRE Model Using Large Language Model
by: Bhattacharya, Kausik, et al.
Published: (2024)
by: Bhattacharya, Kausik, et al.
Published: (2024)
Creating a Causally Grounded Rating Method for Assessing the Robustness of AI Models for Time-Series Forecasting
by: Lakkaraju, Kausik, et al.
Published: (2025)
by: Lakkaraju, Kausik, et al.
Published: (2025)
ICLAD: In-Context Learning with Comparison-Guidance for Audio Deepfake Detection
by: Chou, Benjamin, et al.
Published: (2026)
by: Chou, Benjamin, et al.
Published: (2026)
SESGO: Spanish Evaluation of Stereotypical Generative Outputs
by: Robles, Melissa, et al.
Published: (2025)
by: Robles, Melissa, et al.
Published: (2025)
TinyScientist: An Interactive, Extensible, and Controllable Framework for Building Research Agents
by: Yu, Haofei, et al.
Published: (2025)
by: Yu, Haofei, et al.
Published: (2025)
Task-Dependent Evaluation of LLM Output Homogenization: A Taxonomy-Guided Framework
by: Jain, Shomik, et al.
Published: (2025)
by: Jain, Shomik, et al.
Published: (2025)
SO-Bench: A Structural Output Evaluation of Multimodal LLMs
by: Feng, Di, et al.
Published: (2025)
by: Feng, Di, et al.
Published: (2025)
Alethia: A Foundational Encoder for Voice Deepfakes
by: Zhu, Yi, et al.
Published: (2026)
by: Zhu, Yi, et al.
Published: (2026)
MapIQ: Evaluating Multimodal Large Language Models for Map Question Answering
by: Srivastava, Varun, et al.
Published: (2025)
by: Srivastava, Varun, et al.
Published: (2025)
EasySteer: A Unified Framework for High-Performance and Extensible LLM Steering
by: Xu, Haolei, et al.
Published: (2025)
by: Xu, Haolei, et al.
Published: (2025)
Evaluating the Evaluation of Diversity in Commonsense Generation
by: Zhang, Tianhui, et al.
Published: (2025)
by: Zhang, Tianhui, et al.
Published: (2025)
YASPS: A Symbolic Framework for Extensible, High-Performance IPC Simulation
by: Tang, Xuan, et al.
Published: (2026)
by: Tang, Xuan, et al.
Published: (2026)
A Principled Framework for Evaluating on Typologically Diverse Languages
by: Ploeger, Esther, et al.
Published: (2024)
by: Ploeger, Esther, et al.
Published: (2024)
Understanding the Effects of Iterative Prompting on Truthfulness
by: Krishna, Satyapriya, et al.
Published: (2024)
by: Krishna, Satyapriya, et al.
Published: (2024)
On the Impact of Fine-Tuning on Chain-of-Thought Reasoning
by: Lobo, Elita, et al.
Published: (2024)
by: Lobo, Elita, et al.
Published: (2024)
MapQaTor: An Extensible Framework for Efficient Annotation of Map-Based QA Datasets
by: Dihan, Mahir Labib, et al.
Published: (2024)
by: Dihan, Mahir Labib, et al.
Published: (2024)
Language of Thought Shapes Output Diversity in Large Language Models
by: Xu, Shaoyang, et al.
Published: (2026)
by: Xu, Shaoyang, et al.
Published: (2026)
A Better LLM Evaluator for Text Generation: The Impact of Prompt Output Sequencing and Optimization
by: Chu, KuanChao, et al.
Published: (2024)
by: Chu, KuanChao, et al.
Published: (2024)
A Comprehensive Analysis of Large Language Model Outputs: Similarity, Diversity, and Bias
by: Smith, Brandon, et al.
Published: (2025)
by: Smith, Brandon, et al.
Published: (2025)
Extensible Embedding: A Flexible Multipler For LLM's Context Length
by: Shao, Ninglu, et al.
Published: (2024)
by: Shao, Ninglu, et al.
Published: (2024)
STED and Consistency Scoring: A Framework for Evaluating LLM Structured Output Reliability
by: Wang, Guanghui, et al.
Published: (2025)
by: Wang, Guanghui, et al.
Published: (2025)
Similar Items
-
SafeChat: A Framework for Building Trustworthy Collaborative Assistants and a Case Study of its Usefulness
by: Srivastava, Biplav, et al.
Published: (2025) -
BEACON: Balancing Convenience and Nutrition in Meals With Long-Term Group Recommendations and Reasoning on Multimodal Recipes
by: Nagpal, Vansh, et al.
Published: (2024) -
The Effect of Human v/s Synthetic Test Data and Round-tripping on Assessment of Sentiment Analysis Systems for Bias
by: Lakkaraju, Kausik, et al.
Published: (2024) -
PLANTS: A Novel Problem and Dataset for Summarization of Planning-Like (PL) Tasks
by: Pallagani, Vishal, et al.
Published: (2024) -
Holistic Explainable AI (H-XAI): Extending Transparency Beyond Developers in AI-Driven Decision Making
by: Lakkaraju, Kausik, et al.
Published: (2025)