CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Roger, Alexis, Humane, Prateek, Kaplan, Daniel Z., Gupta, Kshitij, Sun, Qi, Adamopoulos, George, Lim, Jonathan Siu Chi, Anthony, Quentin, Fennell, Edwin, Rish, Irina |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Influence Functions for Efficient Data Selection in Reasoning
di: Humane, Prateek, et al.
Pubblicazione: (2025)
di: Humane, Prateek, et al.
Pubblicazione: (2025)
LLM Pretraining Shapes a Generalizable Manifold: Insights into Cross-Modal Transfer to Time Series
di: Roger, Alexis, et al.
Pubblicazione: (2026)
di: Roger, Alexis, et al.
Pubblicazione: (2026)
Random Initialization Can't Catch Up: The Advantage of Language Model Transfer for Time Series Forecasting
di: Riachi, Roland, et al.
Pubblicazione: (2025)
di: Riachi, Roland, et al.
Pubblicazione: (2025)
Towards ethical multimodal systems
di: Roger, Alexis, et al.
Pubblicazione: (2023)
di: Roger, Alexis, et al.
Pubblicazione: (2023)
Simple and Scalable Strategies to Continually Pre-train Large Language Models
di: Ibrahim, Adam, et al.
Pubblicazione: (2024)
di: Ibrahim, Adam, et al.
Pubblicazione: (2024)
Towards Adversarially Robust Vision-Language Models: Insights from Design Choices and Prompt Formatting Techniques
di: Bhagwatkar, Rishika, et al.
Pubblicazione: (2024)
di: Bhagwatkar, Rishika, et al.
Pubblicazione: (2024)
Image Tiling for High-Resolution Reasoning: Balancing Local Detail with Global Context
di: de Margerie, Anatole Jacquin, et al.
Pubblicazione: (2025)
di: de Margerie, Anatole Jacquin, et al.
Pubblicazione: (2025)
Small Vocabularies, Big Gains: Pretraining and Tokenization in Time Series Models
di: Roger, Alexis, et al.
Pubblicazione: (2025)
di: Roger, Alexis, et al.
Pubblicazione: (2025)
Amalgamated CHIRP and OFDM for ISAC
di: Kumar, Pankaj, et al.
Pubblicazione: (2026)
di: Kumar, Pankaj, et al.
Pubblicazione: (2026)
Interpretability in Action: Exploratory Analysis of VPT, a Minecraft Agent
di: Jucys, Karolis, et al.
Pubblicazione: (2024)
di: Jucys, Karolis, et al.
Pubblicazione: (2024)
Fine-Grained Unambiguous Measurements
di: Buzet, Quentin, et al.
Pubblicazione: (2025)
di: Buzet, Quentin, et al.
Pubblicazione: (2025)
WHODUNIT: Evaluation benchmark for culprit detection in mystery stories
di: Gupta, Kshitij
Pubblicazione: (2025)
di: Gupta, Kshitij
Pubblicazione: (2025)
Fine-Grained Privacy Guarantees for Coverage Problems
di: Dhulipala, Laxman, et al.
Pubblicazione: (2024)
di: Dhulipala, Laxman, et al.
Pubblicazione: (2024)
Warming Up for Zeroth-Order Federated Pre-Training with Low Resource Clients
di: Legate, Gwen, et al.
Pubblicazione: (2025)
di: Legate, Gwen, et al.
Pubblicazione: (2025)
Fine-Grained Open-Vocabulary Object Detection with Fined-Grained Prompts: Task, Dataset and Benchmark
di: Liu, Ying, et al.
Pubblicazione: (2025)
di: Liu, Ying, et al.
Pubblicazione: (2025)
OpenSIR: Open-Ended Self-Improving Reasoner
di: Kwan, Wai-Chung, et al.
Pubblicazione: (2025)
di: Kwan, Wai-Chung, et al.
Pubblicazione: (2025)
Computational Framework for White Hole Detection in Gravitational Waves and CMB Data
di: Valamontes, Antonios, et al.
Pubblicazione: (2025)
di: Valamontes, Antonios, et al.
Pubblicazione: (2025)
Symposium on Misallocation and Structural Transformation: Introduction
di: Tasso Adamopoulos, et al.
Pubblicazione: (2024)
di: Tasso Adamopoulos, et al.
Pubblicazione: (2024)
On Creativity and Open-Endedness
di: Soros, L. B., et al.
Pubblicazione: (2024)
di: Soros, L. B., et al.
Pubblicazione: (2024)
The common agricultural policy of the European community / Rosemary Fennell
di: Fennell, Rosemary
Pubblicazione: (1979)
di: Fennell, Rosemary
Pubblicazione: (1979)
Craftax: A Lightning-Fast Benchmark for Open-Ended Reinforcement Learning
di: Matthews, Michael, et al.
Pubblicazione: (2024)
di: Matthews, Michael, et al.
Pubblicazione: (2024)
PuzzleWorld: A Benchmark for Multimodal, Open-Ended Reasoning in Puzzlehunts
di: Li, Hengzhi, et al.
Pubblicazione: (2025)
di: Li, Hengzhi, et al.
Pubblicazione: (2025)
MORQA: Benchmarking Evaluation Metrics for Medical Open-Ended Question Answering
di: Yim, Wen-wai, et al.
Pubblicazione: (2025)
di: Yim, Wen-wai, et al.
Pubblicazione: (2025)
Fine-Grained Computation in 3-Space: Matrix Multiplication and Graph Problems
di: Stout, Quentin F.
Pubblicazione: (2024)
di: Stout, Quentin F.
Pubblicazione: (2024)
OpenEarthSensing: Large-Scale Fine-Grained Benchmark for Open-World Remote Sensing
di: Xiang, Xiang, et al.
Pubblicazione: (2025)
di: Xiang, Xiang, et al.
Pubblicazione: (2025)
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
di: Xie, Tianbao, et al.
Pubblicazione: (2024)
di: Xie, Tianbao, et al.
Pubblicazione: (2024)
GUIDE: A Benchmark for Understanding and Assisting Users in Open-Ended GUI Tasks
di: Yang, Saelyne, et al.
Pubblicazione: (2026)
di: Yang, Saelyne, et al.
Pubblicazione: (2026)
Assessing Bias in Metric Models for LLM Open-Ended Generation Bias Benchmarks
di: Demchak, Nathaniel, et al.
Pubblicazione: (2024)
di: Demchak, Nathaniel, et al.
Pubblicazione: (2024)
Beyond MCQ: An Open-Ended Arabic Cultural QA Benchmark with Dialect Variants
di: Bhatti, Hunzalah Hassan, et al.
Pubblicazione: (2025)
di: Bhatti, Hunzalah Hassan, et al.
Pubblicazione: (2025)
TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision
di: Gupta, Ayush, et al.
Pubblicazione: (2025)
di: Gupta, Ayush, et al.
Pubblicazione: (2025)
Knowledge Distillation for Federated Learning: a Practical Guide
di: Mora, Alessio, et al.
Pubblicazione: (2022)
di: Mora, Alessio, et al.
Pubblicazione: (2022)
On the Adversarial Robustness of Discrete Image Tokenizers
di: Bhagwatkar, Rishika, et al.
Pubblicazione: (2026)
di: Bhagwatkar, Rishika, et al.
Pubblicazione: (2026)
Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference
di: Riemer, Matthew, et al.
Pubblicazione: (2024)
di: Riemer, Matthew, et al.
Pubblicazione: (2024)
FOCUS: Bridging Fine-Grained Recognition and Open-World Discovery across Domains
di: Rathore, Vaibhav, et al.
Pubblicazione: (2026)
di: Rathore, Vaibhav, et al.
Pubblicazione: (2026)
Language, Literature, and the Negotiation of Identity
di: Fennell, Barbara A.
Pubblicazione: (2020)
di: Fennell, Barbara A.
Pubblicazione: (2020)
Chapter Introduction
di: C. Fennell, Christopher
Pubblicazione: (2026)
di: C. Fennell, Christopher
Pubblicazione: (2026)
An examination of values and environmental attitudes among ecotourists : a descriptive study involving three samples / David A. Fennell, Agnes M. K. Nowaczek
di: Fennell, David A
di: Fennell, David A
Resource Centre Assessment; A Speech Presented in Sarnia, 1978.
di: Fennell, Doris Pauline
Pubblicazione: (1978)
di: Fennell, Doris Pauline
Pubblicazione: (1978)
Staff Development: The Enduring Investment.
di: Fennell, Janice C.
Pubblicazione: (1995)
di: Fennell, Janice C.
Pubblicazione: (1995)
Quantum Walks and Exact RG in de Sitter Space
di: Green, Daniel, et al.
Pubblicazione: (2025)
di: Green, Daniel, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Influence Functions for Efficient Data Selection in Reasoning
di: Humane, Prateek, et al.
Pubblicazione: (2025) -
LLM Pretraining Shapes a Generalizable Manifold: Insights into Cross-Modal Transfer to Time Series
di: Roger, Alexis, et al.
Pubblicazione: (2026) -
Random Initialization Can't Catch Up: The Advantage of Language Model Transfer for Time Series Forecasting
di: Riachi, Roland, et al.
Pubblicazione: (2025) -
Towards ethical multimodal systems
di: Roger, Alexis, et al.
Pubblicazione: (2023) -
Simple and Scalable Strategies to Continually Pre-train Large Language Models
di: Ibrahim, Adam, et al.
Pubblicazione: (2024)