The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Seungone, Suk, Juyoung, Cho, Ji Yong, Longpre, Shayne, Kim, Chaeeun, Yoon, Dongkeun, Son, Guijin, Cho, Yejin, Shafayat, Sheikh, Baek, Jinheon, Park, Sue Hyun, Hwang, Hyeonbin, Jo, Jinkyung, Cho, Hyowon, Shin, Haebin, Lee, Seongyun, Oh, Hanseok, Lee, Noah, Ho, Namgyu, Joo, Se June, Ko, Miyoung, Lee, Yoonjoo, Chae, Hyungjoo, Shin, Jamin, Jang, Joel, Ye, Seonghyeon, Lin, Bill Yuchen, Welleck, Sean, Neubig, Graham, Lee, Moontae, Lee, Kyungjae, Seo, Minjoon |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
KTRL+F: Knowledge-Augmented In-Document Search
by: Oh, Hanseok, et al.
Published: (2023)
by: Oh, Hanseok, et al.
Published: (2023)
Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
by: Kim, Seungone, et al.
Published: (2024)
by: Kim, Seungone, et al.
Published: (2024)
The CoT Encyclopedia: Analyzing, Predicting, and Controlling how a Reasoning Model will Think
by: Lee, Seongyun, et al.
Published: (2025)
by: Lee, Seongyun, et al.
Published: (2025)
Lost in the Noise: How Reasoning Models Fail with Contextual Distractors
by: Lee, Seongyun, et al.
Published: (2026)
by: Lee, Seongyun, et al.
Published: (2026)
Aligning to Thousands of Preferences via System Message Generalization
by: Lee, Seongyun, et al.
Published: (2024)
by: Lee, Seongyun, et al.
Published: (2024)
INSTRUCTIR: A Benchmark for Instruction Following of Information Retrieval Models
by: Oh, Hanseok, et al.
Published: (2024)
by: Oh, Hanseok, et al.
Published: (2024)
Learning to Explore and Select for Coverage-Conditioned Retrieval-Augmented Generation
by: Kim, Takyoung, et al.
Published: (2024)
by: Kim, Takyoung, et al.
Published: (2024)
Prometheus: Inducing Fine-grained Evaluation Capability in Language Models
by: Kim, Seungone, et al.
Published: (2023)
by: Kim, Seungone, et al.
Published: (2023)
Prometheus-Vision: Vision-Language Model as a Judge for Fine-Grained Evaluation
by: Lee, Seongyun, et al.
Published: (2024)
by: Lee, Seongyun, et al.
Published: (2024)
Mind the Blind Spots: A Focus-Level Evaluation Framework for LLM Reviews
by: Shin, Hyungyu, et al.
Published: (2025)
by: Shin, Hyungyu, et al.
Published: (2025)
Mitigating Resolution-Drift in Federated Learning: Case of Keypoint Detection
by: Lim, Taeheon, et al.
Published: (2025)
by: Lim, Taeheon, et al.
Published: (2025)
Partial-Multivariate Model for Forecasting
by: Lee, Jaehoon, et al.
Published: (2024)
by: Lee, Jaehoon, et al.
Published: (2024)
LangBridge: Multilingual Reasoning Without Multilingual Supervision
by: Yoon, Dongkeun, et al.
Published: (2024)
by: Yoon, Dongkeun, et al.
Published: (2024)
PANORAMA: A Dataset and Benchmarks Capturing Decision Trails and Rationales in Patent Examination
by: Lim, Hyunseung, et al.
Published: (2025)
by: Lim, Hyunseung, et al.
Published: (2025)
Characterizing Pattern Matching and Its Limits on Compositional Task Structures
by: Chang, Hoyeon, et al.
Published: (2025)
by: Chang, Hoyeon, et al.
Published: (2025)
LG AI Research & KAIST at EHRSQL 2024: Self-Training Large Language Models with Pseudo-Labeled Unanswerable Questions for a Reliable Text-to-SQL System on EHRs
by: Jo, Yongrae, et al.
Published: (2024)
by: Jo, Yongrae, et al.
Published: (2024)
Reinforcement Learning from Reflective Feedback (RLRF): Aligning and Improving LLMs via Fine-Grained Self-Reflection
by: Lee, Kyungjae, et al.
Published: (2024)
by: Lee, Kyungjae, et al.
Published: (2024)
Towards Robust and Parameter-Efficient Knowledge Unlearning for LLMs
by: Cha, Sungmin, et al.
Published: (2024)
by: Cha, Sungmin, et al.
Published: (2024)
Efficient Long Context Language Model Retrieval with Compression
by: Seo, Minju, et al.
Published: (2024)
by: Seo, Minju, et al.
Published: (2024)
Paper2Code: Automating Code Generation from Scientific Papers in Machine Learning
by: Seo, Minju, et al.
Published: (2025)
by: Seo, Minju, et al.
Published: (2025)
Knowledge Entropy Decay during Language Model Pretraining Hinders New Knowledge Acquisition
by: Kim, Jiyeon, et al.
Published: (2024)
by: Kim, Jiyeon, et al.
Published: (2024)
K-BrowseComp: A Web Browsing Agent Benchmark Grounded in Korean Contexts
by: Lee, Nahyun, et al.
Published: (2026)
by: Lee, Nahyun, et al.
Published: (2026)
Differential Information Distribution: A Bayesian Perspective on Direct Preference Optimization
by: Won, Yunjae, et al.
Published: (2025)
by: Won, Yunjae, et al.
Published: (2025)
Synchronisation-Oriented Design Approach for Adaptive Control
by: Cho, Namhoon, et al.
Published: (2024)
by: Cho, Namhoon, et al.
Published: (2024)
Prediction of Highway Traffic Flow Based on Artificial Intelligence Algorithms Using California Traffic Data
by: Lee, Junseong, et al.
Published: (2025)
by: Lee, Junseong, et al.
Published: (2025)
Foundation Model-Driven Framework for Human-Object Interaction Prediction with Segmentation Mask Integration
by: Park, Juhan, et al.
Published: (2025)
by: Park, Juhan, et al.
Published: (2025)
Practical and Reproducible Symbolic Music Generation by Large Language Models with Structural Embeddings
by: Rhyu, Seungyeon, et al.
Published: (2024)
by: Rhyu, Seungyeon, et al.
Published: (2024)
Learning to Unlearn: Instance-wise Unlearning for Pre-trained Classifiers
by: Cha, Sungmin, et al.
Published: (2023)
by: Cha, Sungmin, et al.
Published: (2023)
Volcano: Mitigating Multimodal Hallucination through Self-Feedback Guided Revision
by: Lee, Seongyun, et al.
Published: (2023)
by: Lee, Seongyun, et al.
Published: (2023)
Bellman Unbiasedness: Toward Provably Efficient Distributional Reinforcement Learning with General Value Function Approximation
by: Cho, Taehyun, et al.
Published: (2024)
by: Cho, Taehyun, et al.
Published: (2024)
Policy-labeled Preference Learning: Is Preference Enough for RLHF?
by: Cho, Taehyun, et al.
Published: (2025)
by: Cho, Taehyun, et al.
Published: (2025)
One Missing Piece for Open-Source Reasoning Models: A Dataset to Mitigate Cold-Starting Short CoT LLMs in RL
by: Chae, Hyungjoo, et al.
Published: (2025)
by: Chae, Hyungjoo, et al.
Published: (2025)
Autonomous Bayesian Optimization‐Based Control System for Droplet Generation
by: Seongsu Cho, et al.
Published: (2025)
by: Seongsu Cho, et al.
Published: (2025)
Video Diffusion Models are Strong Video Inpainter
by: Lee, Minhyeok, et al.
Published: (2024)
by: Lee, Minhyeok, et al.
Published: (2024)
G2L:From Giga-Scale to Cancer-Specific Large-Scale Pathology Foundation Models via Knowledge Distillation
by: Cho, Yesung, et al.
Published: (2025)
by: Cho, Yesung, et al.
Published: (2025)
Widening Regional Economic Disparities and Fertility Decline in South Korea: A Cross‐Level Analysis
by: Kyungjae Lee, et al.
Published: (2026)
by: Kyungjae Lee, et al.
Published: (2026)
Evidence-Focused Fact Summarization for Knowledge-Augmented Zero-Shot Question Answering
by: Ko, Sungho, et al.
Published: (2024)
by: Ko, Sungho, et al.
Published: (2024)
BEOL‐Compatible Liquid‐Metal‐Printing of Ultrathin 2D Oxide Memtransistors and Its Applications in Neuromorphic Computing
by: Sanghyun Moon, et al.
Published: (2026)
by: Sanghyun Moon, et al.
Published: (2026)
High‐Efficiency Quantum Dot Permeable Electrode Light‐Emitting Triodes for Visible Light Communications and on‐Device Data Encryption (Adv. Mater. 38/2025)
by: Seungmin Shin, et al.
Published: (2025)
by: Seungmin Shin, et al.
Published: (2025)
High‐Efficiency Quantum Dot Permeable Electrode Light‐Emitting Triodes for Visible Light Communications and on‐Device Data Encryption
by: Seungmin Shin, et al.
Published: (2025)
by: Seungmin Shin, et al.
Published: (2025)
Similar Items
-
KTRL+F: Knowledge-Augmented In-Document Search
by: Oh, Hanseok, et al.
Published: (2023) -
Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
by: Kim, Seungone, et al.
Published: (2024) -
The CoT Encyclopedia: Analyzing, Predicting, and Controlling how a Reasoning Model will Think
by: Lee, Seongyun, et al.
Published: (2025) -
Lost in the Noise: How Reasoning Models Fail with Contextual Distractors
by: Lee, Seongyun, et al.
Published: (2026) -
Aligning to Thousands of Preferences via System Message Generalization
by: Lee, Seongyun, et al.
Published: (2024)