Towards a Holistic and Automated Evaluation Framework for Multi-Level Comprehension of LLMs in Book-Length Contexts
Fuente:
arXiv
Saved in:
| Main Authors: | Deng, Jiaqi, Lee, Yuho, Kim, Nicole Hee-Yeon, Min, Hyangsuk, Yun, Taewon, Ban, Minjeong, Yul, Kim, Song, Hwanjun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Multi-dimensional Evaluation of LLM Summarization across Domains and Languages
by: Min, Hyangsuk, et al.
Published: (2025)
by: Min, Hyangsuk, et al.
Published: (2025)
Completing Missing Annotation: Multi-Agent Debate for Accurate and Scalable Relevant Assessment for IR Benchmarks
by: Ban, Minjeong, et al.
Published: (2026)
by: Ban, Minjeong, et al.
Published: (2026)
ReFeed: Multi-dimensional Summarization Refinement with Reflective Reasoning on Feedback
by: Yun, Taewon, et al.
Published: (2025)
by: Yun, Taewon, et al.
Published: (2025)
UniSumEval: Towards Unified, Fine-Grained, Multi-Dimensional Summarization Evaluation for LLMs
by: Lee, Yuho, et al.
Published: (2024)
by: Lee, Yuho, et al.
Published: (2024)
Learning to Verify Summary Facts with Fine-Grained LLM Feedback
by: Oh, Jihwan, et al.
Published: (2024)
by: Oh, Jihwan, et al.
Published: (2024)
What Makes a Sale? Rethinking End-to-End Seller--Buyer Retail Dynamics with LLM Agents
by: Choi, Jeonghwan, et al.
Published: (2026)
by: Choi, Jeonghwan, et al.
Published: (2026)
Learning to Summarize from LLM-generated Feedback
by: Song, Hwanjun, et al.
Published: (2024)
by: Song, Hwanjun, et al.
Published: (2024)
Distilling Long-CoT Reasoning through Collaborative Step-wise Multi-Teacher Decoding
by: Yun, Taewon, et al.
Published: (2026)
by: Yun, Taewon, et al.
Published: (2026)
Relationship between perceived human resource management system strength, thriving at work, and employee's turnover intention in Chinese hotels
by: Jing Guo, et al.
Published: (2024)
by: Jing Guo, et al.
Published: (2024)
Modality-Agnostic Style Transfer for Holistic Feature Imputation
by: Baek, Seunghun, et al.
Published: (2025)
by: Baek, Seunghun, et al.
Published: (2025)
How Do LLMs See Charts? A Comparative Study on High-Level Visualization Comprehension in Humans and LLMs
by: Jeon, Hyotaek, et al.
Published: (2026)
by: Jeon, Hyotaek, et al.
Published: (2026)
Associations between the tissue bacterial microbiome and keratinocyte cancer
by: Yul Hee Kim, et al.
Published: (2024)
by: Yul Hee Kim, et al.
Published: (2024)
Staphylococcus Enrichment in Cutaneous Melanoma
by: Hyoung Soo Park, et al.
Published: (2025)
by: Hyoung Soo Park, et al.
Published: (2025)
Towards Holistic Surgical Scene Graph
by: Shin, Jongmin, et al.
Published: (2025)
by: Shin, Jongmin, et al.
Published: (2025)
Aligning Extraction and Generation for Robust Retrieval-Augmented Generation
by: Song, Hwanjun, et al.
Published: (2025)
by: Song, Hwanjun, et al.
Published: (2025)
Validity and Reliability of the Korean Version of the Self‐Care Self‐Efficacy Scale for Patients With Heart Failure: A Psychometric Evaluation
by: JinShil Kim, et al.
Published: (2025)
by: JinShil Kim, et al.
Published: (2025)
Alignment Tuning for Large Language Models: A Data-Centric Lens on Alignment Data Pipelines
by: Song, Hwanjun
Published: (2026)
by: Song, Hwanjun
Published: (2026)
Towards Reliable Test-Time Adaptation: Style Invariance as a Correctness Likelihood
by: Nam, Gilhyun, et al.
Published: (2025)
by: Nam, Gilhyun, et al.
Published: (2025)
Prescription Trends of Initial Pharmacotherapy for Benign Prostatic Hyperplasia Among Treatment‐Naïve Patients in South Korea: A Retrospective Analysis
by: Yeon Hee Kim, et al.
Published: (2025)
by: Yeon Hee Kim, et al.
Published: (2025)
In-Context Learning with Noisy Labels
by: Kang, Junyong, et al.
Published: (2024)
by: Kang, Junyong, et al.
Published: (2024)
SecureMCP: A Policy-Enforced LLM Data Access Framework for AIoT Systems via Model Context Protocol
by: Kim, Wonbae, et al.
Published: (2026)
by: Kim, Wonbae, et al.
Published: (2026)
Towards a Holistic Evaluation of LLMs on Factual Knowledge Recall
by: Yuan, Jiaqing, et al.
Published: (2024)
by: Yuan, Jiaqing, et al.
Published: (2024)
Making Sense of Korean Sentences: A Comprehensive Evaluation of LLMs through KoSEnd Dataset
by: Yu, Seunguk, et al.
Published: (2025)
by: Yu, Seunguk, et al.
Published: (2025)
Efficient generation of recombinant anti‐HER2 scFv with high yield and purity using a simple method
by: Hanool Yun, et al.
Published: (2024)
by: Hanool Yun, et al.
Published: (2024)
Expanding on EnCLAP with Auxiliary Retrieval Model for Automated Audio Captioning
by: Kim, Jaeyeon, et al.
Published: (2024)
by: Kim, Jaeyeon, et al.
Published: (2024)
Designing Ethical Learning for Agentic AI: Toegye Yi Hwang's Ethical Emotion Regulation Framework
by: Kim, Ji Yeon
Published: (2026)
by: Kim, Ji Yeon
Published: (2026)
Titulador potenciométrico digital automático para experimentos de equilibrios en disolución
by: Yul Goncalves
Published: (2010)
by: Yul Goncalves
Published: (2010)
Point2Act: Efficient 3D Distillation of Multimodal LLMs for Zero-Shot Context-Aware Grasping
by: Kim, Sang Min, et al.
Published: (2025)
by: Kim, Sang Min, et al.
Published: (2025)
PPA-Plan: Proactive Pitfall Avoidance for Reliable Planning in Long-Context LLM Reasoning
by: Kim, Byeongjin, et al.
Published: (2026)
by: Kim, Byeongjin, et al.
Published: (2026)
The Physiological Response of the Fiddler Crab to Anthropogenic Low-Frequency Substrate-Borne Vibrations.
by: Joo, Soobin, et al.
Published: (2025)
by: Joo, Soobin, et al.
Published: (2025)
ViTA-PAR: Visual and Textual Attribute Alignment with Attribute Prompting for Pedestrian Attribute Recognition
by: Park, Minjeong, et al.
Published: (2025)
by: Park, Minjeong, et al.
Published: (2025)
Ageism and Voting Behaviour in a Fictitious Election Situation in Japan
by: Yuho Shimizu
Published: (2025)
by: Yuho Shimizu
Published: (2025)
Towards Fully-Automated Materials Discovery via Large-Scale Synthesis Dataset and Expert-Level LLM-as-a-Judge
by: Kim, Heegyu, et al.
Published: (2025)
by: Kim, Heegyu, et al.
Published: (2025)
Character-Centered Dialogue Generation from Scene-Level Prompts
by: Kang, Taewon, et al.
Published: (2025)
by: Kang, Taewon, et al.
Published: (2025)
Crystal structure and photoluminescence properties of CuBrxI1−x(melamine) (0 ≤ x ≤ 1) complexes
by: Juhyun Kim, et al.
Published: (2025)
by: Juhyun Kim, et al.
Published: (2025)
Holistic Evaluation of Multimodal LLMs on Spatial Intelligence
by: Cai, Zhongang, et al.
Published: (2025)
by: Cai, Zhongang, et al.
Published: (2025)
Comparison in treatment effect according to the IPL treatment area in meibomian gland dysfunction patients
by: Ji Sang Min, et al.
Published: (2024)
by: Ji Sang Min, et al.
Published: (2024)
The Impact of Perceived Tone, Age, and Gender on Voice Assistant Persuasiveness in the Context of Product Recommendations
by: Pias, Sabid Bin Habib, et al.
Published: (2024)
by: Pias, Sabid Bin Habib, et al.
Published: (2024)
Addressing selectivity challenges in seawater splitting: Catalyst design for oxygen and chlorine evolution reactions
by: Gisang Park, et al.
Published: (2025)
by: Gisang Park, et al.
Published: (2025)
Intrinsic Entropy of Context Length Scaling in LLMs
by: Shi, Jingzhe, et al.
Published: (2025)
by: Shi, Jingzhe, et al.
Published: (2025)
Similar Items
-
Towards Multi-dimensional Evaluation of LLM Summarization across Domains and Languages
by: Min, Hyangsuk, et al.
Published: (2025) -
Completing Missing Annotation: Multi-Agent Debate for Accurate and Scalable Relevant Assessment for IR Benchmarks
by: Ban, Minjeong, et al.
Published: (2026) -
ReFeed: Multi-dimensional Summarization Refinement with Reflective Reasoning on Feedback
by: Yun, Taewon, et al.
Published: (2025) -
UniSumEval: Towards Unified, Fine-Grained, Multi-Dimensional Summarization Evaluation for LLMs
by: Lee, Yuho, et al.
Published: (2024) -
Learning to Verify Summary Facts with Fine-Grained LLM Feedback
by: Oh, Jihwan, et al.
Published: (2024)