Benchmarking Foundation Models on Exceptional Cases: Dataset Creation and Validation
Fuente:
arXiv
Saved in:
| Main Authors: | Kang, Suho, Park, Jungyang, Ha, Joonseo, Kim, SoMin, Kim, JinHyeong, Park, Subeen, Song, Kyungwoo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Tracing Mathematical Proficiency Through Problem-Solving Processes
by: Park, Jungyang, et al.
Published: (2025)
by: Park, Jungyang, et al.
Published: (2025)
Spurious Correlation-Aware Embedding Regularization for Worst-Group Robustness
by: Park, Subeen, et al.
Published: (2025)
by: Park, Subeen, et al.
Published: (2025)
Sufficient Invariant Learning for Distribution Shift
by: Kim, Taero, et al.
Published: (2022)
by: Kim, Taero, et al.
Published: (2022)
Adaptive Task Vectors for Large Language Models
by: Kang, Joonseong, et al.
Published: (2025)
by: Kang, Joonseong, et al.
Published: (2025)
Bounded Hyperbolic Tangent: A Stable and Efficient Alternative to Pre-Layer Normalization in Large Language Models
by: Byun, Hoyoon, et al.
Published: (2025)
by: Byun, Hoyoon, et al.
Published: (2025)
MIDUS: Memory-Infused Depth Up-Scaling
by: Kim, Taero, et al.
Published: (2025)
by: Kim, Taero, et al.
Published: (2025)
MELT: Materials-aware Continued Pre-training for Language Model Adaptation to Materials Science
by: Kim, Junho, et al.
Published: (2024)
by: Kim, Junho, et al.
Published: (2024)
RoAD Benchmark: How LiDAR Models Fail under Coupled Domain Shifts and Label Evolution
by: Lee, Subeen, et al.
Published: (2026)
by: Lee, Subeen, et al.
Published: (2026)
Contrastive Language Prompting to Ease False Positives in Medical Anomaly Detection
by: Park, YeongHyeon, et al.
Published: (2024)
by: Park, YeongHyeon, et al.
Published: (2024)
Feature Attenuation of Defective Representation Can Resolve Incomplete Masking on Anomaly Detection
by: Park, YeongHyeon, et al.
Published: (2024)
by: Park, YeongHyeon, et al.
Published: (2024)
CAMEL-CLIP: Channel-aware Multimodal Electroencephalography-text Alignment for Generalizable Brain Foundation Models
by: Choi, Hanseul, et al.
Published: (2026)
by: Choi, Hanseul, et al.
Published: (2026)
RenderMem: Rendering as Spatial Memory Retrieval
by: Park, JooHyun, et al.
Published: (2026)
by: Park, JooHyun, et al.
Published: (2026)
Introducing VaDA: Novel Image Segmentation Model for Maritime Object Segmentation Using New Dataset
by: Kim, Yongjin, et al.
Published: (2024)
by: Kim, Yongjin, et al.
Published: (2024)
Multi-LLM Adaptive Conformal Inference for Reliable LLM Responses
by: Noh, Kangjun, et al.
Published: (2026)
by: Noh, Kangjun, et al.
Published: (2026)
CAdam: Context-Adaptive Moment Estimation for 3D Gaussian Densification in Generative Distillation
by: Chung, SeungJeh, et al.
Published: (2026)
by: Chung, SeungJeh, et al.
Published: (2026)
Cross-lingual Collapse: How Language-Centric Foundation Models Shape Reasoning in Large Language Models
by: Park, Cheonbok, et al.
Published: (2025)
by: Park, Cheonbok, et al.
Published: (2025)
MV-CLAM: Multi-View Molecular Interpretation with Cross-Modal Projection via Language Model
by: Ha, Sumin, et al.
Published: (2025)
by: Ha, Sumin, et al.
Published: (2025)
Social Human Robot Embodied Conversation (SHREC) Dataset: Benchmarking Foundational Models' Social Reasoning
by: Lee, Dong Won, et al.
Published: (2025)
by: Lee, Dong Won, et al.
Published: (2025)
Semi-Supervised Preference Optimization with Limited Feedback
by: Lee, Seonggyun, et al.
Published: (2025)
by: Lee, Seonggyun, et al.
Published: (2025)
GUARD-CAN: Graph-Understanding and Recurrent Architecture for CAN Anomaly Detection
by: Kim, Hyeong Seon, et al.
Published: (2025)
by: Kim, Hyeong Seon, et al.
Published: (2025)
360 in the Wild: Dataset for Depth Prediction and View Synthesis
by: Park, Kibaek, et al.
Published: (2024)
by: Park, Kibaek, et al.
Published: (2024)
ToxReason: A Benchmark for Mechanistic Chemical Toxicity Reasoning via Adverse Outcome Pathway
by: Park, Jueon, et al.
Published: (2026)
by: Park, Jueon, et al.
Published: (2026)
KITE: A Benchmark for Evaluating Korean Instruction-Following Abilities in Large Language Models
by: Kim, Dongjun, et al.
Published: (2025)
by: Kim, Dongjun, et al.
Published: (2025)
Pix2Next: Leveraging Vision Foundation Models for RGB to NIR Image Translation
by: Jin, Youngwan, et al.
Published: (2024)
by: Jin, Youngwan, et al.
Published: (2024)
SurgMLLMBench: A Multimodal Large Language Model Benchmark Dataset for Surgical Scene Understanding
by: Choi, Tae-Min, et al.
Published: (2025)
by: Choi, Tae-Min, et al.
Published: (2025)
EPLKG: Efficient Prompt Learning with Knowledge Graph
by: Lim, YongTaek, et al.
Published: (2023)
by: Lim, YongTaek, et al.
Published: (2023)
ARGOS: Who, Where, and When in Agentic Multi-Camera Person Search
by: Kim, Myungchul, et al.
Published: (2026)
by: Kim, Myungchul, et al.
Published: (2026)
Eigen-Value: Efficient Domain-Robust Data Valuation via Eigenvalue-Based Approach
by: Choi, Youngjun, et al.
Published: (2025)
by: Choi, Youngjun, et al.
Published: (2025)
InstaTrans: An Instruction-Aware Translation Framework for Non-English Instruction Datasets
by: Kim, Yungi, et al.
Published: (2024)
by: Kim, Yungi, et al.
Published: (2024)
Ko-MuSR: A Multistep Soft Reasoning Benchmark for LLMs Capable of Understanding Korean
by: Park, Chanwoo, et al.
Published: (2025)
by: Park, Chanwoo, et al.
Published: (2025)
VGenST-Bench: A Benchmark for Spatio-Temporal Reasoning via Active Video Synthesis
by: Park, Jinho, et al.
Published: (2026)
by: Park, Jinho, et al.
Published: (2026)
Explain with Visual Keypoints Like a Real Mentor! A Benchmark for Multimodal Solution Explanation
by: Park, Jaewoo, et al.
Published: (2025)
by: Park, Jaewoo, et al.
Published: (2025)
Bridging Dynamic Factor Models and Neural Controlled Differential Equations for Nowcasting GDP
by: Lim, Seonkyu, et al.
Published: (2024)
by: Lim, Seonkyu, et al.
Published: (2024)
Personalize-then-Store: Benchmarking and Learning Personalized Memory for Long-horizon Agents
by: In, Yeonjun, et al.
Published: (2026)
by: In, Yeonjun, et al.
Published: (2026)
Example-Based Concept Analysis Framework for Deep Weather Forecast Models
by: Kim, Soyeon, et al.
Published: (2025)
by: Kim, Soyeon, et al.
Published: (2025)
Text-Aware Image Restoration with Diffusion Models
by: Min, Jaewon, et al.
Published: (2025)
by: Min, Jaewon, et al.
Published: (2025)
AuditoryBench++: Can Language Models Understand Auditory Knowledge without Hearing?
by: Ok, Hyunjong, et al.
Published: (2025)
by: Ok, Hyunjong, et al.
Published: (2025)
Autoencoder-based Denoising Defense against Adversarial Attacks on Object Detection
by: Song, Min Geun, et al.
Published: (2025)
by: Song, Min Geun, et al.
Published: (2025)
Open Ko-LLM Leaderboard2: Bridging Foundational and Practical Evaluation for Korean LLMs
by: Kim, Hyeonwoo, et al.
Published: (2024)
by: Kim, Hyeonwoo, et al.
Published: (2024)
Representing the Under-Represented: Cultural and Core Capability Benchmarks for Developing Thai Large Language Models
by: Kim, Dahyun, et al.
Published: (2024)
by: Kim, Dahyun, et al.
Published: (2024)
Similar Items
-
Tracing Mathematical Proficiency Through Problem-Solving Processes
by: Park, Jungyang, et al.
Published: (2025) -
Spurious Correlation-Aware Embedding Regularization for Worst-Group Robustness
by: Park, Subeen, et al.
Published: (2025) -
Sufficient Invariant Learning for Distribution Shift
by: Kim, Taero, et al.
Published: (2022) -
Adaptive Task Vectors for Large Language Models
by: Kang, Joonseong, et al.
Published: (2025) -
Bounded Hyperbolic Tangent: A Stable and Efficient Alternative to Pre-Layer Normalization in Large Language Models
by: Byun, Hoyoon, et al.
Published: (2025)