Is a Peeled Apple Still Red? Evaluating LLMs' Ability for Conceptual Combination with Property Type
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Song, Seokwon, Lee, Taehyun, Ahn, Jaewoo, Sung, Jae Hyuk, Kim, Gunhee |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TimeChara: Evaluating Point-in-Time Character Hallucination of Role-Playing Large Language Models
von: Ahn, Jaewoo, et al.
Veröffentlicht: (2024)
von: Ahn, Jaewoo, et al.
Veröffentlicht: (2024)
ChartCap: Mitigating Hallucination of Dense Chart Captioning
von: Lim, Junyoung, et al.
Veröffentlicht: (2025)
von: Lim, Junyoung, et al.
Veröffentlicht: (2025)
Who Wrote this Code? Watermarking for Code Generation
von: Lee, Taehyun, et al.
Veröffentlicht: (2023)
von: Lee, Taehyun, et al.
Veröffentlicht: (2023)
Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates
von: Ahn, Jaewoo, et al.
Veröffentlicht: (2025)
von: Ahn, Jaewoo, et al.
Veröffentlicht: (2025)
MAVIS: A Benchmark for Multimodal Source Attribution in Long-form Visual Question Answering
von: Song, Seokwon, et al.
Veröffentlicht: (2025)
von: Song, Seokwon, et al.
Veröffentlicht: (2025)
FlashAdventure: A Benchmark for GUI Agents Solving Full Story Arcs in Diverse Adventure Games
von: Ahn, Jaewoo, et al.
Veröffentlicht: (2025)
von: Ahn, Jaewoo, et al.
Veröffentlicht: (2025)
Optimal Replenishment Strategy for Satellite Constellation with Dual Supply Modes
von: Kim, Jaewoo, et al.
Veröffentlicht: (2024)
von: Kim, Jaewoo, et al.
Veröffentlicht: (2024)
GrowOVER: How Can LLMs Adapt to Growing Real-World Knowledge?
von: Ko, Dayoon, et al.
Veröffentlicht: (2024)
von: Ko, Dayoon, et al.
Veröffentlicht: (2024)
Evaluating the Pre-Consultation Ability of LLMs using Diagnostic Guidelines
von: Seo, Jean, et al.
Veröffentlicht: (2026)
von: Seo, Jean, et al.
Veröffentlicht: (2026)
Mitigating Quantization Errors Due to Activation Spikes in GLU-Based LLMs
von: Yang, Jaewoo, et al.
Veröffentlicht: (2024)
von: Yang, Jaewoo, et al.
Veröffentlicht: (2024)
Can Language Models Laugh at YouTube Short-form Videos?
von: Ko, Dayoon, et al.
Veröffentlicht: (2023)
von: Ko, Dayoon, et al.
Veröffentlicht: (2023)
Learning from Negative Samples in Biomedical Generative Entity Linking
von: Kim, Chanhwi, et al.
Veröffentlicht: (2024)
von: Kim, Chanhwi, et al.
Veröffentlicht: (2024)
Still Between Us? Evaluating and Improving Voice Assistant Robustness to Third-Party Interruptions
von: Lee, Dongwook, et al.
Veröffentlicht: (2026)
von: Lee, Dongwook, et al.
Veröffentlicht: (2026)
ETHIC: Evaluating Large Language Models on Long-Context Tasks with High Information Coverage
von: Lee, Taewhoo, et al.
Veröffentlicht: (2024)
von: Lee, Taewhoo, et al.
Veröffentlicht: (2024)
How Language Directions Align with Token Geometry in Multilingual LLMs
von: Kim, JaeSeong, et al.
Veröffentlicht: (2025)
von: Kim, JaeSeong, et al.
Veröffentlicht: (2025)
Scaling Personality Control in LLMs with Big Five Scaler Prompts
von: Cho, Gunhee, et al.
Veröffentlicht: (2025)
von: Cho, Gunhee, et al.
Veröffentlicht: (2025)
Extracting and Combining Abilities For Building Multi-lingual Ability-enhanced Large Language Models
von: Chen, Zhipeng, et al.
Veröffentlicht: (2024)
von: Chen, Zhipeng, et al.
Veröffentlicht: (2024)
Language Models Don't Learn the Physical Manifestation of Language
von: Lee, Bruce W., et al.
Veröffentlicht: (2024)
von: Lee, Bruce W., et al.
Veröffentlicht: (2024)
Joint Replenishment Strategy for Multiple Satellite Constellations with Shared Launch Opportunities
von: Kim, Jaewoo, et al.
Veröffentlicht: (2025)
von: Kim, Jaewoo, et al.
Veröffentlicht: (2025)
On-Orbit Servicing-Integrated Maintenance Strategy for Satellite Constellation
von: Kim, Jaewoo, et al.
Veröffentlicht: (2025)
von: Kim, Jaewoo, et al.
Veröffentlicht: (2025)
Culinary Class Wars: Evaluating LLMs using ASH in Cuisine Transfer Task
von: Lee, Hoonick, et al.
Veröffentlicht: (2024)
von: Lee, Hoonick, et al.
Veröffentlicht: (2024)
DynamicER: Resolving Emerging Mentions to Dynamic Entities for RAG
von: Kim, Jinyoung, et al.
Veröffentlicht: (2024)
von: Kim, Jinyoung, et al.
Veröffentlicht: (2024)
Evaluating Large Language Models on the 2026 Korean CSAT Mathematics Exam: Measuring Mathematical Ability in a Zero-Data-Leakage Setting
von: Pyeon, Goun, et al.
Veröffentlicht: (2025)
von: Pyeon, Goun, et al.
Veröffentlicht: (2025)
Still No Lie Detector for Language Models: Probing Empirical and Conceptual Roadblocks
von: Levinstein, B. A., et al.
Veröffentlicht: (2023)
von: Levinstein, B. A., et al.
Veröffentlicht: (2023)
MEME: Multi-entity & Evolving Memory Evaluation
von: Jung, Seokwon, et al.
Veröffentlicht: (2026)
von: Jung, Seokwon, et al.
Veröffentlicht: (2026)
When Should Dense Retrievers Be Updated in Evolving Corpora? Detecting Out-of-Distribution Corpora Using GradNormIR
von: Ko, Dayoon, et al.
Veröffentlicht: (2025)
von: Ko, Dayoon, et al.
Veröffentlicht: (2025)
Evaluating the Ability of LLMs to Solve Semantics-Aware Process Mining Tasks
von: Rebmann, Adrian, et al.
Veröffentlicht: (2024)
von: Rebmann, Adrian, et al.
Veröffentlicht: (2024)
SIMPLOT: Enhancing Chart Question Answering by Distilling Essentials
von: Kim, Wonjoong, et al.
Veröffentlicht: (2024)
von: Kim, Wonjoong, et al.
Veröffentlicht: (2024)
Help Me Write a Story: Evaluating LLMs' Ability to Generate Writing Feedback
von: Rashkin, Hannah, et al.
Veröffentlicht: (2025)
von: Rashkin, Hannah, et al.
Veröffentlicht: (2025)
Think, Verbalize, then Speak: Bridging Complex Thoughts and Comprehensible Speech
von: Woo, Sang Hoon, et al.
Veröffentlicht: (2025)
von: Woo, Sang Hoon, et al.
Veröffentlicht: (2025)
KOMBO: Korean Character Representations Based on the Combination Rules of Subcharacters
von: Kim, SungHo, et al.
Veröffentlicht: (2026)
von: Kim, SungHo, et al.
Veröffentlicht: (2026)
What's Not Said Still Hurts: A Description-Based Evaluation Framework for Measuring Social Bias in LLMs
von: Pan, Jinhao, et al.
Veröffentlicht: (2025)
von: Pan, Jinhao, et al.
Veröffentlicht: (2025)
Is the Top Still Spinning? Evaluating Subjectivity in Narrative Understanding
von: Subbiah, Melanie, et al.
Veröffentlicht: (2025)
von: Subbiah, Melanie, et al.
Veröffentlicht: (2025)
MARS: Benchmarking the Metaphysical Reasoning Abilities of Language Models with a Multi-task Evaluation Dataset
von: Wang, Weiqi, et al.
Veröffentlicht: (2024)
von: Wang, Weiqi, et al.
Veröffentlicht: (2024)
AetherCode: Evaluating LLMs' Ability to Win In Premier Programming Competitions
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
Smoothie-Qwen: Post-Hoc Smoothing to Reduce Language Bias in Multilingual LLMs
von: Ji, SeungWon, et al.
Veröffentlicht: (2025)
von: Ji, SeungWon, et al.
Veröffentlicht: (2025)
Teaching Language Models to Think in Code
von: Hwang, Hyeon, et al.
Veröffentlicht: (2026)
von: Hwang, Hyeon, et al.
Veröffentlicht: (2026)
Neuro-RIT: Neuron-Guided Instruction Tuning for Robust Retrieval-Augmented Language Model
von: Kim, Jaemin, et al.
Veröffentlicht: (2026)
von: Kim, Jaemin, et al.
Veröffentlicht: (2026)
Detecting Conceptual Abstraction in LLMs
von: Regneri, Michaela, et al.
Veröffentlicht: (2024)
von: Regneri, Michaela, et al.
Veröffentlicht: (2024)
The Curious Case of Analogies: Investigating Analogical Reasoning in Large Language Models
von: Lee, Taewhoo, et al.
Veröffentlicht: (2025)
von: Lee, Taewhoo, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
TimeChara: Evaluating Point-in-Time Character Hallucination of Role-Playing Large Language Models
von: Ahn, Jaewoo, et al.
Veröffentlicht: (2024) -
ChartCap: Mitigating Hallucination of Dense Chart Captioning
von: Lim, Junyoung, et al.
Veröffentlicht: (2025) -
Who Wrote this Code? Watermarking for Code Generation
von: Lee, Taehyun, et al.
Veröffentlicht: (2023) -
Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates
von: Ahn, Jaewoo, et al.
Veröffentlicht: (2025) -
MAVIS: A Benchmark for Multimodal Source Attribution in Long-form Visual Question Answering
von: Song, Seokwon, et al.
Veröffentlicht: (2025)