Revisiting Compositional Generalization Capability of Large Language Models Considering Instruction Following Ability
Fuente:
arXiv
Saved in:
| Main Authors: | Sakai, Yusuke, Kamigaito, Hidetaka, Watanabe, Taro |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Toward the Evaluation of Large Language Models Considering Score Variance across Instruction Templates
by: Sakai, Yusuke, et al.
Published: (2024)
by: Sakai, Yusuke, et al.
Published: (2024)
mCSQA: Multilingual Commonsense Reasoning Dataset with Unified Creation Strategy by Language Models and Humans
by: Sakai, Yusuke, et al.
Published: (2024)
by: Sakai, Yusuke, et al.
Published: (2024)
Multilinguality of Large Language Models From a Structural Perspective
by: Sakajo, Haruki, et al.
Published: (2026)
by: Sakajo, Haruki, et al.
Published: (2026)
Tonguescape: Exploring Language Models Understanding of Vowel Articulation
by: Sakajo, Haruki, et al.
Published: (2025)
by: Sakajo, Haruki, et al.
Published: (2025)
HalluCitation Matters: Revealing the Impact of Hallucinated References with 300 Hallucinated Papers in ACL Conferences
by: Sakai, Yusuke, et al.
Published: (2026)
by: Sakai, Yusuke, et al.
Published: (2026)
HalluCiteChecker: A Lightweight Toolkit for Hallucinated Citation Detection and Verification in the Era of AI Scientists
by: Sakai, Yusuke, et al.
Published: (2026)
by: Sakai, Yusuke, et al.
Published: (2026)
Simultaneous Interpretation Corpus Construction by Large Language Models in Distant Language Pair
by: Sakai, Yusuke, et al.
Published: (2024)
by: Sakai, Yusuke, et al.
Published: (2024)
Does Pre-trained Language Model Actually Infer Unseen Links in Knowledge Graph Completion?
by: Sakai, Yusuke, et al.
Published: (2023)
by: Sakai, Yusuke, et al.
Published: (2023)
StructLens: A Structural Lens for Language Models via Maximum Spanning Trees
by: Sakajo, Haruki, et al.
Published: (2026)
by: Sakajo, Haruki, et al.
Published: (2026)
Dictionaries to the Rescue: Cross-Lingual Vocabulary Transfer for Low-Resource Languages Using Bilingual Dictionaries
by: Sakajo, Haruki, et al.
Published: (2025)
by: Sakajo, Haruki, et al.
Published: (2025)
Agreement-Constrained Probabilistic Minimum Bayes Risk Decoding
by: Natsumi, Koki, et al.
Published: (2025)
by: Natsumi, Koki, et al.
Published: (2025)
Efficient Nearest Neighbor based Uncertainty Estimation for Natural Language Processing Tasks
by: Hashimoto, Wataru, et al.
Published: (2024)
by: Hashimoto, Wataru, et al.
Published: (2024)
How to Make the Most of LLMs' Grammatical Knowledge for Acceptability Judgments
by: Ide, Yusuke, et al.
Published: (2024)
by: Ide, Yusuke, et al.
Published: (2024)
Dependency-Aware Semi-Structured Sparsity of GLU Variants in Large Language Models
by: Guo, Zhiyu, et al.
Published: (2024)
by: Guo, Zhiyu, et al.
Published: (2024)
Toward Automatic Safe Driving Instruction: A Large-Scale Vision Language Model Approach
by: Sakajo, Haruki, et al.
Published: (2025)
by: Sakajo, Haruki, et al.
Published: (2025)
Are Data Augmentation Methods in Named Entity Recognition Applicable for Uncertainty Estimation?
by: Hashimoto, Wataru, et al.
Published: (2024)
by: Hashimoto, Wataru, et al.
Published: (2024)
Model-based Subsampling for Knowledge Graph Completion
by: Feng, Xincan, et al.
Published: (2023)
by: Feng, Xincan, et al.
Published: (2023)
InstructCMP: Length Control in Sentence Compression through Instruction-based Large Language Models
by: Juseon-Do, et al.
Published: (2024)
by: Juseon-Do, et al.
Published: (2024)
IRR: Image Review Ranking Framework for Evaluating Vision-Language Models
by: Hayashi, Kazuki, et al.
Published: (2024)
by: Hayashi, Kazuki, et al.
Published: (2024)
BQA: Body Language Question Answering Dataset for Video Large Language Models
by: Ozaki, Shintaro, et al.
Published: (2024)
by: Ozaki, Shintaro, et al.
Published: (2024)
mbrs: A Library for Minimum Bayes Risk Decoding
by: Deguchi, Hiroyuki, et al.
Published: (2024)
by: Deguchi, Hiroyuki, et al.
Published: (2024)
Towards Cross-Lingual Explanation of Artwork in Large-scale Vision Language Models
by: Ozaki, Shintaro, et al.
Published: (2024)
by: Ozaki, Shintaro, et al.
Published: (2024)
Diagnosing Vision Language Models' Perception by Leveraging Human Methods for Color Vision Deficiencies
by: Hayashi, Kazuki, et al.
Published: (2025)
by: Hayashi, Kazuki, et al.
Published: (2025)
Decoding Uncertainty: The Impact of Decoding Strategies for Uncertainty Estimation in Large Language Models
by: Hashimoto, Wataru, et al.
Published: (2025)
by: Hashimoto, Wataru, et al.
Published: (2025)
IMPARA-GED: Grammatical Error Detection is Boosting Reference-free Grammatical Error Quality Estimator
by: Sakai, Yusuke, et al.
Published: (2025)
by: Sakai, Yusuke, et al.
Published: (2025)
InFoBench: Evaluating Instruction Following Ability in Large Language Models
by: Qin, Yiwei, et al.
Published: (2024)
by: Qin, Yiwei, et al.
Published: (2024)
Considering Length Diversity in Retrieval-Augmented Summarization
by: Juseon-Do, et al.
Published: (2025)
by: Juseon-Do, et al.
Published: (2025)
Diversity Explains Inference Scaling Laws: Through a Case Study of Minimum Bayes Risk Decoding
by: Kamigaito, Hidetaka, et al.
Published: (2024)
by: Kamigaito, Hidetaka, et al.
Published: (2024)
KITE: A Benchmark for Evaluating Korean Instruction-Following Abilities in Large Language Models
by: Kim, Dongjun, et al.
Published: (2025)
by: Kim, Dongjun, et al.
Published: (2025)
Diversity of Transformer Layers: One Aspect of Parameter Scaling Laws
by: Kamigaito, Hidetaka, et al.
Published: (2025)
by: Kamigaito, Hidetaka, et al.
Published: (2025)
Conifer: Improving Complex Constrained Instruction-Following Ability of Large Language Models
by: Sun, Haoran, et al.
Published: (2024)
by: Sun, Haoran, et al.
Published: (2024)
Revisiting the Reliability of Language Models in Instruction-Following
by: Dong, Jianshuo, et al.
Published: (2025)
by: Dong, Jianshuo, et al.
Published: (2025)
AdParaphrase: Paraphrase Dataset for Analyzing Linguistic Features toward Generating Attractive Ad Texts
by: Murakami, Soichiro, et al.
Published: (2025)
by: Murakami, Soichiro, et al.
Published: (2025)
Oogiri-Master: Benchmarking Humor Understanding via Oogiri
by: Murakami, Soichiro, et al.
Published: (2025)
by: Murakami, Soichiro, et al.
Published: (2025)
Routing by Analogy: kNN-Augmented Expert Assignment for Mixture-of-Experts
by: Lyu, Boxuan, et al.
Published: (2026)
by: Lyu, Boxuan, et al.
Published: (2026)
Who Laughs with Whom? Disentangling Influential Factors in Humor Preferences across User Clusters and LLMs
by: Murakami, Soichiro, et al.
Published: (2026)
by: Murakami, Soichiro, et al.
Published: (2026)
Unveiling the Power of Source: Source-based Minimum Bayes Risk Decoding for Neural Machine Translation
by: Lyu, Boxuan, et al.
Published: (2024)
by: Lyu, Boxuan, et al.
Published: (2024)
Revisiting Non-Verbatim Memorization in Large Language Models: The Role of Entity Surface Forms
by: Nishida, Yuto, et al.
Published: (2026)
by: Nishida, Yuto, et al.
Published: (2026)
Centroid-Based Efficient Minimum Bayes Risk Decoding
by: Deguchi, Hiroyuki, et al.
Published: (2024)
by: Deguchi, Hiroyuki, et al.
Published: (2024)
Constructing Multilingual Visual-Text Datasets Revealing Visual Multilingual Ability of Vision Language Models
by: Atuhurra, Jesse, et al.
Published: (2024)
by: Atuhurra, Jesse, et al.
Published: (2024)
Similar Items
-
Toward the Evaluation of Large Language Models Considering Score Variance across Instruction Templates
by: Sakai, Yusuke, et al.
Published: (2024) -
mCSQA: Multilingual Commonsense Reasoning Dataset with Unified Creation Strategy by Language Models and Humans
by: Sakai, Yusuke, et al.
Published: (2024) -
Multilinguality of Large Language Models From a Structural Perspective
by: Sakajo, Haruki, et al.
Published: (2026) -
Tonguescape: Exploring Language Models Understanding of Vowel Articulation
by: Sakajo, Haruki, et al.
Published: (2025) -
HalluCitation Matters: Revealing the Impact of Hallucinated References with 300 Hallucinated Papers in ACL Conferences
by: Sakai, Yusuke, et al.
Published: (2026)