What Makes Language Models Good-enough?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Asami, Daiki, Sugawara, Saku |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Dual-Task Paradigm to Investigate Sentence Comprehension Strategies in Language Models
von: Emura, Rei, et al.
Veröffentlicht: (2026)
von: Emura, Rei, et al.
Veröffentlicht: (2026)
CxMP: A Linguistic Minimal-Pair Benchmark for Evaluating Constructional Understanding in Language Models
von: Oba, Miyu, et al.
Veröffentlicht: (2026)
von: Oba, Miyu, et al.
Veröffentlicht: (2026)
C2: Scalable Rubric-Augmented Reward Modeling from Binary Preferences
von: Kawabata, Akira, et al.
Veröffentlicht: (2026)
von: Kawabata, Akira, et al.
Veröffentlicht: (2026)
Rationale-Aware Answer Verification by Pairwise Self-Evaluation
von: Kawabata, Akira, et al.
Veröffentlicht: (2024)
von: Kawabata, Akira, et al.
Veröffentlicht: (2024)
Specification-Aware Machine Translation and Evaluation for Purpose Alignment
von: Kayano, Yoko, et al.
Veröffentlicht: (2025)
von: Kayano, Yoko, et al.
Veröffentlicht: (2025)
Are Checklists Really Useful for Automatic Evaluation of Generative Tasks?
von: Furuhashi, Momoka, et al.
Veröffentlicht: (2025)
von: Furuhashi, Momoka, et al.
Veröffentlicht: (2025)
Can Language Models Induce Grammatical Knowledge from Indirect Evidence?
von: Oba, Miyu, et al.
Veröffentlicht: (2024)
von: Oba, Miyu, et al.
Veröffentlicht: (2024)
What Makes a Good Natural Language Prompt?
von: Long, Do Xuan, et al.
Veröffentlicht: (2025)
von: Long, Do Xuan, et al.
Veröffentlicht: (2025)
TactfulToM: Do LLMs Have the Theory of Mind Ability to Understand White Lies?
von: Liu, Yiwei, et al.
Veröffentlicht: (2025)
von: Liu, Yiwei, et al.
Veröffentlicht: (2025)
MoreHopQA: More Than Multi-hop Reasoning
von: Schnitzler, Julian, et al.
Veröffentlicht: (2024)
von: Schnitzler, Julian, et al.
Veröffentlicht: (2024)
Which Feedback Works for Whom? Differential Effects of LLM-Generated Feedback Elements Across Learner Profiles
von: Furuhashi, Momoka, et al.
Veröffentlicht: (2026)
von: Furuhashi, Momoka, et al.
Veröffentlicht: (2026)
Quality Text, Robust Vision: The Role of Language in Enhancing Visual Robustness of Vision-Language Models
von: Waseda, Futa, et al.
Veröffentlicht: (2025)
von: Waseda, Futa, et al.
Veröffentlicht: (2025)
What Makes Good Instruction-Tuning Data? An In-Context Learning Perspective
von: Han, Guangzeng, et al.
Veröffentlicht: (2026)
von: Han, Guangzeng, et al.
Veröffentlicht: (2026)
Label Effects: Shared Heuristic Reliance in Trust Assessment by Humans and LLM-as-a-Judge
von: Sun, Xin, et al.
Veröffentlicht: (2026)
von: Sun, Xin, et al.
Veröffentlicht: (2026)
What Makes Two Language Models Think Alike?
von: Salle, Jeanne, et al.
Veröffentlicht: (2024)
von: Salle, Jeanne, et al.
Veröffentlicht: (2024)
What Makes a Good Doctor Response? A Study on Text-Based Telemedicine
von: Cosma, Adrian, et al.
Veröffentlicht: (2026)
von: Cosma, Adrian, et al.
Veröffentlicht: (2026)
What Makes a Reward Model a Good Teacher? An Optimization Perspective
von: Razin, Noam, et al.
Veröffentlicht: (2025)
von: Razin, Noam, et al.
Veröffentlicht: (2025)
Automatic Feedback Generation for Short Answer Questions using Answer Diagnostic Graphs
von: Furuhashi, Momoka, et al.
Veröffentlicht: (2025)
von: Furuhashi, Momoka, et al.
Veröffentlicht: (2025)
Measuring Human Involvement in AI-Generated Text: A Case Study on Academic Writing
von: Guo, Yuchen, et al.
Veröffentlicht: (2025)
von: Guo, Yuchen, et al.
Veröffentlicht: (2025)
Large Language Models Know What Makes Exemplary Contexts
von: Long, Quanyu, et al.
Veröffentlicht: (2024)
von: Long, Quanyu, et al.
Veröffentlicht: (2024)
What Makes Diffusion Language Models Super Data Learners?
von: Gao, Zitian, et al.
Veröffentlicht: (2025)
von: Gao, Zitian, et al.
Veröffentlicht: (2025)
What Makes a Good Response? An Empirical Analysis of Quality in Qualitative Interviews
von: Ivey, Jonathan, et al.
Veröffentlicht: (2026)
von: Ivey, Jonathan, et al.
Veröffentlicht: (2026)
What Makes Good Multilingual Reasoning? Disentangling Reasoning Traces with Measurable Features
von: Ki, Dayeon, et al.
Veröffentlicht: (2026)
von: Ki, Dayeon, et al.
Veröffentlicht: (2026)
Sample Design Engineering: An Empirical Study of What Makes Good Downstream Fine-Tuning Samples for LLMs
von: Guo, Biyang, et al.
Veröffentlicht: (2024)
von: Guo, Biyang, et al.
Veröffentlicht: (2024)
Guiding Through Complexity: What Makes Good Supervision for Hard Math Reasoning Tasks?
von: He, Xuan, et al.
Veröffentlicht: (2024)
von: He, Xuan, et al.
Veröffentlicht: (2024)
What Makes an LLM a Good Optimizer? A Trajectory Analysis of LLM-Guided Evolutionary Search
von: Zhang, Xinhao, et al.
Veröffentlicht: (2026)
von: Zhang, Xinhao, et al.
Veröffentlicht: (2026)
What Makes Large Language Models Reason in (Multi-Turn) Code Generation?
von: Zheng, Kunhao, et al.
Veröffentlicht: (2024)
von: Zheng, Kunhao, et al.
Veröffentlicht: (2024)
Can Large Language Models Robustly Perform Natural Language Inference for Japanese Comparatives?
von: Mikami, Yosuke, et al.
Veröffentlicht: (2025)
von: Mikami, Yosuke, et al.
Veröffentlicht: (2025)
What Makes for Good Visual Instructions? Synthesizing Complex Visual Reasoning Instructions for Visual Instruction Tuning
von: Du, Yifan, et al.
Veröffentlicht: (2023)
von: Du, Yifan, et al.
Veröffentlicht: (2023)
What Makes a Good Query? Measuring the Impact of Human-Confusing Linguistic Features on LLM Performance
von: Watson, William, et al.
Veröffentlicht: (2026)
von: Watson, William, et al.
Veröffentlicht: (2026)
How Many Languages Make Good Multilingual Instruction Tuning? A Case Study on BLOOM
von: Ji, Shaoxiong, et al.
Veröffentlicht: (2024)
von: Ji, Shaoxiong, et al.
Veröffentlicht: (2024)
Are Large Language Models Good Data Preprocessors?
von: Meguellati, Elyas, et al.
Veröffentlicht: (2025)
von: Meguellati, Elyas, et al.
Veröffentlicht: (2025)
Intention Analysis Makes LLMs A Good Jailbreak Defender
von: Zhang, Yuqi, et al.
Veröffentlicht: (2024)
von: Zhang, Yuqi, et al.
Veröffentlicht: (2024)
Is analogy enough to draw novel adjective-noun inferences?
von: Ross, Hayley, et al.
Veröffentlicht: (2025)
von: Ross, Hayley, et al.
Veröffentlicht: (2025)
Are Large Language Models Good Statisticians?
von: Zhu, Yizhang, et al.
Veröffentlicht: (2024)
von: Zhu, Yizhang, et al.
Veröffentlicht: (2024)
What Makes Good Data for Alignment? A Comprehensive Study of Automatic Data Selection in Instruction Tuning
von: Liu, Wei, et al.
Veröffentlicht: (2023)
von: Liu, Wei, et al.
Veröffentlicht: (2023)
A Good Plan is Hard to Find: Aligning Models with Preferences is Misaligned with What Helps Users
von: Balepur, Nishant, et al.
Veröffentlicht: (2025)
von: Balepur, Nishant, et al.
Veröffentlicht: (2025)
Bias Vector: Mitigating Biases in Language Models with Task Arithmetic Approach
von: Shirafuji, Daiki, et al.
Veröffentlicht: (2024)
von: Shirafuji, Daiki, et al.
Veröffentlicht: (2024)
Are Large Language Models Actually Good at Text Style Transfer?
von: Mukherjee, Sourabrata, et al.
Veröffentlicht: (2024)
von: Mukherjee, Sourabrata, et al.
Veröffentlicht: (2024)
One prompt is not enough: Instruction Sensitivity Undermines Embedding Model Evaluation
von: Kostiuk, Yevhen, et al.
Veröffentlicht: (2026)
von: Kostiuk, Yevhen, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
A Dual-Task Paradigm to Investigate Sentence Comprehension Strategies in Language Models
von: Emura, Rei, et al.
Veröffentlicht: (2026) -
CxMP: A Linguistic Minimal-Pair Benchmark for Evaluating Constructional Understanding in Language Models
von: Oba, Miyu, et al.
Veröffentlicht: (2026) -
C2: Scalable Rubric-Augmented Reward Modeling from Binary Preferences
von: Kawabata, Akira, et al.
Veröffentlicht: (2026) -
Rationale-Aware Answer Verification by Pairwise Self-Evaluation
von: Kawabata, Akira, et al.
Veröffentlicht: (2024) -
Specification-Aware Machine Translation and Evaluation for Purpose Alignment
von: Kayano, Yoko, et al.
Veröffentlicht: (2025)