What Makes Language Models Good-enough?
Fuente:
arXiv
Salvato in:
| Autori principali: | Asami, Daiki, Sugawara, Saku |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A Dual-Task Paradigm to Investigate Sentence Comprehension Strategies in Language Models
di: Emura, Rei, et al.
Pubblicazione: (2026)
di: Emura, Rei, et al.
Pubblicazione: (2026)
CxMP: A Linguistic Minimal-Pair Benchmark for Evaluating Constructional Understanding in Language Models
di: Oba, Miyu, et al.
Pubblicazione: (2026)
di: Oba, Miyu, et al.
Pubblicazione: (2026)
C2: Scalable Rubric-Augmented Reward Modeling from Binary Preferences
di: Kawabata, Akira, et al.
Pubblicazione: (2026)
di: Kawabata, Akira, et al.
Pubblicazione: (2026)
Rationale-Aware Answer Verification by Pairwise Self-Evaluation
di: Kawabata, Akira, et al.
Pubblicazione: (2024)
di: Kawabata, Akira, et al.
Pubblicazione: (2024)
Specification-Aware Machine Translation and Evaluation for Purpose Alignment
di: Kayano, Yoko, et al.
Pubblicazione: (2025)
di: Kayano, Yoko, et al.
Pubblicazione: (2025)
Are Checklists Really Useful for Automatic Evaluation of Generative Tasks?
di: Furuhashi, Momoka, et al.
Pubblicazione: (2025)
di: Furuhashi, Momoka, et al.
Pubblicazione: (2025)
Can Language Models Induce Grammatical Knowledge from Indirect Evidence?
di: Oba, Miyu, et al.
Pubblicazione: (2024)
di: Oba, Miyu, et al.
Pubblicazione: (2024)
What Makes a Good Natural Language Prompt?
di: Long, Do Xuan, et al.
Pubblicazione: (2025)
di: Long, Do Xuan, et al.
Pubblicazione: (2025)
TactfulToM: Do LLMs Have the Theory of Mind Ability to Understand White Lies?
di: Liu, Yiwei, et al.
Pubblicazione: (2025)
di: Liu, Yiwei, et al.
Pubblicazione: (2025)
MoreHopQA: More Than Multi-hop Reasoning
di: Schnitzler, Julian, et al.
Pubblicazione: (2024)
di: Schnitzler, Julian, et al.
Pubblicazione: (2024)
Which Feedback Works for Whom? Differential Effects of LLM-Generated Feedback Elements Across Learner Profiles
di: Furuhashi, Momoka, et al.
Pubblicazione: (2026)
di: Furuhashi, Momoka, et al.
Pubblicazione: (2026)
Quality Text, Robust Vision: The Role of Language in Enhancing Visual Robustness of Vision-Language Models
di: Waseda, Futa, et al.
Pubblicazione: (2025)
di: Waseda, Futa, et al.
Pubblicazione: (2025)
What Makes Good Instruction-Tuning Data? An In-Context Learning Perspective
di: Han, Guangzeng, et al.
Pubblicazione: (2026)
di: Han, Guangzeng, et al.
Pubblicazione: (2026)
Label Effects: Shared Heuristic Reliance in Trust Assessment by Humans and LLM-as-a-Judge
di: Sun, Xin, et al.
Pubblicazione: (2026)
di: Sun, Xin, et al.
Pubblicazione: (2026)
What Makes Two Language Models Think Alike?
di: Salle, Jeanne, et al.
Pubblicazione: (2024)
di: Salle, Jeanne, et al.
Pubblicazione: (2024)
What Makes a Good Doctor Response? A Study on Text-Based Telemedicine
di: Cosma, Adrian, et al.
Pubblicazione: (2026)
di: Cosma, Adrian, et al.
Pubblicazione: (2026)
What Makes a Reward Model a Good Teacher? An Optimization Perspective
di: Razin, Noam, et al.
Pubblicazione: (2025)
di: Razin, Noam, et al.
Pubblicazione: (2025)
Automatic Feedback Generation for Short Answer Questions using Answer Diagnostic Graphs
di: Furuhashi, Momoka, et al.
Pubblicazione: (2025)
di: Furuhashi, Momoka, et al.
Pubblicazione: (2025)
Measuring Human Involvement in AI-Generated Text: A Case Study on Academic Writing
di: Guo, Yuchen, et al.
Pubblicazione: (2025)
di: Guo, Yuchen, et al.
Pubblicazione: (2025)
Large Language Models Know What Makes Exemplary Contexts
di: Long, Quanyu, et al.
Pubblicazione: (2024)
di: Long, Quanyu, et al.
Pubblicazione: (2024)
What Makes Diffusion Language Models Super Data Learners?
di: Gao, Zitian, et al.
Pubblicazione: (2025)
di: Gao, Zitian, et al.
Pubblicazione: (2025)
What Makes a Good Response? An Empirical Analysis of Quality in Qualitative Interviews
di: Ivey, Jonathan, et al.
Pubblicazione: (2026)
di: Ivey, Jonathan, et al.
Pubblicazione: (2026)
What Makes Good Multilingual Reasoning? Disentangling Reasoning Traces with Measurable Features
di: Ki, Dayeon, et al.
Pubblicazione: (2026)
di: Ki, Dayeon, et al.
Pubblicazione: (2026)
Sample Design Engineering: An Empirical Study of What Makes Good Downstream Fine-Tuning Samples for LLMs
di: Guo, Biyang, et al.
Pubblicazione: (2024)
di: Guo, Biyang, et al.
Pubblicazione: (2024)
Guiding Through Complexity: What Makes Good Supervision for Hard Math Reasoning Tasks?
di: He, Xuan, et al.
Pubblicazione: (2024)
di: He, Xuan, et al.
Pubblicazione: (2024)
What Makes an LLM a Good Optimizer? A Trajectory Analysis of LLM-Guided Evolutionary Search
di: Zhang, Xinhao, et al.
Pubblicazione: (2026)
di: Zhang, Xinhao, et al.
Pubblicazione: (2026)
What Makes Large Language Models Reason in (Multi-Turn) Code Generation?
di: Zheng, Kunhao, et al.
Pubblicazione: (2024)
di: Zheng, Kunhao, et al.
Pubblicazione: (2024)
Can Large Language Models Robustly Perform Natural Language Inference for Japanese Comparatives?
di: Mikami, Yosuke, et al.
Pubblicazione: (2025)
di: Mikami, Yosuke, et al.
Pubblicazione: (2025)
What Makes for Good Visual Instructions? Synthesizing Complex Visual Reasoning Instructions for Visual Instruction Tuning
di: Du, Yifan, et al.
Pubblicazione: (2023)
di: Du, Yifan, et al.
Pubblicazione: (2023)
What Makes a Good Query? Measuring the Impact of Human-Confusing Linguistic Features on LLM Performance
di: Watson, William, et al.
Pubblicazione: (2026)
di: Watson, William, et al.
Pubblicazione: (2026)
How Many Languages Make Good Multilingual Instruction Tuning? A Case Study on BLOOM
di: Ji, Shaoxiong, et al.
Pubblicazione: (2024)
di: Ji, Shaoxiong, et al.
Pubblicazione: (2024)
Are Large Language Models Good Data Preprocessors?
di: Meguellati, Elyas, et al.
Pubblicazione: (2025)
di: Meguellati, Elyas, et al.
Pubblicazione: (2025)
Intention Analysis Makes LLMs A Good Jailbreak Defender
di: Zhang, Yuqi, et al.
Pubblicazione: (2024)
di: Zhang, Yuqi, et al.
Pubblicazione: (2024)
Is analogy enough to draw novel adjective-noun inferences?
di: Ross, Hayley, et al.
Pubblicazione: (2025)
di: Ross, Hayley, et al.
Pubblicazione: (2025)
Are Large Language Models Good Statisticians?
di: Zhu, Yizhang, et al.
Pubblicazione: (2024)
di: Zhu, Yizhang, et al.
Pubblicazione: (2024)
What Makes Good Data for Alignment? A Comprehensive Study of Automatic Data Selection in Instruction Tuning
di: Liu, Wei, et al.
Pubblicazione: (2023)
di: Liu, Wei, et al.
Pubblicazione: (2023)
A Good Plan is Hard to Find: Aligning Models with Preferences is Misaligned with What Helps Users
di: Balepur, Nishant, et al.
Pubblicazione: (2025)
di: Balepur, Nishant, et al.
Pubblicazione: (2025)
Bias Vector: Mitigating Biases in Language Models with Task Arithmetic Approach
di: Shirafuji, Daiki, et al.
Pubblicazione: (2024)
di: Shirafuji, Daiki, et al.
Pubblicazione: (2024)
Are Large Language Models Actually Good at Text Style Transfer?
di: Mukherjee, Sourabrata, et al.
Pubblicazione: (2024)
di: Mukherjee, Sourabrata, et al.
Pubblicazione: (2024)
One prompt is not enough: Instruction Sensitivity Undermines Embedding Model Evaluation
di: Kostiuk, Yevhen, et al.
Pubblicazione: (2026)
di: Kostiuk, Yevhen, et al.
Pubblicazione: (2026)
Documenti analoghi
-
A Dual-Task Paradigm to Investigate Sentence Comprehension Strategies in Language Models
di: Emura, Rei, et al.
Pubblicazione: (2026) -
CxMP: A Linguistic Minimal-Pair Benchmark for Evaluating Constructional Understanding in Language Models
di: Oba, Miyu, et al.
Pubblicazione: (2026) -
C2: Scalable Rubric-Augmented Reward Modeling from Binary Preferences
di: Kawabata, Akira, et al.
Pubblicazione: (2026) -
Rationale-Aware Answer Verification by Pairwise Self-Evaluation
di: Kawabata, Akira, et al.
Pubblicazione: (2024) -
Specification-Aware Machine Translation and Evaluation for Purpose Alignment
di: Kayano, Yoko, et al.
Pubblicazione: (2025)