Awes, Laws, and Flaws From Today's LLM Research
Fuente:
arXiv
Saved in:
| Main Author: | de Wynter, Adrian |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
If Eleanor Rigby Had Met ChatGPT: A Study on Loneliness in a Post-LLM World
by: de Wynter, Adrian
Published: (2024)
by: de Wynter, Adrian
Published: (2024)
If LLMs Have Human-Like Attributes, Then So Does Age of Empires II
by: de Wynter, Adrian
Published: (2026)
by: de Wynter, Adrian
Published: (2026)
Will GPT-4 Run DOOM?
by: de Wynter, Adrian
Published: (2024)
by: de Wynter, Adrian
Published: (2024)
Is In-Context Learning Learning?
by: de Wynter, Adrian
Published: (2025)
by: de Wynter, Adrian
Published: (2025)
Algorithmically Establishing Trust in Evaluators
by: de Wynter, Adrian
Published: (2025)
by: de Wynter, Adrian
Published: (2025)
"I'd Like to Have an Argument, Please": Argumentative Reasoning in Large Language Models
by: de Wynter, Adrian, et al.
Published: (2023)
by: de Wynter, Adrian, et al.
Published: (2023)
The Hrunting of AI: Where and How to Improve English Dialectal Fairness
by: Li, Wei, et al.
Published: (2026)
by: Li, Wei, et al.
Published: (2026)
The Thin Line Between Comprehension and Persuasion in LLMs
by: de Wynter, Adrian, et al.
Published: (2025)
by: de Wynter, Adrian, et al.
Published: (2025)
Collaboratively adding new knowledge to an LLM
by: Lee, Rhui Dih, et al.
Published: (2024)
by: Lee, Rhui Dih, et al.
Published: (2024)
FindTheFlaws: Annotated Errors for Detecting Flawed Reasoning and Scalable Oversight Research
by: Recchia, Gabriel, et al.
Published: (2025)
by: Recchia, Gabriel, et al.
Published: (2025)
A Multilingual, Culture-First Approach to Addressing Misgendering in LLM Applications
by: Sitaram, Sunayana, et al.
Published: (2025)
by: Sitaram, Sunayana, et al.
Published: (2025)
Evaluating Style-Personalized Text Generation: Challenges and Directions
by: Jangra, Anubhav, et al.
Published: (2025)
by: Jangra, Anubhav, et al.
Published: (2025)
Biased or Flawed? Mitigating Stereotypes in Generative Language Models by Addressing Task-Specific Flaws
by: Jha, Akshita, et al.
Published: (2024)
by: Jha, Akshita, et al.
Published: (2024)
LLM as a Mastermind: A Survey of Strategic Reasoning with Large Language Models
by: Zhang, Yadong, et al.
Published: (2024)
by: Zhang, Yadong, et al.
Published: (2024)
Causal Reasoning Favors Encoders: On The Limits of Decoder-Only Models
by: Roy, Amartya, et al.
Published: (2025)
by: Roy, Amartya, et al.
Published: (2025)
Scaling Flaws of Verifier-Guided Search in Mathematical Reasoning
by: Yu, Fei, et al.
Published: (2025)
by: Yu, Fei, et al.
Published: (2025)
The Flaw of Averages: Quantifying Uniformity of Performance on Benchmarks
by: Uzunoglu, Arda, et al.
Published: (2025)
by: Uzunoglu, Arda, et al.
Published: (2025)
On Meta-Prompting
by: de Wynter, Adrian, et al.
Published: (2023)
by: de Wynter, Adrian, et al.
Published: (2023)
Turing Completeness and Sid Meier's Civilization
by: de Wynter, Adrian
Published: (2021)
by: de Wynter, Adrian
Published: (2021)
An Evaluation on Large Language Model Outputs: Discourse and Memorization
by: de Wynter, Adrian, et al.
Published: (2023)
by: de Wynter, Adrian, et al.
Published: (2023)
Are Large Language Model-based Evaluators the Solution to Scaling Up Multilingual Evaluation?
by: Hada, Rishav, et al.
Published: (2023)
by: Hada, Rishav, et al.
Published: (2023)
BenchMarker: An Education-Inspired Toolkit for Highlighting Flaws in Multiple-Choice Benchmarks
by: Balepur, Nishant, et al.
Published: (2026)
by: Balepur, Nishant, et al.
Published: (2026)
Textual Gradients are a Flawed Metaphor for Automatic Prompt Optimization
by: Melcer, Daniel, et al.
Published: (2025)
by: Melcer, Daniel, et al.
Published: (2025)
Flee the Flaw: Annotating the Underlying Logic of Fallacious Arguments Through Templates and Slot-filling
by: Robbani, Irfan, et al.
Published: (2024)
by: Robbani, Irfan, et al.
Published: (2024)
Unveiling the Flaws: Exploring Imperfections in Synthetic Data and Mitigation Strategies for Large Language Models
by: Chen, Jie, et al.
Published: (2024)
by: Chen, Jie, et al.
Published: (2024)
How We Refute Claims: Automatic Fact-Checking through Flaw Identification and Explanation
by: Kao, Wei-Yu, et al.
Published: (2024)
by: Kao, Wei-Yu, et al.
Published: (2024)
Stop! In the Name of Flaws: Disentangling Personal Names and Sociodemographic Attributes in NLP
by: Gautam, Vagrant, et al.
Published: (2024)
by: Gautam, Vagrant, et al.
Published: (2024)
The Impact of Item-Writing Flaws on Difficulty and Discrimination in Item Response Theory
by: Schmucker, Robin, et al.
Published: (2025)
by: Schmucker, Robin, et al.
Published: (2025)
Flaw or Artifact? Rethinking Prompt Sensitivity in Evaluating LLMs
by: Hua, Andong, et al.
Published: (2025)
by: Hua, Andong, et al.
Published: (2025)
Flexible and Effective Mixing of Large Language Models into a Mixture of Domain Experts
by: Lee, Rhui Dih, et al.
Published: (2024)
by: Lee, Rhui Dih, et al.
Published: (2024)
Finding Flawed Fictions: Evaluating Complex Reasoning in Language Models via Plot Hole Detection
by: Ahuja, Kabir, et al.
Published: (2025)
by: Ahuja, Kabir, et al.
Published: (2025)
Jailbreaking Large Language Diffusion Models: Revealing Hidden Safety Flaws in Diffusion-Based Text Generation
by: Zhang, Yuanhe, et al.
Published: (2025)
by: Zhang, Yuanhe, et al.
Published: (2025)
GenAI vs. Human Fact-Checkers: Accurate Ratings, Flawed Rationales
by: Tai, Yuehong Cassandra, et al.
Published: (2025)
by: Tai, Yuehong Cassandra, et al.
Published: (2025)
Architectural Flaw Detection in Civil Engineering Using GPT-4
by: Kumar, Saket, et al.
Published: (2024)
by: Kumar, Saket, et al.
Published: (2024)
Token-Weighted RNN-T for Learning from Flawed Data
by: Keren, Gil, et al.
Published: (2024)
by: Keren, Gil, et al.
Published: (2024)
CitaLaw: Enhancing LLM with Citations in Legal Domain
by: Zhang, Kepu, et al.
Published: (2024)
by: Zhang, Kepu, et al.
Published: (2024)
Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the Above
by: Balepur, Nishant, et al.
Published: (2025)
by: Balepur, Nishant, et al.
Published: (2025)
The State of Multilingual LLM Safety Research: From Measuring the Language Gap to Mitigating It
by: Yong, Zheng-Xin, et al.
Published: (2025)
by: Yong, Zheng-Xin, et al.
Published: (2025)
The Law of Knowledge Overshadowing: Towards Understanding, Predicting, and Preventing LLM Hallucination
by: Zhang, Yuji, et al.
Published: (2025)
by: Zhang, Yuji, et al.
Published: (2025)
The Scaling Laws of Skills in LLM Agent Systems
by: Chen, Charles, et al.
Published: (2026)
by: Chen, Charles, et al.
Published: (2026)
Similar Items
-
If Eleanor Rigby Had Met ChatGPT: A Study on Loneliness in a Post-LLM World
by: de Wynter, Adrian
Published: (2024) -
If LLMs Have Human-Like Attributes, Then So Does Age of Empires II
by: de Wynter, Adrian
Published: (2026) -
Will GPT-4 Run DOOM?
by: de Wynter, Adrian
Published: (2024) -
Is In-Context Learning Learning?
by: de Wynter, Adrian
Published: (2025) -
Algorithmically Establishing Trust in Evaluators
by: de Wynter, Adrian
Published: (2025)