Awes, Laws, and Flaws From Today's LLM Research
Fuente:
arXiv
Salvato in:
| Autore principale: | de Wynter, Adrian |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
If Eleanor Rigby Had Met ChatGPT: A Study on Loneliness in a Post-LLM World
di: de Wynter, Adrian
Pubblicazione: (2024)
di: de Wynter, Adrian
Pubblicazione: (2024)
If LLMs Have Human-Like Attributes, Then So Does Age of Empires II
di: de Wynter, Adrian
Pubblicazione: (2026)
di: de Wynter, Adrian
Pubblicazione: (2026)
Will GPT-4 Run DOOM?
di: de Wynter, Adrian
Pubblicazione: (2024)
di: de Wynter, Adrian
Pubblicazione: (2024)
Is In-Context Learning Learning?
di: de Wynter, Adrian
Pubblicazione: (2025)
di: de Wynter, Adrian
Pubblicazione: (2025)
Algorithmically Establishing Trust in Evaluators
di: de Wynter, Adrian
Pubblicazione: (2025)
di: de Wynter, Adrian
Pubblicazione: (2025)
"I'd Like to Have an Argument, Please": Argumentative Reasoning in Large Language Models
di: de Wynter, Adrian, et al.
Pubblicazione: (2023)
di: de Wynter, Adrian, et al.
Pubblicazione: (2023)
The Hrunting of AI: Where and How to Improve English Dialectal Fairness
di: Li, Wei, et al.
Pubblicazione: (2026)
di: Li, Wei, et al.
Pubblicazione: (2026)
The Thin Line Between Comprehension and Persuasion in LLMs
di: de Wynter, Adrian, et al.
Pubblicazione: (2025)
di: de Wynter, Adrian, et al.
Pubblicazione: (2025)
Collaboratively adding new knowledge to an LLM
di: Lee, Rhui Dih, et al.
Pubblicazione: (2024)
di: Lee, Rhui Dih, et al.
Pubblicazione: (2024)
FindTheFlaws: Annotated Errors for Detecting Flawed Reasoning and Scalable Oversight Research
di: Recchia, Gabriel, et al.
Pubblicazione: (2025)
di: Recchia, Gabriel, et al.
Pubblicazione: (2025)
A Multilingual, Culture-First Approach to Addressing Misgendering in LLM Applications
di: Sitaram, Sunayana, et al.
Pubblicazione: (2025)
di: Sitaram, Sunayana, et al.
Pubblicazione: (2025)
Evaluating Style-Personalized Text Generation: Challenges and Directions
di: Jangra, Anubhav, et al.
Pubblicazione: (2025)
di: Jangra, Anubhav, et al.
Pubblicazione: (2025)
Biased or Flawed? Mitigating Stereotypes in Generative Language Models by Addressing Task-Specific Flaws
di: Jha, Akshita, et al.
Pubblicazione: (2024)
di: Jha, Akshita, et al.
Pubblicazione: (2024)
LLM as a Mastermind: A Survey of Strategic Reasoning with Large Language Models
di: Zhang, Yadong, et al.
Pubblicazione: (2024)
di: Zhang, Yadong, et al.
Pubblicazione: (2024)
Causal Reasoning Favors Encoders: On The Limits of Decoder-Only Models
di: Roy, Amartya, et al.
Pubblicazione: (2025)
di: Roy, Amartya, et al.
Pubblicazione: (2025)
Scaling Flaws of Verifier-Guided Search in Mathematical Reasoning
di: Yu, Fei, et al.
Pubblicazione: (2025)
di: Yu, Fei, et al.
Pubblicazione: (2025)
The Flaw of Averages: Quantifying Uniformity of Performance on Benchmarks
di: Uzunoglu, Arda, et al.
Pubblicazione: (2025)
di: Uzunoglu, Arda, et al.
Pubblicazione: (2025)
On Meta-Prompting
di: de Wynter, Adrian, et al.
Pubblicazione: (2023)
di: de Wynter, Adrian, et al.
Pubblicazione: (2023)
Turing Completeness and Sid Meier's Civilization
di: de Wynter, Adrian
Pubblicazione: (2021)
di: de Wynter, Adrian
Pubblicazione: (2021)
An Evaluation on Large Language Model Outputs: Discourse and Memorization
di: de Wynter, Adrian, et al.
Pubblicazione: (2023)
di: de Wynter, Adrian, et al.
Pubblicazione: (2023)
Are Large Language Model-based Evaluators the Solution to Scaling Up Multilingual Evaluation?
di: Hada, Rishav, et al.
Pubblicazione: (2023)
di: Hada, Rishav, et al.
Pubblicazione: (2023)
BenchMarker: An Education-Inspired Toolkit for Highlighting Flaws in Multiple-Choice Benchmarks
di: Balepur, Nishant, et al.
Pubblicazione: (2026)
di: Balepur, Nishant, et al.
Pubblicazione: (2026)
Textual Gradients are a Flawed Metaphor for Automatic Prompt Optimization
di: Melcer, Daniel, et al.
Pubblicazione: (2025)
di: Melcer, Daniel, et al.
Pubblicazione: (2025)
Flee the Flaw: Annotating the Underlying Logic of Fallacious Arguments Through Templates and Slot-filling
di: Robbani, Irfan, et al.
Pubblicazione: (2024)
di: Robbani, Irfan, et al.
Pubblicazione: (2024)
Unveiling the Flaws: Exploring Imperfections in Synthetic Data and Mitigation Strategies for Large Language Models
di: Chen, Jie, et al.
Pubblicazione: (2024)
di: Chen, Jie, et al.
Pubblicazione: (2024)
How We Refute Claims: Automatic Fact-Checking through Flaw Identification and Explanation
di: Kao, Wei-Yu, et al.
Pubblicazione: (2024)
di: Kao, Wei-Yu, et al.
Pubblicazione: (2024)
Stop! In the Name of Flaws: Disentangling Personal Names and Sociodemographic Attributes in NLP
di: Gautam, Vagrant, et al.
Pubblicazione: (2024)
di: Gautam, Vagrant, et al.
Pubblicazione: (2024)
The Impact of Item-Writing Flaws on Difficulty and Discrimination in Item Response Theory
di: Schmucker, Robin, et al.
Pubblicazione: (2025)
di: Schmucker, Robin, et al.
Pubblicazione: (2025)
Flaw or Artifact? Rethinking Prompt Sensitivity in Evaluating LLMs
di: Hua, Andong, et al.
Pubblicazione: (2025)
di: Hua, Andong, et al.
Pubblicazione: (2025)
Flexible and Effective Mixing of Large Language Models into a Mixture of Domain Experts
di: Lee, Rhui Dih, et al.
Pubblicazione: (2024)
di: Lee, Rhui Dih, et al.
Pubblicazione: (2024)
Finding Flawed Fictions: Evaluating Complex Reasoning in Language Models via Plot Hole Detection
di: Ahuja, Kabir, et al.
Pubblicazione: (2025)
di: Ahuja, Kabir, et al.
Pubblicazione: (2025)
Jailbreaking Large Language Diffusion Models: Revealing Hidden Safety Flaws in Diffusion-Based Text Generation
di: Zhang, Yuanhe, et al.
Pubblicazione: (2025)
di: Zhang, Yuanhe, et al.
Pubblicazione: (2025)
GenAI vs. Human Fact-Checkers: Accurate Ratings, Flawed Rationales
di: Tai, Yuehong Cassandra, et al.
Pubblicazione: (2025)
di: Tai, Yuehong Cassandra, et al.
Pubblicazione: (2025)
Architectural Flaw Detection in Civil Engineering Using GPT-4
di: Kumar, Saket, et al.
Pubblicazione: (2024)
di: Kumar, Saket, et al.
Pubblicazione: (2024)
Token-Weighted RNN-T for Learning from Flawed Data
di: Keren, Gil, et al.
Pubblicazione: (2024)
di: Keren, Gil, et al.
Pubblicazione: (2024)
CitaLaw: Enhancing LLM with Citations in Legal Domain
di: Zhang, Kepu, et al.
Pubblicazione: (2024)
di: Zhang, Kepu, et al.
Pubblicazione: (2024)
Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the Above
di: Balepur, Nishant, et al.
Pubblicazione: (2025)
di: Balepur, Nishant, et al.
Pubblicazione: (2025)
The State of Multilingual LLM Safety Research: From Measuring the Language Gap to Mitigating It
di: Yong, Zheng-Xin, et al.
Pubblicazione: (2025)
di: Yong, Zheng-Xin, et al.
Pubblicazione: (2025)
The Law of Knowledge Overshadowing: Towards Understanding, Predicting, and Preventing LLM Hallucination
di: Zhang, Yuji, et al.
Pubblicazione: (2025)
di: Zhang, Yuji, et al.
Pubblicazione: (2025)
The Scaling Laws of Skills in LLM Agent Systems
di: Chen, Charles, et al.
Pubblicazione: (2026)
di: Chen, Charles, et al.
Pubblicazione: (2026)
Documenti analoghi
-
If Eleanor Rigby Had Met ChatGPT: A Study on Loneliness in a Post-LLM World
di: de Wynter, Adrian
Pubblicazione: (2024) -
If LLMs Have Human-Like Attributes, Then So Does Age of Empires II
di: de Wynter, Adrian
Pubblicazione: (2026) -
Will GPT-4 Run DOOM?
di: de Wynter, Adrian
Pubblicazione: (2024) -
Is In-Context Learning Learning?
di: de Wynter, Adrian
Pubblicazione: (2025) -
Algorithmically Establishing Trust in Evaluators
di: de Wynter, Adrian
Pubblicazione: (2025)