Salvato in:
| Autori principali: | Lior, Gili, Nacchace, Liron, Stanovsky, Gabriel |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2502.17091 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Comparing Humans and Models on a Similar Scale: Towards Cognitive Gender Bias Evaluation in Coreference Resolution
di: Lior, Gili, et al.
Pubblicazione: (2023)
di: Lior, Gili, et al.
Pubblicazione: (2023)
Leveraging Collection-Wide Similarities for Unsupervised Document Structure Extraction
di: Lior, Gili, et al.
Pubblicazione: (2024)
di: Lior, Gili, et al.
Pubblicazione: (2024)
Planted in Pretraining, Swayed by Finetuning: A Case Study on the Origins of Cognitive Biases in LLMs
di: Itzhak, Itay, et al.
Pubblicazione: (2025)
di: Itzhak, Itay, et al.
Pubblicazione: (2025)
Can LLMs Help Uncover Insights about LLMs? A Large-Scale, Evolving Literature Analysis of Frontier LLMs
di: Park, Jungsoo, et al.
Pubblicazione: (2025)
di: Park, Jungsoo, et al.
Pubblicazione: (2025)
PromptSuite: A Task-Agnostic Framework for Multi-Prompt Generation
di: Habba, Eliya, et al.
Pubblicazione: (2025)
di: Habba, Eliya, et al.
Pubblicazione: (2025)
From Feelings to Metrics: Understanding and Formalizing How Users Vibe-Test LLMs
di: Itzhak, Itay, et al.
Pubblicazione: (2026)
di: Itzhak, Itay, et al.
Pubblicazione: (2026)
Time to Talk: LLM Agents for Asynchronous Group Communication in Mafia Games
di: Eckhaus, Niv, et al.
Pubblicazione: (2025)
di: Eckhaus, Niv, et al.
Pubblicazione: (2025)
WildIFEval: Instruction Following in the Wild
di: Lior, Gili, et al.
Pubblicazione: (2025)
di: Lior, Gili, et al.
Pubblicazione: (2025)
ReliableEval: A Recipe for Stochastic LLM Evaluation via Method of Moments
di: Lior, Gili, et al.
Pubblicazione: (2025)
di: Lior, Gili, et al.
Pubblicazione: (2025)
A Nurse is Blue and Elephant is Rugby: Cross Domain Alignment in Large Language Models Reveal Human-like Patterns
di: Yehudai, Asaf, et al.
Pubblicazione: (2024)
di: Yehudai, Asaf, et al.
Pubblicazione: (2024)
Anticipatory Evaluation of Language Models
di: Park, Jungsoo, et al.
Pubblicazione: (2025)
di: Park, Jungsoo, et al.
Pubblicazione: (2025)
Beyond Benchmarks: On The False Promise of AI Regulation
di: Stanovsky, Gabriel, et al.
Pubblicazione: (2025)
di: Stanovsky, Gabriel, et al.
Pubblicazione: (2025)
From Tokens to Thoughts: How LLMs and Humans Trade Compression for Meaning
di: Shani, Chen, et al.
Pubblicazione: (2025)
di: Shani, Chen, et al.
Pubblicazione: (2025)
SEAM: A Stochastic Benchmark for Multi-Document Tasks
di: Lior, Gili, et al.
Pubblicazione: (2024)
di: Lior, Gili, et al.
Pubblicazione: (2024)
Text Annotation via Inductive Coding: Comparing Human Experts to LLMs in Qualitative Data Analysis
di: Parfenova, Angelina, et al.
Pubblicazione: (2025)
di: Parfenova, Angelina, et al.
Pubblicazione: (2025)
Cooking Up Creativity: Enhancing LLM Creativity through Structured Recombination
di: Mizrahi, Moran, et al.
Pubblicazione: (2025)
di: Mizrahi, Moran, et al.
Pubblicazione: (2025)
Estimating Causal Effects of Text Interventions Leveraging LLMs
di: Guo, Siyi, et al.
Pubblicazione: (2024)
di: Guo, Siyi, et al.
Pubblicazione: (2024)
A Comparative Analysis of Instruction Fine-Tuning LLMs for Financial Text Classification
di: Fatemi, Sorouralsadat, et al.
Pubblicazione: (2024)
di: Fatemi, Sorouralsadat, et al.
Pubblicazione: (2024)
DeFrame: Debiasing Large Language Models Against Framing Effects
di: Lim, Kahee, et al.
Pubblicazione: (2026)
di: Lim, Kahee, et al.
Pubblicazione: (2026)
Framing Political Bias in Multilingual LLMs Across Pakistani Languages
di: Nadeem, Afrozah, et al.
Pubblicazione: (2025)
di: Nadeem, Afrozah, et al.
Pubblicazione: (2025)
Forget What You Know about LLMs Evaluations -- LLMs are Like a Chameleon
di: Cohen-Inger, Nurit, et al.
Pubblicazione: (2025)
di: Cohen-Inger, Nurit, et al.
Pubblicazione: (2025)
LLMs and Cultural Values: the Impact of Prompt Language and Explicit Cultural Framing
di: Bulté, Bram, et al.
Pubblicazione: (2025)
di: Bulté, Bram, et al.
Pubblicazione: (2025)
Small Language Models can Outperform Humans in Short Creative Writing: A Study Comparing SLMs with Humans and LLMs
di: Marco, Guillermo, et al.
Pubblicazione: (2024)
di: Marco, Guillermo, et al.
Pubblicazione: (2024)
PRILoRA: Pruned and Rank-Increasing Low-Rank Adaptation
di: Benedek, Nadav, et al.
Pubblicazione: (2024)
di: Benedek, Nadav, et al.
Pubblicazione: (2024)
Measuring Embedding Sensitivity to Authorial Style in French: Comparing Literary Texts with Language Model Rewritings
di: Icard, Benjamin, et al.
Pubblicazione: (2026)
di: Icard, Benjamin, et al.
Pubblicazione: (2026)
VietMix: A Naturally-Occurring Parallel Corpus and Augmentation Framework for Vietnamese-English Code-Mixed Machine Translation
di: Tran, Hieu, et al.
Pubblicazione: (2025)
di: Tran, Hieu, et al.
Pubblicazione: (2025)
MuTSE: A Human-in-the-Loop Multi-use Text Simplification Evaluator
di: Roscan, Rares-Alexandru, et al.
Pubblicazione: (2026)
di: Roscan, Rares-Alexandru, et al.
Pubblicazione: (2026)
Can LLMs Narrate Tabular Data? An Evaluation Framework for Natural Language Representations of Text-to-SQL System Outputs
di: Singh, Jyotika, et al.
Pubblicazione: (2025)
di: Singh, Jyotika, et al.
Pubblicazione: (2025)
Comparing LLM Text Annotation Skills: A Study on Human Rights Violations in Social Media Data
di: Nemkova, Poli Apollinaire, et al.
Pubblicazione: (2025)
di: Nemkova, Poli Apollinaire, et al.
Pubblicazione: (2025)
SUPERNOVA: Eliciting General Reasoning in LLMs with Reinforcement Learning on Natural Instructions
di: Suvarna, Ashima, et al.
Pubblicazione: (2026)
di: Suvarna, Ashima, et al.
Pubblicazione: (2026)
LLMs Enable Bag-of-Texts Representations for Short-Text Clustering
di: Lin, I-Fan, et al.
Pubblicazione: (2025)
di: Lin, I-Fan, et al.
Pubblicazione: (2025)
TextQuests: How Good are LLMs at Text-Based Video Games?
di: Phan, Long, et al.
Pubblicazione: (2025)
di: Phan, Long, et al.
Pubblicazione: (2025)
How English Print Media Frames Human-Elephant Conflicts in India
di: Punith, Bonala Sai, et al.
Pubblicazione: (2026)
di: Punith, Bonala Sai, et al.
Pubblicazione: (2026)
Faithfulness Serum: Mitigating the Faithfulness Gap in Textual Explanations of LLM Decisions via Attribution Guidance
di: Alon, Bar, et al.
Pubblicazione: (2026)
di: Alon, Bar, et al.
Pubblicazione: (2026)
Conditioning LLMs to Generate Code-Switched Text
di: Heredia, Maite, et al.
Pubblicazione: (2025)
di: Heredia, Maite, et al.
Pubblicazione: (2025)
LLMs as Strategic Actors: Behavioral Alignment, Risk Calibration, and Argumentation Framing in Geopolitical Simulations
di: Solopova, Veronika, et al.
Pubblicazione: (2026)
di: Solopova, Veronika, et al.
Pubblicazione: (2026)
TextAtari: 100K Frames Game Playing with Language Agents
di: Li, Wenhao, et al.
Pubblicazione: (2025)
di: Li, Wenhao, et al.
Pubblicazione: (2025)
NATURAL PLAN: Benchmarking LLMs on Natural Language Planning
di: Zheng, Huaixiu Steven, et al.
Pubblicazione: (2024)
di: Zheng, Huaixiu Steven, et al.
Pubblicazione: (2024)
Text or Pixels? It Takes Half: On the Token Efficiency of Visual Text Inputs in Multimodal LLMs
di: Li, Yanhong, et al.
Pubblicazione: (2025)
di: Li, Yanhong, et al.
Pubblicazione: (2025)
Humans and LLMs Diverge on Probabilistic Inferences
di: Kamath, Gaurav, et al.
Pubblicazione: (2026)
di: Kamath, Gaurav, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Comparing Humans and Models on a Similar Scale: Towards Cognitive Gender Bias Evaluation in Coreference Resolution
di: Lior, Gili, et al.
Pubblicazione: (2023) -
Leveraging Collection-Wide Similarities for Unsupervised Document Structure Extraction
di: Lior, Gili, et al.
Pubblicazione: (2024) -
Planted in Pretraining, Swayed by Finetuning: A Case Study on the Origins of Cognitive Biases in LLMs
di: Itzhak, Itay, et al.
Pubblicazione: (2025) -
Can LLMs Help Uncover Insights about LLMs? A Large-Scale, Evolving Literature Analysis of Frontier LLMs
di: Park, Jungsoo, et al.
Pubblicazione: (2025) -
PromptSuite: A Task-Agnostic Framework for Multi-Prompt Generation
di: Habba, Eliya, et al.
Pubblicazione: (2025)