Evaluating Language Model Agency through Negotiations
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Davidson, Tim R., Veselovsky, Veniamin, Josifoski, Martin, Peyrard, Maxime, Bosselut, Antoine, Kosinski, Michal, West, Robert |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Grammar-Constrained Decoding for Structured NLP Tasks without Finetuning
von: Geng, Saibo, et al.
Veröffentlicht: (2023)
von: Geng, Saibo, et al.
Veröffentlicht: (2023)
Self-Recognition in Language Models
von: Davidson, Tim R., et al.
Veröffentlicht: (2024)
von: Davidson, Tim R., et al.
Veröffentlicht: (2024)
Agentic AI: The Era of Semantic Decoding
von: Peyrard, Maxime, et al.
Veröffentlicht: (2024)
von: Peyrard, Maxime, et al.
Veröffentlicht: (2024)
A Glitch in the Matrix? Locating and Detecting Language Model Grounding with Fakepedia
von: Monea, Giovanni, et al.
Veröffentlicht: (2023)
von: Monea, Giovanni, et al.
Veröffentlicht: (2023)
Interactive Evaluation of Large Language Models for Multi-Requirement Software Engineering Tasks
von: Rontogiannis, Dimitrios, et al.
Veröffentlicht: (2025)
von: Rontogiannis, Dimitrios, et al.
Veröffentlicht: (2025)
Separating Tongue from Thought: Activation Patching Reveals Language-Agnostic Concept Representations in Transformers
von: Dumas, Clément, et al.
Veröffentlicht: (2024)
von: Dumas, Clément, et al.
Veröffentlicht: (2024)
Symbolic Autoencoding for Self-Supervised Sequence Learning
von: Amani, Mohammad Hossein, et al.
Veröffentlicht: (2024)
von: Amani, Mohammad Hossein, et al.
Veröffentlicht: (2024)
Identifying and Mitigating the Influence of the Prior Distribution in Large Language Models
von: Zhang, Liyi, et al.
Veröffentlicht: (2025)
von: Zhang, Liyi, et al.
Veröffentlicht: (2025)
What is a Number, That a Large Language Model May Know It?
von: Marjieh, Raja, et al.
Veröffentlicht: (2025)
von: Marjieh, Raja, et al.
Veröffentlicht: (2025)
Localized Cultural Knowledge is Conserved and Controllable in Large Language Models
von: Veselovsky, Veniamin, et al.
Veröffentlicht: (2025)
von: Veselovsky, Veniamin, et al.
Veröffentlicht: (2025)
Mechanistic Interpretability as Statistical Estimation: A Variance Analysis
von: Méloux, Maxime, et al.
Veröffentlicht: (2025)
von: Méloux, Maxime, et al.
Veröffentlicht: (2025)
Everything, Everywhere, All at Once: Is Mechanistic Interpretability Identifiable?
von: Méloux, Maxime, et al.
Veröffentlicht: (2025)
von: Méloux, Maxime, et al.
Veröffentlicht: (2025)
A Logical Fallacy-Informed Framework for Argument Generation
von: Mouchel, Luca, et al.
Veröffentlicht: (2024)
von: Mouchel, Luca, et al.
Veröffentlicht: (2024)
Revisiting Multilingual Data Mixtures in Language Model Pretraining
von: Foroutan, Negar, et al.
Veröffentlicht: (2025)
von: Foroutan, Negar, et al.
Veröffentlicht: (2025)
Reliable Evaluation and Benchmarks for Statement Autoformalization
von: Poiroux, Auguste, et al.
Veröffentlicht: (2024)
von: Poiroux, Auguste, et al.
Veröffentlicht: (2024)
Discovering Knowledge-Critical Subnetworks in Pretrained Language Models
von: Bayazit, Deniz, et al.
Veröffentlicht: (2023)
von: Bayazit, Deniz, et al.
Veröffentlicht: (2023)
Crosscoding Through Time: Tracking Emergence & Consolidation Of Linguistic Representations Throughout LLM Pretraining
von: Bayazit, Deniz, et al.
Veröffentlicht: (2025)
von: Bayazit, Deniz, et al.
Veröffentlicht: (2025)
ConLID: Supervised Contrastive Learning for Low-Resource Language Identification
von: Foroutan, Negar, et al.
Veröffentlicht: (2025)
von: Foroutan, Negar, et al.
Veröffentlicht: (2025)
Rational Metareasoning for Large Language Models
von: De Sabbata, C. Nicolò, et al.
Veröffentlicht: (2024)
von: De Sabbata, C. Nicolò, et al.
Veröffentlicht: (2024)
Do Llamas Work in English? On the Latent Language of Multilingual Transformers
von: Wendler, Chris, et al.
Veröffentlicht: (2024)
von: Wendler, Chris, et al.
Veröffentlicht: (2024)
What Really Controls Temporal Reasoning in Large Language Models: Tokenisation or Representation of Time?
von: Bhatia, Gagan, et al.
Veröffentlicht: (2026)
von: Bhatia, Gagan, et al.
Veröffentlicht: (2026)
Meta-Statistical Learning: Supervised Learning of Statistical Estimators
von: Peyrard, Maxime, et al.
Veröffentlicht: (2025)
von: Peyrard, Maxime, et al.
Veröffentlicht: (2025)
LLMs Are In-Context Bandit Reinforcement Learners
von: Monea, Giovanni, et al.
Veröffentlicht: (2024)
von: Monea, Giovanni, et al.
Veröffentlicht: (2024)
RLMEval: Evaluating Research-Level Neural Theorem Proving
von: Poiroux, Auguste, et al.
Veröffentlicht: (2025)
von: Poiroux, Auguste, et al.
Veröffentlicht: (2025)
Date Fragments: A Hidden Bottleneck of Tokenization for Temporal Reasoning
von: Bhatia, Gagan, et al.
Veröffentlicht: (2025)
von: Bhatia, Gagan, et al.
Veröffentlicht: (2025)
Reasoning-Driven Synthetic Data Generation and Evaluation
von: Davidson, Tim R., et al.
Veröffentlicht: (2026)
von: Davidson, Tim R., et al.
Veröffentlicht: (2026)
REFINER: Reasoning Feedback on Intermediate Representations
von: Paul, Debjit, et al.
Veröffentlicht: (2023)
von: Paul, Debjit, et al.
Veröffentlicht: (2023)
Levels of Analysis for Large Language Models
von: Ku, Alexander Y., et al.
Veröffentlicht: (2025)
von: Ku, Alexander Y., et al.
Veröffentlicht: (2025)
Evaluating Morphological Compositional Generalization in Large Language Models
von: Ismayilzada, Mete, et al.
Veröffentlicht: (2024)
von: Ismayilzada, Mete, et al.
Veröffentlicht: (2024)
Parity-Aware Byte-Pair Encoding: Improving Cross-lingual Fairness in Tokenization
von: Foroutan, Negar, et al.
Veröffentlicht: (2025)
von: Foroutan, Negar, et al.
Veröffentlicht: (2025)
The Collaboration Gap
von: Davidson, Tim R., et al.
Veröffentlicht: (2025)
von: Davidson, Tim R., et al.
Veröffentlicht: (2025)
Complex Reasoning over Logical Queries on Commonsense Knowledge Graphs
von: Fang, Tianqing, et al.
Veröffentlicht: (2024)
von: Fang, Tianqing, et al.
Veröffentlicht: (2024)
PICLe: Pseudo-Annotations for In-Context Learning in Low-Resource Named Entity Detection
von: Mamooler, Sepideh, et al.
Veröffentlicht: (2024)
von: Mamooler, Sepideh, et al.
Veröffentlicht: (2024)
Revisiting Prompt Sensitivity in Large Language Models for Text Classification: The Role of Prompt Underspecification
von: Pecher, Branislav, et al.
Veröffentlicht: (2026)
von: Pecher, Branislav, et al.
Veröffentlicht: (2026)
Transforming Agency. On the mode of existence of Large Language Models
von: Barandiaran, Xabier E., et al.
Veröffentlicht: (2024)
von: Barandiaran, Xabier E., et al.
Veröffentlicht: (2024)
Getting Serious about Humor: Crafting Humor Datasets with Unfunny Large Language Models
von: Horvitz, Zachary, et al.
Veröffentlicht: (2024)
von: Horvitz, Zachary, et al.
Veröffentlicht: (2024)
Evaluating Large Language Models in Theory of Mind Tasks
von: Kosinski, Michal
Veröffentlicht: (2023)
von: Kosinski, Michal
Veröffentlicht: (2023)
Your Finetuned Large Language Model is Already a Powerful Out-of-distribution Detector
von: Zhang, Andi, et al.
Veröffentlicht: (2024)
von: Zhang, Andi, et al.
Veröffentlicht: (2024)
Measuring Evaluation-Context Divergence in Open-Weight LLMs: A Paired-Prompt Protocol with Pilot Evidence of Alignment-Pipeline-Specific Heterogeneity
von: Burnat, Florian A. D., et al.
Veröffentlicht: (2026)
von: Burnat, Florian A. D., et al.
Veröffentlicht: (2026)
Creativity in AI: Progresses and Challenges
von: Ismayilzada, Mete, et al.
Veröffentlicht: (2024)
von: Ismayilzada, Mete, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Grammar-Constrained Decoding for Structured NLP Tasks without Finetuning
von: Geng, Saibo, et al.
Veröffentlicht: (2023) -
Self-Recognition in Language Models
von: Davidson, Tim R., et al.
Veröffentlicht: (2024) -
Agentic AI: The Era of Semantic Decoding
von: Peyrard, Maxime, et al.
Veröffentlicht: (2024) -
A Glitch in the Matrix? Locating and Detecting Language Model Grounding with Fakepedia
von: Monea, Giovanni, et al.
Veröffentlicht: (2023) -
Interactive Evaluation of Large Language Models for Multi-Requirement Software Engineering Tasks
von: Rontogiannis, Dimitrios, et al.
Veröffentlicht: (2025)