MT-PingEval: Evaluating Multi-Turn Collaboration with Private Information Games
Fuente:
arXiv
Guardado en:
| Autores principales: | Eisenstein, Jacob, Huot, Fantine, Fisch, Adam, Berant, Jonathan, Lapata, Mirella |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Don't lie to your friends: Learning what you know from collaborative self-play
por: Eisenstein, Jacob, et al.
Publicado: (2025)
por: Eisenstein, Jacob, et al.
Publicado: (2025)
Learning Steerable Clarification Policies with Collaborative Self-play
por: Berant, Jonathan, et al.
Publicado: (2025)
por: Berant, Jonathan, et al.
Publicado: (2025)
Agents' Room: Narrative Generation through Multi-step Collaboration
por: Huot, Fantine, et al.
Publicado: (2024)
por: Huot, Fantine, et al.
Publicado: (2024)
Help Me Write a Story: Evaluating LLMs' Ability to Generate Writing Feedback
por: Rashkin, Hannah, et al.
Publicado: (2025)
por: Rashkin, Hannah, et al.
Publicado: (2025)
Low-Rank Adaptation for Multilingual Summarization: An Empirical Study
por: Whitehouse, Chenxi, et al.
Publicado: (2023)
por: Whitehouse, Chenxi, et al.
Publicado: (2023)
Evaluating LLMs for Targeted Concept Simplification for Domain-Specific Texts
por: Asthana, Sumit, et al.
Publicado: (2024)
por: Asthana, Sumit, et al.
Publicado: (2024)
Robust Preference Optimization through Reward Model Distillation
por: Fisch, Adam, et al.
Publicado: (2024)
por: Fisch, Adam, et al.
Publicado: (2024)
Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning
por: Setlur, Amrith, et al.
Publicado: (2024)
por: Setlur, Amrith, et al.
Publicado: (2024)
Cost-Optimal Active AI Model Evaluation
por: Angelopoulos, Anastasios N., et al.
Publicado: (2025)
por: Angelopoulos, Anastasios N., et al.
Publicado: (2025)
Long-Form Information Alignment Evaluation Beyond Atomic Facts
por: Zheng, Danna, et al.
Publicado: (2025)
por: Zheng, Danna, et al.
Publicado: (2025)
DOLOMITES: Domain-Specific Long-Form Methodical Tasks
por: Malaviya, Chaitanya, et al.
Publicado: (2024)
por: Malaviya, Chaitanya, et al.
Publicado: (2024)
ALTA: Compiler-Based Analysis of Transformers
por: Shaw, Peter, et al.
Publicado: (2024)
por: Shaw, Peter, et al.
Publicado: (2024)
Plantain: Plan-Answer Interleaved Reasoning
por: Liang, Anthony, et al.
Publicado: (2025)
por: Liang, Anthony, et al.
Publicado: (2025)
Think Before you Write: QA-Guided Reasoning for Character Descriptions in Books
por: Papoudakis, Argyrios, et al.
Publicado: (2026)
por: Papoudakis, Argyrios, et al.
Publicado: (2026)
BookWorm: A Dataset for Character Description and Analysis
por: Papoudakis, Argyrios, et al.
Publicado: (2024)
por: Papoudakis, Argyrios, et al.
Publicado: (2024)
Theoretical guarantees on the best-of-n alignment policy
por: Beirami, Ahmad, et al.
Publicado: (2024)
por: Beirami, Ahmad, et al.
Publicado: (2024)
Meta-Adaptive Prompt Distillation for Few-Shot Visual Question Answering
por: Gupta, Akash, et al.
Publicado: (2025)
por: Gupta, Akash, et al.
Publicado: (2025)
Learning to Plan and Generate Text with Citations
por: Fierro, Constanza, et al.
Publicado: (2024)
por: Fierro, Constanza, et al.
Publicado: (2024)
Never Train from Scratch: Fair Comparison of Long-Sequence Models Requires Data-Driven Priors
por: Amos, Ido, et al.
Publicado: (2023)
por: Amos, Ido, et al.
Publicado: (2023)
Explanatory Summarization with Discourse-Driven Planning
por: Liu, Dongqi, et al.
Publicado: (2025)
por: Liu, Dongqi, et al.
Publicado: (2025)
Long-Context Reasoning Through Proxy-Based Chain-of-Thought Tuning
por: Li, Miao, et al.
Publicado: (2026)
por: Li, Miao, et al.
Publicado: (2026)
$μ$PLAN: Summarizing using a Content Plan as Cross-Lingual Bridge
por: Huot, Fantine, et al.
Publicado: (2023)
por: Huot, Fantine, et al.
Publicado: (2023)
ToolACE-MT: Non-Autoregressive Generation for Agentic Multi-Turn Interaction
por: Zeng, Xingshan, et al.
Publicado: (2025)
por: Zeng, Xingshan, et al.
Publicado: (2025)
SEMQA: Semi-Extractive Multi-Source Question Answering
por: Schuster, Tal, et al.
Publicado: (2023)
por: Schuster, Tal, et al.
Publicado: (2023)
InfAlign: Inference-aware language model alignment
por: Balashankar, Ananth, et al.
Publicado: (2024)
por: Balashankar, Ananth, et al.
Publicado: (2024)
APIGen-MT: Agentic Pipeline for Multi-Turn Data Generation via Simulated Agent-Human Interplay
por: Prabhakar, Akshara, et al.
Publicado: (2025)
por: Prabhakar, Akshara, et al.
Publicado: (2025)
MT-Eval: A Multi-Turn Capabilities Evaluation Benchmark for Large Language Models
por: Kwan, Wai-Chung, et al.
Publicado: (2024)
por: Kwan, Wai-Chung, et al.
Publicado: (2024)
Learning to Reason for Long-Form Story Generation
por: Gurung, Alexander, et al.
Publicado: (2025)
por: Gurung, Alexander, et al.
Publicado: (2025)
A Modular Approach for Multimodal Summarization of TV Shows
por: Mahon, Louis, et al.
Publicado: (2024)
por: Mahon, Louis, et al.
Publicado: (2024)
AMBROSIA: A Benchmark for Parsing Ambiguous Questions into Database Queries
por: Saparina, Irina, et al.
Publicado: (2024)
por: Saparina, Irina, et al.
Publicado: (2024)
CHIRON: Rich Character Representations in Long-Form Narratives
por: Gurung, Alexander, et al.
Publicado: (2024)
por: Gurung, Alexander, et al.
Publicado: (2024)
Integrating Large Language Models with Graph-based Reasoning for Conversational Question Answering
por: Jain, Parag, et al.
Publicado: (2024)
por: Jain, Parag, et al.
Publicado: (2024)
Context-Aware Hierarchical Merging for Long Document Summarization
por: Ou, Litu, et al.
Publicado: (2025)
por: Ou, Litu, et al.
Publicado: (2025)
Improving Generalization in Semantic Parsing by Increasing Natural Language Variation
por: Saparina, Irina, et al.
Publicado: (2024)
por: Saparina, Irina, et al.
Publicado: (2024)
Honest Students from Untrusted Teachers: Learning an Interpretable Question-Answering Pipeline from a Pretrained Language Model
por: Eisenstein, Jacob, et al.
Publicado: (2022)
por: Eisenstein, Jacob, et al.
Publicado: (2022)
Multimodal Latent Reasoning via Predictive Embeddings
por: Adhikari, Ashutosh, et al.
Publicado: (2026)
por: Adhikari, Ashutosh, et al.
Publicado: (2026)
Eliciting Behaviors in Multi-Turn Conversations
por: Huang, Jing, et al.
Publicado: (2025)
por: Huang, Jing, et al.
Publicado: (2025)
Bridging Kolmogorov Complexity and Deep Learning: Asymptotically Optimal Description Length Objectives for Transformers
por: Shaw, Peter, et al.
Publicado: (2025)
por: Shaw, Peter, et al.
Publicado: (2025)
Relaxed Recursive Transformers: Effective Parameter Sharing with Layer-wise LoRA
por: Bae, Sangmin, et al.
Publicado: (2024)
por: Bae, Sangmin, et al.
Publicado: (2024)
Conformal Language Modeling
por: Quach, Victor, et al.
Publicado: (2023)
por: Quach, Victor, et al.
Publicado: (2023)
Ejemplares similares
-
Don't lie to your friends: Learning what you know from collaborative self-play
por: Eisenstein, Jacob, et al.
Publicado: (2025) -
Learning Steerable Clarification Policies with Collaborative Self-play
por: Berant, Jonathan, et al.
Publicado: (2025) -
Agents' Room: Narrative Generation through Multi-step Collaboration
por: Huot, Fantine, et al.
Publicado: (2024) -
Help Me Write a Story: Evaluating LLMs' Ability to Generate Writing Feedback
por: Rashkin, Hannah, et al.
Publicado: (2025) -
Low-Rank Adaptation for Multilingual Summarization: An Empirical Study
por: Whitehouse, Chenxi, et al.
Publicado: (2023)