From Tools to Teammates: Evaluating LLMs in Multi-Session Coding Interactions
Fuente:
arXiv
Guardado en:
| Autores principales: | Rakotonirina, Nathanaël Carraz, Hamdy, Mohammed, Campos, Jon Ander, Weber, Lucas, Testoni, Alberto, Fadaee, Marzieh, Pezzelle, Sandro, Del Tredici, Marco |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
MemoryPrompt: A Light Wrapper to Improve Context Tracking in Pre-trained Language Models
por: Rakotonirina, Nathanaël Carraz, et al.
Publicado: (2024)
por: Rakotonirina, Nathanaël Carraz, et al.
Publicado: (2024)
Naming, Describing, and Quantifying Visual Objects in Humans and LLMs
por: Testoni, Alberto, et al.
Publicado: (2024)
por: Testoni, Alberto, et al.
Publicado: (2024)
Evil twins are not that evil: Qualitative insights into machine-generated prompts
por: Rakotonirina, Nathanaël Carraz, et al.
Publicado: (2024)
por: Rakotonirina, Nathanaël Carraz, et al.
Publicado: (2024)
Optimizing with Low Budgets: a Comparison on the Black-box Optimization Benchmarking Suite and OpenAI Gym
por: Raponi, Elena, et al.
Publicado: (2023)
por: Raponi, Elena, et al.
Publicado: (2023)
Correct, Concise and Complete: Multi-stage Training For Adaptive Reasoning
por: Rakotonirina, Nathanaël Carraz, et al.
Publicado: (2026)
por: Rakotonirina, Nathanaël Carraz, et al.
Publicado: (2026)
Automatic WordNet Construction Using Markov Chain Monte Carlo
por: Marzieh Fadaee
Publicado: (2013)
por: Marzieh Fadaee
Publicado: (2013)
Ginkgos and People: A Thousand Years of Interaction
por: Del Tredici, Peter
Publicado: (1991)
por: Del Tredici, Peter
Publicado: (1991)
Chapter Introduzione
por: Carocci, Sandro, et al.
Publicado: (2023)
por: Carocci, Sandro, et al.
Publicado: (2023)
They want to pretend not to understand: The Limits of Current LLMs in Interpreting Implicit Content of Political Discourse
por: Paci, Walter, et al.
Publicado: (2025)
por: Paci, Walter, et al.
Publicado: (2025)
The Upright White Pine
por: Del Tredici, Peter
Publicado: (1993)
por: Del Tredici, Peter
Publicado: (1993)
What's in a Leaf?
por: Del Tredici, Peter
Publicado: (1985)
por: Del Tredici, Peter
Publicado: (1985)
Chapter Signorie rurali e poteri superiori in Italia settentrionale (secoli XIV-XV)
por: Del Tredici, Federico
Publicado: (2023)
por: Del Tredici, Federico
Publicado: (2023)
Chapter L’estensione del dominio dell’amicizia. Signori e amici in Lombardia e Italia centro-settentrionale, secoli XIV-XV
por: Del Tredici, Federico
Publicado: (2022)
por: Del Tredici, Federico
Publicado: (2022)
Valores en conflicto: análisis de representaciones y estrategias de resiliencia climática en la Reserva Natural Urbana San Martín
por: Romina Del Tredici
Publicado: (2023)
por: Romina Del Tredici
Publicado: (2023)
Comprando paz social: la distribución de planes sociales durante los gobiernos de Cristina Kirchner y Mauricio Macri*
por: Romina Del Tredici
Publicado: (2023)
por: Romina Del Tredici
Publicado: (2023)
Capital social redistributivo: los efectos de las redes de compromiso en la desigualdad subnacional en Argentina
por: Romina Del Tredici
Publicado: (2022)
por: Romina Del Tredici
Publicado: (2022)
To LLM, or Not to LLM: How Designers and Developers Navigate LLMs as Tools or Teammates
por: Vishwarupe, Varad, et al.
Publicado: (2026)
por: Vishwarupe, Varad, et al.
Publicado: (2026)
When to Retrieve: Teaching LLMs to Utilize Information Retrieval Effectively
por: Labruna, Tiziano, et al.
Publicado: (2024)
por: Labruna, Tiziano, et al.
Publicado: (2024)
AI as Teammate or Tool? A Review of Human-AI Interaction in Decision Support
por: Samu, Most. Sharmin Sultana, et al.
Publicado: (2026)
por: Samu, Most. Sharmin Sultana, et al.
Publicado: (2026)
Transportation networks and competition in the market for corporate control
por: Marco Testoni
Publicado: (2024)
por: Marco Testoni
Publicado: (2024)
Calycanthus chinensis: The Chinese Sweetshrub
por: Li, Jianhua, et al.
Publicado: (2005)
por: Li, Jianhua, et al.
Publicado: (2005)
Verification Limits Code LLM Training
por: Gureja, Srishti, et al.
Publicado: (2025)
por: Gureja, Srishti, et al.
Publicado: (2025)
Are formal and functional linguistic mechanisms dissociated in language models?
por: Hanna, Michael, et al.
Publicado: (2025)
por: Hanna, Michael, et al.
Publicado: (2025)
Have Faith in Faithfulness: Going Beyond Circuit Overlap When Finding Model Mechanisms
por: Hanna, Michael, et al.
Publicado: (2024)
por: Hanna, Michael, et al.
Publicado: (2024)
Who is the richest club in the championship? Detecting and Rewriting Underspecified Questions Improve QA Performance
por: Huang, Yunchong, et al.
Publicado: (2026)
por: Huang, Yunchong, et al.
Publicado: (2026)
Describing Images $\textit{Fast and Slow}$: Quantifying and Predicting the Variation in Human Signals during Visuo-Linguistic Processes
por: Takmaz, Ece, et al.
Publicado: (2024)
por: Takmaz, Ece, et al.
Publicado: (2024)
How Language Models Conflate Logical Validity with Plausibility: A Representational Analysis of Content Effects
por: Bertolazzi, Leonardo, et al.
Publicado: (2025)
por: Bertolazzi, Leonardo, et al.
Publicado: (2025)
Beyond Divergent Creativity: A Human-Based Evaluation of Creativity in Large Language Models
por: Nakajima, Kumiko, et al.
Publicado: (2026)
por: Nakajima, Kumiko, et al.
Publicado: (2026)
Do Pre-Trained Language Models Detect and Understand Semantic Underspecification? Ask the DUST!
por: Wildenburg, Frank, et al.
Publicado: (2024)
por: Wildenburg, Frank, et al.
Publicado: (2024)
The BLA Benchmark: Investigating Basic Language Abilities of Pre-Trained Multimodal Models
por: Chen, Xinyi, et al.
Publicado: (2023)
por: Chen, Xinyi, et al.
Publicado: (2023)
Is my model perplexed for the right reason? Contrasting LLMs' Benchmark Behavior with Token-Level Perplexity
por: Prins, Zoë, et al.
Publicado: (2026)
por: Prins, Zoë, et al.
Publicado: (2026)
RAcQUEt: Unveiling the Dangers of Overlooked Referential Ambiguity in Visual LLMs
por: Testoni, Alberto, et al.
Publicado: (2024)
por: Testoni, Alberto, et al.
Publicado: (2024)
Learning to Ask Informative Questions: Enhancing LLMs with Preference Optimization and Expected Information Gain
por: Mazzaccara, Davide, et al.
Publicado: (2024)
por: Mazzaccara, Davide, et al.
Publicado: (2024)
From Tool to Teammate: LLM Coding Agents as Collaborative Partners for Behavioral Labeling in Educational Dialogue Analysis
por: Chen, Eason, et al.
Publicado: (2026)
por: Chen, Eason, et al.
Publicado: (2026)
A Post-trainer's Guide to Multilingual Training Data: Uncovering Cross-lingual Transfer Dynamics
por: Shimabucoro, Luisa, et al.
Publicado: (2025)
por: Shimabucoro, Luisa, et al.
Publicado: (2025)
Diversify and Conquer: Diversity-Centric Data Selection with Iterative Refinement
por: Yu, Simon, et al.
Publicado: (2024)
por: Yu, Simon, et al.
Publicado: (2024)
Not (yet) the whole story: Evaluating Visual Storytelling Requires More than Measuring Coherence, Grounding, and Repetition
por: Surikuchi, Aditya K, et al.
Publicado: (2024)
por: Surikuchi, Aditya K, et al.
Publicado: (2024)
Natural Language Generation from Visual Events: State-of-the-Art and Key Open Questions
por: Surikuchi, Aditya K, et al.
Publicado: (2025)
por: Surikuchi, Aditya K, et al.
Publicado: (2025)
Where is the multimodal goal post? On the Ability of Foundation Models to Recognize Contextually Important Moments
por: Surikuchi, Aditya K, et al.
Publicado: (2026)
por: Surikuchi, Aditya K, et al.
Publicado: (2026)
To Code, or Not To Code? Exploring Impact of Code in Pre-training
por: Aryabumi, Viraat, et al.
Publicado: (2024)
por: Aryabumi, Viraat, et al.
Publicado: (2024)
Ejemplares similares
-
MemoryPrompt: A Light Wrapper to Improve Context Tracking in Pre-trained Language Models
por: Rakotonirina, Nathanaël Carraz, et al.
Publicado: (2024) -
Naming, Describing, and Quantifying Visual Objects in Humans and LLMs
por: Testoni, Alberto, et al.
Publicado: (2024) -
Evil twins are not that evil: Qualitative insights into machine-generated prompts
por: Rakotonirina, Nathanaël Carraz, et al.
Publicado: (2024) -
Optimizing with Low Budgets: a Comparison on the Black-box Optimization Benchmarking Suite and OpenAI Gym
por: Raponi, Elena, et al.
Publicado: (2023) -
Correct, Concise and Complete: Multi-stage Training For Adaptive Reasoning
por: Rakotonirina, Nathanaël Carraz, et al.
Publicado: (2026)