From Tools to Teammates: Evaluating LLMs in Multi-Session Coding Interactions
Fuente:
arXiv
Salvato in:
| Autori principali: | Rakotonirina, Nathanaël Carraz, Hamdy, Mohammed, Campos, Jon Ander, Weber, Lucas, Testoni, Alberto, Fadaee, Marzieh, Pezzelle, Sandro, Del Tredici, Marco |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MemoryPrompt: A Light Wrapper to Improve Context Tracking in Pre-trained Language Models
di: Rakotonirina, Nathanaël Carraz, et al.
Pubblicazione: (2024)
di: Rakotonirina, Nathanaël Carraz, et al.
Pubblicazione: (2024)
Naming, Describing, and Quantifying Visual Objects in Humans and LLMs
di: Testoni, Alberto, et al.
Pubblicazione: (2024)
di: Testoni, Alberto, et al.
Pubblicazione: (2024)
Evil twins are not that evil: Qualitative insights into machine-generated prompts
di: Rakotonirina, Nathanaël Carraz, et al.
Pubblicazione: (2024)
di: Rakotonirina, Nathanaël Carraz, et al.
Pubblicazione: (2024)
Optimizing with Low Budgets: a Comparison on the Black-box Optimization Benchmarking Suite and OpenAI Gym
di: Raponi, Elena, et al.
Pubblicazione: (2023)
di: Raponi, Elena, et al.
Pubblicazione: (2023)
Correct, Concise and Complete: Multi-stage Training For Adaptive Reasoning
di: Rakotonirina, Nathanaël Carraz, et al.
Pubblicazione: (2026)
di: Rakotonirina, Nathanaël Carraz, et al.
Pubblicazione: (2026)
Automatic WordNet Construction Using Markov Chain Monte Carlo
di: Marzieh Fadaee
Pubblicazione: (2013)
di: Marzieh Fadaee
Pubblicazione: (2013)
Ginkgos and People: A Thousand Years of Interaction
di: Del Tredici, Peter
Pubblicazione: (1991)
di: Del Tredici, Peter
Pubblicazione: (1991)
Chapter Introduzione
di: Carocci, Sandro, et al.
Pubblicazione: (2023)
di: Carocci, Sandro, et al.
Pubblicazione: (2023)
They want to pretend not to understand: The Limits of Current LLMs in Interpreting Implicit Content of Political Discourse
di: Paci, Walter, et al.
Pubblicazione: (2025)
di: Paci, Walter, et al.
Pubblicazione: (2025)
The Upright White Pine
di: Del Tredici, Peter
Pubblicazione: (1993)
di: Del Tredici, Peter
Pubblicazione: (1993)
What's in a Leaf?
di: Del Tredici, Peter
Pubblicazione: (1985)
di: Del Tredici, Peter
Pubblicazione: (1985)
Chapter Signorie rurali e poteri superiori in Italia settentrionale (secoli XIV-XV)
di: Del Tredici, Federico
Pubblicazione: (2023)
di: Del Tredici, Federico
Pubblicazione: (2023)
Chapter L’estensione del dominio dell’amicizia. Signori e amici in Lombardia e Italia centro-settentrionale, secoli XIV-XV
di: Del Tredici, Federico
Pubblicazione: (2022)
di: Del Tredici, Federico
Pubblicazione: (2022)
Valores en conflicto: análisis de representaciones y estrategias de resiliencia climática en la Reserva Natural Urbana San Martín
di: Romina Del Tredici
Pubblicazione: (2023)
di: Romina Del Tredici
Pubblicazione: (2023)
Comprando paz social: la distribución de planes sociales durante los gobiernos de Cristina Kirchner y Mauricio Macri*
di: Romina Del Tredici
Pubblicazione: (2023)
di: Romina Del Tredici
Pubblicazione: (2023)
Capital social redistributivo: los efectos de las redes de compromiso en la desigualdad subnacional en Argentina
di: Romina Del Tredici
Pubblicazione: (2022)
di: Romina Del Tredici
Pubblicazione: (2022)
To LLM, or Not to LLM: How Designers and Developers Navigate LLMs as Tools or Teammates
di: Vishwarupe, Varad, et al.
Pubblicazione: (2026)
di: Vishwarupe, Varad, et al.
Pubblicazione: (2026)
When to Retrieve: Teaching LLMs to Utilize Information Retrieval Effectively
di: Labruna, Tiziano, et al.
Pubblicazione: (2024)
di: Labruna, Tiziano, et al.
Pubblicazione: (2024)
AI as Teammate or Tool? A Review of Human-AI Interaction in Decision Support
di: Samu, Most. Sharmin Sultana, et al.
Pubblicazione: (2026)
di: Samu, Most. Sharmin Sultana, et al.
Pubblicazione: (2026)
Transportation networks and competition in the market for corporate control
di: Marco Testoni
Pubblicazione: (2024)
di: Marco Testoni
Pubblicazione: (2024)
Calycanthus chinensis: The Chinese Sweetshrub
di: Li, Jianhua, et al.
Pubblicazione: (2005)
di: Li, Jianhua, et al.
Pubblicazione: (2005)
Verification Limits Code LLM Training
di: Gureja, Srishti, et al.
Pubblicazione: (2025)
di: Gureja, Srishti, et al.
Pubblicazione: (2025)
Are formal and functional linguistic mechanisms dissociated in language models?
di: Hanna, Michael, et al.
Pubblicazione: (2025)
di: Hanna, Michael, et al.
Pubblicazione: (2025)
Have Faith in Faithfulness: Going Beyond Circuit Overlap When Finding Model Mechanisms
di: Hanna, Michael, et al.
Pubblicazione: (2024)
di: Hanna, Michael, et al.
Pubblicazione: (2024)
Who is the richest club in the championship? Detecting and Rewriting Underspecified Questions Improve QA Performance
di: Huang, Yunchong, et al.
Pubblicazione: (2026)
di: Huang, Yunchong, et al.
Pubblicazione: (2026)
Describing Images $\textit{Fast and Slow}$: Quantifying and Predicting the Variation in Human Signals during Visuo-Linguistic Processes
di: Takmaz, Ece, et al.
Pubblicazione: (2024)
di: Takmaz, Ece, et al.
Pubblicazione: (2024)
How Language Models Conflate Logical Validity with Plausibility: A Representational Analysis of Content Effects
di: Bertolazzi, Leonardo, et al.
Pubblicazione: (2025)
di: Bertolazzi, Leonardo, et al.
Pubblicazione: (2025)
Beyond Divergent Creativity: A Human-Based Evaluation of Creativity in Large Language Models
di: Nakajima, Kumiko, et al.
Pubblicazione: (2026)
di: Nakajima, Kumiko, et al.
Pubblicazione: (2026)
Do Pre-Trained Language Models Detect and Understand Semantic Underspecification? Ask the DUST!
di: Wildenburg, Frank, et al.
Pubblicazione: (2024)
di: Wildenburg, Frank, et al.
Pubblicazione: (2024)
The BLA Benchmark: Investigating Basic Language Abilities of Pre-Trained Multimodal Models
di: Chen, Xinyi, et al.
Pubblicazione: (2023)
di: Chen, Xinyi, et al.
Pubblicazione: (2023)
Is my model perplexed for the right reason? Contrasting LLMs' Benchmark Behavior with Token-Level Perplexity
di: Prins, Zoë, et al.
Pubblicazione: (2026)
di: Prins, Zoë, et al.
Pubblicazione: (2026)
RAcQUEt: Unveiling the Dangers of Overlooked Referential Ambiguity in Visual LLMs
di: Testoni, Alberto, et al.
Pubblicazione: (2024)
di: Testoni, Alberto, et al.
Pubblicazione: (2024)
Learning to Ask Informative Questions: Enhancing LLMs with Preference Optimization and Expected Information Gain
di: Mazzaccara, Davide, et al.
Pubblicazione: (2024)
di: Mazzaccara, Davide, et al.
Pubblicazione: (2024)
From Tool to Teammate: LLM Coding Agents as Collaborative Partners for Behavioral Labeling in Educational Dialogue Analysis
di: Chen, Eason, et al.
Pubblicazione: (2026)
di: Chen, Eason, et al.
Pubblicazione: (2026)
A Post-trainer's Guide to Multilingual Training Data: Uncovering Cross-lingual Transfer Dynamics
di: Shimabucoro, Luisa, et al.
Pubblicazione: (2025)
di: Shimabucoro, Luisa, et al.
Pubblicazione: (2025)
Diversify and Conquer: Diversity-Centric Data Selection with Iterative Refinement
di: Yu, Simon, et al.
Pubblicazione: (2024)
di: Yu, Simon, et al.
Pubblicazione: (2024)
Not (yet) the whole story: Evaluating Visual Storytelling Requires More than Measuring Coherence, Grounding, and Repetition
di: Surikuchi, Aditya K, et al.
Pubblicazione: (2024)
di: Surikuchi, Aditya K, et al.
Pubblicazione: (2024)
Natural Language Generation from Visual Events: State-of-the-Art and Key Open Questions
di: Surikuchi, Aditya K, et al.
Pubblicazione: (2025)
di: Surikuchi, Aditya K, et al.
Pubblicazione: (2025)
Where is the multimodal goal post? On the Ability of Foundation Models to Recognize Contextually Important Moments
di: Surikuchi, Aditya K, et al.
Pubblicazione: (2026)
di: Surikuchi, Aditya K, et al.
Pubblicazione: (2026)
To Code, or Not To Code? Exploring Impact of Code in Pre-training
di: Aryabumi, Viraat, et al.
Pubblicazione: (2024)
di: Aryabumi, Viraat, et al.
Pubblicazione: (2024)
Documenti analoghi
-
MemoryPrompt: A Light Wrapper to Improve Context Tracking in Pre-trained Language Models
di: Rakotonirina, Nathanaël Carraz, et al.
Pubblicazione: (2024) -
Naming, Describing, and Quantifying Visual Objects in Humans and LLMs
di: Testoni, Alberto, et al.
Pubblicazione: (2024) -
Evil twins are not that evil: Qualitative insights into machine-generated prompts
di: Rakotonirina, Nathanaël Carraz, et al.
Pubblicazione: (2024) -
Optimizing with Low Budgets: a Comparison on the Black-box Optimization Benchmarking Suite and OpenAI Gym
di: Raponi, Elena, et al.
Pubblicazione: (2023) -
Correct, Concise and Complete: Multi-stage Training For Adaptive Reasoning
di: Rakotonirina, Nathanaël Carraz, et al.
Pubblicazione: (2026)