Talking with Oompa Loompas: A novel framework for evaluating linguistic acquisition of LLM agents
Fuente:
arXiv
Saved in:
| Main Authors: | Swain, Sankalp Tattwadarshi, Krishnatray, Anshika, Kumar, Dhruv, Challa, Jagat Sesh |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SAC: A Framework for Measuring and Inducing Personality Traits in LLMs with Dynamic Intensity Control
by: Chittem, Adithya, et al.
Published: (2025)
by: Chittem, Adithya, et al.
Published: (2025)
"It's not like Jarvis, but it's pretty close!" -- Examining ChatGPT's Usage among Undergraduate Students in Computer Science
by: Joshi, Ishika, et al.
Published: (2023)
by: Joshi, Ishika, et al.
Published: (2023)
The Impact of Large Language Models on K-12 Education in Rural India: A Thematic Analysis of Student Volunteer's Perspectives
by: Goyal, Harshita, et al.
Published: (2025)
by: Goyal, Harshita, et al.
Published: (2025)
Sakshm AI: Advancing AI-Assisted Coding Education for Engineering Students in India Through Socratic Tutoring and Comprehensive Feedback
by: Gupta, Raj, et al.
Published: (2025)
by: Gupta, Raj, et al.
Published: (2025)
CBEval: A framework for evaluating and interpreting cognitive biases in LLMs
by: Shaikh, Ammar, et al.
Published: (2024)
by: Shaikh, Ammar, et al.
Published: (2024)
Talk Less, Call Right: Enhancing Role-Play LLM Agents with Automatic Prompt Optimization and Role Prompting
by: Ruangtanusak, Saksorn, et al.
Published: (2025)
by: Ruangtanusak, Saksorn, et al.
Published: (2025)
Talking to Machines: do you read me?
by: Rojas-Barahona, Lina M.
Published: (2024)
by: Rojas-Barahona, Lina M.
Published: (2024)
Generation Z's Ability to Discriminate Between AI-generated and Human-Authored Text on Discord
by: Ramu, Dhruv, et al.
Published: (2023)
by: Ramu, Dhruv, et al.
Published: (2023)
MimiTalk: Revolutionizing Qualitative Research with Dual-Agent AI
by: Liu, Fengming, et al.
Published: (2025)
by: Liu, Fengming, et al.
Published: (2025)
The Art of Audience Engagement: LLM-Based Thin-Slicing of Scientific Talks
by: Schmälzle, Ralf, et al.
Published: (2025)
by: Schmälzle, Ralf, et al.
Published: (2025)
AI on My Shoulder: Supporting Emotional Labor in Front-Office Roles with an LLM-based Empathetic Coworker
by: Swain, Vedant Das, et al.
Published: (2024)
by: Swain, Vedant Das, et al.
Published: (2024)
PsychBench: A comprehensive and professional benchmark for evaluating the performance of LLM-assisted psychiatric clinical practice
by: Liu, Shuyu, et al.
Published: (2025)
by: Liu, Shuyu, et al.
Published: (2025)
Detecting Deceptive Dark Patterns in E-commerce Platforms
by: Ramteke, Arya, et al.
Published: (2024)
by: Ramteke, Arya, et al.
Published: (2024)
Do We Talk to Robots Like Therapists, and Do They Respond Accordingly? Language Alignment in AI Emotional Support
by: Chiang, Sophie, et al.
Published: (2025)
by: Chiang, Sophie, et al.
Published: (2025)
Writing as a testbed for open ended agents
by: Gooding, Sian, et al.
Published: (2025)
by: Gooding, Sian, et al.
Published: (2025)
PhysicsSolutionAgent: Towards Multimodal Explanations for Numerical Physics Problem Solving
by: Thole, Aditya, et al.
Published: (2026)
by: Thole, Aditya, et al.
Published: (2026)
TUX: Measuring Human--AI Tacit Understanding
by: Li, Yueshen, et al.
Published: (2026)
by: Li, Yueshen, et al.
Published: (2026)
Talking to Robots: A Practical Examination of Speech Foundation Models for HRI Applications
by: Rosin, Theresa Pekarek, et al.
Published: (2025)
by: Rosin, Theresa Pekarek, et al.
Published: (2025)
AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments
by: Schmidgall, Samuel, et al.
Published: (2024)
by: Schmidgall, Samuel, et al.
Published: (2024)
PhDGPT: Introducing a psychometric and linguistic dataset about how large language models perceive graduate students and professors in psychology
by: De Duro, Edoardo Sebastiano, et al.
Published: (2024)
by: De Duro, Edoardo Sebastiano, et al.
Published: (2024)
Multi-agent KTO: Reinforcing Strategic Interactions of Large Language Model in Language Game
by: Ye, Rong, et al.
Published: (2025)
by: Ye, Rong, et al.
Published: (2025)
User Willingness-aware Sales Talk Dataset
by: Hentona, Asahi, et al.
Published: (2024)
by: Hentona, Asahi, et al.
Published: (2024)
Automated stereotactic radiosurgery planning using a human-in-the-loop reasoning large language model agent
by: Nusrat, Humza, et al.
Published: (2025)
by: Nusrat, Humza, et al.
Published: (2025)
User Perceptions vs. Proxy LLM Judges: Privacy and Helpfulness in LLM Responses to Privacy-Sensitive Scenarios
by: Wu, Xiaoyuan, et al.
Published: (2025)
by: Wu, Xiaoyuan, et al.
Published: (2025)
Benchmarking LLM Tool-Use in the Wild
by: Yu, Peijie, et al.
Published: (2026)
by: Yu, Peijie, et al.
Published: (2026)
Game Development as Human-LLM Interaction
by: Hong, Jiale, et al.
Published: (2024)
by: Hong, Jiale, et al.
Published: (2024)
Grading Scale Impact on LLM-as-a-Judge: Human-LLM Alignment Is Highest on 0-5 Grading Scale
by: Li, Weiyue, et al.
Published: (2026)
by: Li, Weiyue, et al.
Published: (2026)
Mediating Modes of Thought: LLM's for design scripting
by: Rietschel, Moritz, et al.
Published: (2024)
by: Rietschel, Moritz, et al.
Published: (2024)
BADGE: BADminton report Generation and Evaluation with LLM
by: Chiang, Shang-Hsuan, et al.
Published: (2024)
by: Chiang, Shang-Hsuan, et al.
Published: (2024)
Enhancing Talk Moves Analysis in Mathematics Tutoring through Classroom Teaching Discourse
by: Cao, Jie, et al.
Published: (2024)
by: Cao, Jie, et al.
Published: (2024)
Overhearing LLM Agents: A Survey, Taxonomy, and Roadmap
by: Zhu, Andrew, et al.
Published: (2025)
by: Zhu, Andrew, et al.
Published: (2025)
Effects of Varying LLM Access on Essay Writing Behavior
by: Christenson, Julia, et al.
Published: (2026)
by: Christenson, Julia, et al.
Published: (2026)
Can Unconfident LLM Annotations Be Used for Confident Conclusions?
by: Gligorić, Kristina, et al.
Published: (2024)
by: Gligorić, Kristina, et al.
Published: (2024)
Creating and Evaluating Personas Using Generative AI: A Scoping Review of 81 Articles
by: Amin, Danial, et al.
Published: (2025)
by: Amin, Danial, et al.
Published: (2025)
Creativity in LLM-based Multi-Agent Systems: A Survey
by: Lin, Yi-Cheng, et al.
Published: (2025)
by: Lin, Yi-Cheng, et al.
Published: (2025)
Understanding Learner-LLM Chatbot Interactions and the Impact of Prompting Guidelines
by: Koyuturk, Cansu, et al.
Published: (2025)
by: Koyuturk, Cansu, et al.
Published: (2025)
Designing LLM Chains by Adapting Techniques from Crowdsourcing Workflows
by: Grunde-McLaughlin, Madeleine, et al.
Published: (2023)
by: Grunde-McLaughlin, Madeleine, et al.
Published: (2023)
Patchview: LLM-Powered Worldbuilding with Generative Dust and Magnet Visualization
by: Chung, John Joon Young, et al.
Published: (2024)
by: Chung, John Joon Young, et al.
Published: (2024)
Does AI Coaching Prepare us for Workplace Negotiations?
by: Duddu, Veda, et al.
Published: (2025)
by: Duddu, Veda, et al.
Published: (2025)
Not My Truce: Personality Differences in AI-Mediated Workplace Negotiation
by: Duddu, Veda, et al.
Published: (2026)
by: Duddu, Veda, et al.
Published: (2026)
Similar Items
-
SAC: A Framework for Measuring and Inducing Personality Traits in LLMs with Dynamic Intensity Control
by: Chittem, Adithya, et al.
Published: (2025) -
"It's not like Jarvis, but it's pretty close!" -- Examining ChatGPT's Usage among Undergraduate Students in Computer Science
by: Joshi, Ishika, et al.
Published: (2023) -
The Impact of Large Language Models on K-12 Education in Rural India: A Thematic Analysis of Student Volunteer's Perspectives
by: Goyal, Harshita, et al.
Published: (2025) -
Sakshm AI: Advancing AI-Assisted Coding Education for Engineering Students in India Through Socratic Tutoring and Comprehensive Feedback
by: Gupta, Raj, et al.
Published: (2025) -
CBEval: A framework for evaluating and interpreting cognitive biases in LLMs
by: Shaikh, Ammar, et al.
Published: (2024)