Data-Prompt Co-Evolution: Growing Test Sets to Refine LLM Behavior
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Minjae, Kahng, Minsuk |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Interactive Prompt Debugging with Sequence Salience
by: Tenney, Ian, et al.
Published: (2024)
by: Tenney, Ian, et al.
Published: (2024)
Who Defines "Best"? Towards Interactive, User-Defined Evaluation of LLM Leaderboards
by: Jung, Minji, et al.
Published: (2026)
by: Jung, Minji, et al.
Published: (2026)
LLM Attributor: Interactive Visual Attribution for LLM Generation
by: Lee, Seongmin, et al.
Published: (2024)
by: Lee, Seongmin, et al.
Published: (2024)
LLM Comparator: Visual Analytics for Side-by-Side Evaluation of Large Language Models
by: Kahng, Minsuk, et al.
Published: (2024)
by: Kahng, Minsuk, et al.
Published: (2024)
VLSlice: Interactive Vision-and-Language Slice Discovery
by: Slyman, Eric, et al.
Published: (2023)
by: Slyman, Eric, et al.
Published: (2023)
The Evolution of LLM Adoption in Industry Data Curation Practices
by: Qian, Crystal, et al.
Published: (2024)
by: Qian, Crystal, et al.
Published: (2024)
Understanding the Dataset Practitioners Behind Large Language Model Development
by: Qian, Crystal, et al.
Published: (2024)
by: Qian, Crystal, et al.
Published: (2024)
From Perception to Decision: Assessing the Role of Chart Types Affordances in High-Level Decision Tasks
by: Li, Yixuan, et al.
Published: (2024)
by: Li, Yixuan, et al.
Published: (2024)
Cross-Lingual Prompt Steerability: Towards Accurate and Robust LLM Behavior across Languages
by: Zhang, Lechen, et al.
Published: (2025)
by: Zhang, Lechen, et al.
Published: (2025)
Assessing Graphical Perception of Image Embedding Models using Channel Effectiveness
by: Lee, Soohyun, et al.
Published: (2024)
by: Lee, Soohyun, et al.
Published: (2024)
An Entropy-Based Test and Development Framework for Uncertainty Modeling in Level-Set Visualizations
by: Sisneros, Robert, et al.
Published: (2024)
by: Sisneros, Robert, et al.
Published: (2024)
Automatic Histograms: Leveraging Language Models for Text Dataset Exploration
by: Reif, Emily, et al.
Published: (2024)
by: Reif, Emily, et al.
Published: (2024)
Data Quality in Crowdsourcing and Spamming Behavior Detection
by: Ba, Yang, et al.
Published: (2024)
by: Ba, Yang, et al.
Published: (2024)
Private Yet Social: How LLM Chatbots Support and Challenge Eating Disorder Recovery
by: Choi, Ryuhaerang, et al.
Published: (2024)
by: Choi, Ryuhaerang, et al.
Published: (2024)
Passive Measurement of Autonomic Arousal in Real-World Settings
by: Abdel-Ghaffar, Samy, et al.
Published: (2025)
by: Abdel-Ghaffar, Samy, et al.
Published: (2025)
Conformal Prediction Sets Improve Human Decision Making
by: Cresswell, Jesse C., et al.
Published: (2024)
by: Cresswell, Jesse C., et al.
Published: (2024)
Behavior Matters: An Alternative Perspective on Promoting Responsible Data Science
by: Dong, Ziwei, et al.
Published: (2024)
by: Dong, Ziwei, et al.
Published: (2024)
Designing for Human-Agent Alignment: Understanding what humans want from their agents
by: Goyal, Nitesh, et al.
Published: (2024)
by: Goyal, Nitesh, et al.
Published: (2024)
Smartwatch-Based Sitting Time Estimation in Real-World Office Settings
by: Zhang, Olivia, et al.
Published: (2026)
by: Zhang, Olivia, et al.
Published: (2026)
Human-LLM Collaborative Feature Engineering for Tabular Data
by: Li, Zhuoyan, et al.
Published: (2026)
by: Li, Zhuoyan, et al.
Published: (2026)
Behavioral Biometrics for Automatic Detection of User Familiarity in VR
by: Zafar, Numan, et al.
Published: (2025)
by: Zafar, Numan, et al.
Published: (2025)
JailbreakHunter: A Visual Analytics Approach for Jailbreak Prompts Discovery from Large-Scale Human-LLM Conversational Datasets
by: Jin, Zhihua, et al.
Published: (2024)
by: Jin, Zhihua, et al.
Published: (2024)
LLM Chatbot-Creation Approaches
by: Mehta, Hemil, et al.
Published: (2025)
by: Mehta, Hemil, et al.
Published: (2025)
Human Delegation Behavior in Human-AI Collaboration: The Effect of Contextual Information
by: Spitzer, Philipp, et al.
Published: (2024)
by: Spitzer, Philipp, et al.
Published: (2024)
Enhancing Adaptive Behavioral Interventions with LLM Inference from Participant-Described States
by: Karine, Karine, et al.
Published: (2025)
by: Karine, Karine, et al.
Published: (2025)
DeforestVis: Behavior Analysis of Machine Learning Models with Surrogate Decision Stumps
by: Chatzimparmpas, Angelos, et al.
Published: (2023)
by: Chatzimparmpas, Angelos, et al.
Published: (2023)
Lightweight Test-Time Adaptation for EMG-Based Gesture Recognition
by: Touko, Nia, et al.
Published: (2026)
by: Touko, Nia, et al.
Published: (2026)
Unobtrusive In-Situ Measurement of Behavior Change by Deep Metric Similarity Learning of Motion Patterns
by: Merz, Christian, et al.
Published: (2025)
by: Merz, Christian, et al.
Published: (2025)
Time2Lang: Bridging Time-Series Foundation Models and Large Language Models for Health Sensing Beyond Prompting
by: Pillai, Arvind, et al.
Published: (2025)
by: Pillai, Arvind, et al.
Published: (2025)
T-TIME: Test-Time Information Maximization Ensemble for Plug-and-Play BCIs
by: Li, Siyang, et al.
Published: (2024)
by: Li, Siyang, et al.
Published: (2024)
CoGen: Creation of Reusable UI Components in Figma via Textual Commands
by: Kanapathipillai, Ishani, et al.
Published: (2026)
by: Kanapathipillai, Ishani, et al.
Published: (2026)
Can LLM Assist in the Evaluation of the Quality of Machine Learning Explanations?
by: Wang, Bo, et al.
Published: (2025)
by: Wang, Bo, et al.
Published: (2025)
Evaluation of LLM-based Explanations for a Learning Analytics Dashboard
by: Deriyeva, Alina, et al.
Published: (2025)
by: Deriyeva, Alina, et al.
Published: (2025)
Defining Effective Engagement For Enhancing Cancer Patients' Well-being with Mobile Digital Behavior Change Interventions
by: Lisowska, Aneta, et al.
Published: (2024)
by: Lisowska, Aneta, et al.
Published: (2024)
From Commands to Prompts: LLM-based Semantic File System for AIOS
by: Shi, Zeru, et al.
Published: (2024)
by: Shi, Zeru, et al.
Published: (2024)
CataractBot: An LLM-Powered Expert-in-the-Loop Chatbot for Cataract Patients
by: Ramjee, Pragnya, et al.
Published: (2024)
by: Ramjee, Pragnya, et al.
Published: (2024)
ChatISA: A Prompt-Engineered, In-House Multi-Modal Generative AI Chatbot for Information Systems Education
by: Megahed, Fadel M., et al.
Published: (2024)
by: Megahed, Fadel M., et al.
Published: (2024)
IMUCoCo: Enabling Flexible On-Body IMU Placement for Human Pose Estimation and Activity Recognition
by: Zhou, Haozhe, et al.
Published: (2025)
by: Zhou, Haozhe, et al.
Published: (2025)
d-DQIVAR: Data-centric Visual Analytics and Reasoning for Data Quality Improvement
by: Hong, Hyein, et al.
Published: (2025)
by: Hong, Hyein, et al.
Published: (2025)
Streaming, Fast and Slow: Cognitive Load-Aware Streaming for Efficient LLM Serving
by: Xiao, Chang, et al.
Published: (2025)
by: Xiao, Chang, et al.
Published: (2025)
Similar Items
-
Interactive Prompt Debugging with Sequence Salience
by: Tenney, Ian, et al.
Published: (2024) -
Who Defines "Best"? Towards Interactive, User-Defined Evaluation of LLM Leaderboards
by: Jung, Minji, et al.
Published: (2026) -
LLM Attributor: Interactive Visual Attribution for LLM Generation
by: Lee, Seongmin, et al.
Published: (2024) -
LLM Comparator: Visual Analytics for Side-by-Side Evaluation of Large Language Models
by: Kahng, Minsuk, et al.
Published: (2024) -
VLSlice: Interactive Vision-and-Language Slice Discovery
by: Slyman, Eric, et al.
Published: (2023)