Investigating Cost-Efficiency of LLM-Generated Training Data for Conversational Semantic Frame Analysis
Fuente:
arXiv
Saved in:
| Main Authors: | Matta, Shiho, Huang, Yin Jou, Cheng, Fei, Kiyomaru, Hirokazu, Murawaki, Yugo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can We Trust LLM Detectors?
by: Sandhan, Jivnesh, et al.
Published: (2026)
by: Sandhan, Jivnesh, et al.
Published: (2026)
Language Lives in Sparse Dimensions: Toward Interpretable and Efficient Multilingual Control for Large Language Models
by: Zhong, Chengzhi, et al.
Published: (2025)
by: Zhong, Chengzhi, et al.
Published: (2025)
RecMind: Japanese Movie Recommendation Dialogue with Seeker's Internal State
by: Kodama, Takashi, et al.
Published: (2024)
by: Kodama, Takashi, et al.
Published: (2024)
Principal Component Analysis as a Sanity Check for Bayesian Phylolinguistic Reconstruction
by: Murawaki, Yugo
Published: (2024)
by: Murawaki, Yugo
Published: (2024)
How Does Cognitive Bias Affect Large Language Models? A Case Study on the Anchoring Effect in Price Negotiation Simulations
by: Takenami, Yoshiki, et al.
Published: (2025)
by: Takenami, Yoshiki, et al.
Published: (2025)
Beyond English-Centric LLMs: What Language Do Multilingual Language Models Think in?
by: Zhong, Chengzhi, et al.
Published: (2024)
by: Zhong, Chengzhi, et al.
Published: (2024)
Beyond Self-Reports: Multi-Observer Agents for Personality Assessment in Large Language Models
by: Huang, Yin Jou, et al.
Published: (2025)
by: Huang, Yin Jou, et al.
Published: (2025)
How Personality Traits Influence Negotiation Outcomes? A Simulation based on Large Language Models
by: Huang, Yin Jou, et al.
Published: (2024)
by: Huang, Yin Jou, et al.
Published: (2024)
Towards Transparency: Exploring LLM Trainings Datasets through Visual Topic Modeling and Semantic Frame
by: de Dampierre, Charles, et al.
Published: (2024)
by: de Dampierre, Charles, et al.
Published: (2024)
Addressing Tokenization Inconsistency in Steganography and Watermarking Based on Large Language Models
by: Yan, Ruiyi, et al.
Published: (2025)
by: Yan, Ruiyi, et al.
Published: (2025)
CAPE: Context-Aware Personality Evaluation Framework for Large Language Models
by: Sandhan, Jivnesh, et al.
Published: (2025)
by: Sandhan, Jivnesh, et al.
Published: (2025)
Persona Jailbreaking in Large Language Models
by: Sandhan, Jivnesh, et al.
Published: (2026)
by: Sandhan, Jivnesh, et al.
Published: (2026)
Efficient Provably Secure Linguistic Steganography via Range Coding
by: Yan, Ruiyi, et al.
Published: (2026)
by: Yan, Ruiyi, et al.
Published: (2026)
DataFrame QA: A Universal LLM Framework on DataFrame Question Answering Without Data Exposure
by: Ye, Junyi, et al.
Published: (2024)
by: Ye, Junyi, et al.
Published: (2024)
Crown, Frame, Reverse: Layer-Wise Scaling Variants for LLM Pre-Training
by: Baroian, Andrei, et al.
Published: (2025)
by: Baroian, Andrei, et al.
Published: (2025)
Text Detoxification: Data Efficiency, Semantic Preservation and Model Generalization
by: Yu, Jing, et al.
Published: (2025)
by: Yu, Jing, et al.
Published: (2025)
Improving Training Efficiency and Reducing Maintenance Costs via Language Specific Model Merging
by: Dmonte, Alphaeus, et al.
Published: (2026)
by: Dmonte, Alphaeus, et al.
Published: (2026)
FrameNet Semantic Role Classification by Analogy
by: Ngo, Van-Duy, et al.
Published: (2026)
by: Ngo, Van-Duy, et al.
Published: (2026)
Anchored Sliding Window: Toward Robust and Imperceptible Linguistic Steganography
by: Yan, Ruiyi, et al.
Published: (2026)
by: Yan, Ruiyi, et al.
Published: (2026)
CostBench: Evaluating Multi-Turn Cost-Optimal Planning and Adaptation in Dynamic Environments for LLM Tool-Use Agents
by: Liu, Jiayu, et al.
Published: (2025)
by: Liu, Jiayu, et al.
Published: (2025)
Automated Question Generation on Tabular Data for Conversational Data Exploration
by: Chaudhuri, Ritwik, et al.
Published: (2024)
by: Chaudhuri, Ritwik, et al.
Published: (2024)
Transformer Enhanced Relation Classification: A Comparative Analysis of Contextuality, Data Efficiency and Sequence Complexity
by: Jing, Bowen, et al.
Published: (2025)
by: Jing, Bowen, et al.
Published: (2025)
An LLM-Based Approach for Insight Generation in Data Analysis
by: Pérez, Alberto Sánchez, et al.
Published: (2025)
by: Pérez, Alberto Sánchez, et al.
Published: (2025)
Next-Token Prediction Task Assumes Optimal Data Ordering for LLM Training in Proof Generation
by: An, Chenyang, et al.
Published: (2024)
by: An, Chenyang, et al.
Published: (2024)
Improving Data Efficiency via Curating LLM-Driven Rating Systems
by: Pang, Jinlong, et al.
Published: (2024)
by: Pang, Jinlong, et al.
Published: (2024)
Beyond Public Access in LLM Pre-Training Data
by: Rosenblat, Sruly, et al.
Published: (2025)
by: Rosenblat, Sruly, et al.
Published: (2025)
MindMerger: Efficient Boosting LLM Reasoning in non-English Languages
by: Huang, Zixian, et al.
Published: (2024)
by: Huang, Zixian, et al.
Published: (2024)
Unifying Structured Data as Graph for Data-to-Text Pre-Training
by: Li, Shujie, et al.
Published: (2024)
by: Li, Shujie, et al.
Published: (2024)
LLM-Agnostic Semantic Representation Attack
by: Lian, Jiawei, et al.
Published: (2026)
by: Lian, Jiawei, et al.
Published: (2026)
Modeling Unified Semantic Discourse Structure for High-quality Headline Generation
by: Xu, Minghui, et al.
Published: (2024)
by: Xu, Minghui, et al.
Published: (2024)
MedAlpaca -- An Open-Source Collection of Medical Conversational AI Models and Training Data
by: Han, Tianyu, et al.
Published: (2023)
by: Han, Tianyu, et al.
Published: (2023)
An Investigation into Value Misalignment in LLM-Generated Texts for Cultural Heritage
by: Bu, Fan, et al.
Published: (2025)
by: Bu, Fan, et al.
Published: (2025)
When Wording Steers the Evaluation: Framing Bias in LLM judges
by: Hwang, Yerin, et al.
Published: (2026)
by: Hwang, Yerin, et al.
Published: (2026)
SlimPajama-DC: Understanding Data Combinations for LLM Training
by: Shen, Zhiqiang, et al.
Published: (2023)
by: Shen, Zhiqiang, et al.
Published: (2023)
Does Safety Training of LLMs Generalize to Semantically Related Natural Prompts?
by: Addepalli, Sravanti, et al.
Published: (2024)
by: Addepalli, Sravanti, et al.
Published: (2024)
Multi-News+: Cost-efficient Dataset Cleansing via LLM-based Data Annotation
by: Choi, Juhwan, et al.
Published: (2024)
by: Choi, Juhwan, et al.
Published: (2024)
Optima: Optimizing Effectiveness and Efficiency for LLM-Based Multi-Agent System
by: Chen, Weize, et al.
Published: (2024)
by: Chen, Weize, et al.
Published: (2024)
HDLCoRe: A Training-Free Framework for Mitigating Hallucinations in LLM-Generated HDL
by: Ping, Heng, et al.
Published: (2025)
by: Ping, Heng, et al.
Published: (2025)
Generating High Quality Synthetic Data for Dutch Medical Conversations
by: Kuan, Cecilia, et al.
Published: (2026)
by: Kuan, Cecilia, et al.
Published: (2026)
Improving Synthetic Data Training for Contextual Biasing Models with a Keyword-Aware Cost Function
by: Kwok, Chin Yuen, et al.
Published: (2025)
by: Kwok, Chin Yuen, et al.
Published: (2025)
Similar Items
-
Can We Trust LLM Detectors?
by: Sandhan, Jivnesh, et al.
Published: (2026) -
Language Lives in Sparse Dimensions: Toward Interpretable and Efficient Multilingual Control for Large Language Models
by: Zhong, Chengzhi, et al.
Published: (2025) -
RecMind: Japanese Movie Recommendation Dialogue with Seeker's Internal State
by: Kodama, Takashi, et al.
Published: (2024) -
Principal Component Analysis as a Sanity Check for Bayesian Phylolinguistic Reconstruction
by: Murawaki, Yugo
Published: (2024) -
How Does Cognitive Bias Affect Large Language Models? A Case Study on the Anchoring Effect in Price Negotiation Simulations
by: Takenami, Yoshiki, et al.
Published: (2025)