TRACE: A taxonomy-grounded synthetic dataset for teaching-program generation and session interpretation in Applied Behavior Analysis
Fuente:
arXiv
Guardado en:
| Autor principal: | Kahunla, Festus |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Automated Bug Triaging using Instruction-Tuned Large Language Models
por: Kiashemshaki, Kiana, et al.
Publicado: (2025)
por: Kiashemshaki, Kiana, et al.
Publicado: (2025)
SiliconMind-V1: Multi-Agent Distillation and Debug-Reasoning Workflows for Verilog Code Generation
por: Chen, Mu-Chi, et al.
Publicado: (2026)
por: Chen, Mu-Chi, et al.
Publicado: (2026)
CIDR: A Large-Scale Industrial Source Code Dataset for Software Engineering Research
por: Savenkov, Vladislav
Publicado: (2026)
por: Savenkov, Vladislav
Publicado: (2026)
ContractBench: Can LLM Agents Preserve Observation Contracts?
por: Wang, Jicheng, et al.
Publicado: (2026)
por: Wang, Jicheng, et al.
Publicado: (2026)
PARNESS: A Paper Harness for End-to-End Automated Scientific Research with Dynamic Workflows, Full-Text Indexing, and Cross-Run Knowledge Accumulation
por: Wang, Yuchen, et al.
Publicado: (2026)
por: Wang, Yuchen, et al.
Publicado: (2026)
AgentAtlas: Beyond Outcome Leaderboards for LLM Agents
por: Mazaheri, Parsa, et al.
Publicado: (2026)
por: Mazaheri, Parsa, et al.
Publicado: (2026)
Software Defined Vehicle Code Generation: A Few-Shot Prompting Approach
por: Nguyen, Quang-Dung, et al.
Publicado: (2025)
por: Nguyen, Quang-Dung, et al.
Publicado: (2025)
The Single-File Test: A Longitudinal Public-Interface Evaluation of First-Output LLM Web Generation with Social Reach Tracking
por: Palacios, Diego Cabezas
Publicado: (2026)
por: Palacios, Diego Cabezas
Publicado: (2026)
AuditRepairBench: A Paired-Execution Trace Corpus for Evaluator-Channel Ranking Instability in Agent Repair
por: Hu, Yuelin, et al.
Publicado: (2026)
por: Hu, Yuelin, et al.
Publicado: (2026)
AgentPulse: A Continuous Multi-Signal Framework for Evaluating AI Agents in Deployment
por: Gao, Yuxuan, et al.
Publicado: (2026)
por: Gao, Yuxuan, et al.
Publicado: (2026)
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
por: Fadli, Samih
Publicado: (2025)
por: Fadli, Samih
Publicado: (2025)
Learning the meanings of function words from grounded language using a visual question answering model
por: Portelance, Eva, et al.
Publicado: (2023)
por: Portelance, Eva, et al.
Publicado: (2023)
LLMCup: Ranking-Enhanced Comment Updating with LLMs
por: Ge, Hua, et al.
Publicado: (2025)
por: Ge, Hua, et al.
Publicado: (2025)
5C Prompt Contracts: A Minimalist, Creative-Friendly, Token-Efficient Design Framework for Individual and SME LLM Usage
por: Ari, Ugur
Publicado: (2025)
por: Ari, Ugur
Publicado: (2025)
GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning
por: Agrawal, Lakshya A, et al.
Publicado: (2025)
por: Agrawal, Lakshya A, et al.
Publicado: (2025)
Predictive Analytics for Collaborators Answers, Code Quality, and Dropout on Stack Overflow
por: Zolduoarrati, Elijah, et al.
Publicado: (2025)
por: Zolduoarrati, Elijah, et al.
Publicado: (2025)
Layer-Aware Embedding Fusion for LLMs in Text Classifications
por: Gwak, Jiho, et al.
Publicado: (2025)
por: Gwak, Jiho, et al.
Publicado: (2025)
KSHSeek: Data-Driven Approaches to Mitigating and Detecting Knowledge-Shortcut Hallucinations in Generative Models
por: Liu, Zhongxin, et al.
Publicado: (2025)
por: Liu, Zhongxin, et al.
Publicado: (2025)
On the Influence of Discourse Relations in Persuasive Texts
por: Turk, Nawar, et al.
Publicado: (2025)
por: Turk, Nawar, et al.
Publicado: (2025)
Calibrated Confidence Estimation for Tabular Question Answering
por: Voss, Lukas
Publicado: (2026)
por: Voss, Lukas
Publicado: (2026)
Towards Alignment-Centric Paradigm: A Survey of Instruction Tuning in Large Language Models
por: Han, Xudong, et al.
Publicado: (2025)
por: Han, Xudong, et al.
Publicado: (2025)
Emergent Lexical Semantics in Neural Language Models: Testing Martin's Law on LLM-Generated Text
por: Kugler, Kai
Publicado: (2025)
por: Kugler, Kai
Publicado: (2025)
Character-Level Transformer for Tajik-Persian Transliteration with a Parallel Lexical Corpus
por: Arabov, Mullosharaf K.
Publicado: (2026)
por: Arabov, Mullosharaf K.
Publicado: (2026)
Mitigating Cross-Lingual Cultural Inconsistencies in LLMs via Consensus-Driven Preference Optimisation
por: Resck, Lucas, et al.
Publicado: (2026)
por: Resck, Lucas, et al.
Publicado: (2026)
Bridging the Gap: An Intermediate Language for Enhanced and Cost-Effective Grapheme-to-Phoneme Conversion with Homographs with Multiple Pronunciations Disambiguation
por: Bertina, Abbas, et al.
Publicado: (2025)
por: Bertina, Abbas, et al.
Publicado: (2025)
Induce, Align, Predict: Zero-Shot Stance Detection via Cognitive Inductive Reasoning
por: Zhang, Bowen, et al.
Publicado: (2025)
por: Zhang, Bowen, et al.
Publicado: (2025)
EmoLoom-2B: Fast Base-Model Screening for Emotion Classification and VAD with Lexicon-Weak Supervision and KV-Off Evaluation
por: Li, Zilin, et al.
Publicado: (2026)
por: Li, Zilin, et al.
Publicado: (2026)
TRiMS: Real-Time Tracking of Minimal Sufficient Length for Efficient Reasoning via RL
por: Bian, Tingcheng, et al.
Publicado: (2026)
por: Bian, Tingcheng, et al.
Publicado: (2026)
MCP: A Control-Theoretic Orchestration Framework for Synergistic Efficiency and Interpretability in Multimodal Large Language Models
por: Zhang, Luyan
Publicado: (2025)
por: Zhang, Luyan
Publicado: (2025)
Mixup Model Merge: Enhancing Model Merging Performance through Randomized Linear Interpolation
por: Zhou, Yue, et al.
Publicado: (2025)
por: Zhou, Yue, et al.
Publicado: (2025)
D-COT: Disciplined Chain-of-Thought Learning for Efficient Reasoning in Small Language Models
por: Ubukata, Shunsuke
Publicado: (2026)
por: Ubukata, Shunsuke
Publicado: (2026)
MOCHA: Multi-Objective Chebyshev Annealing for Agent Skill Optimization
por: Tanjim, Md Mehrab, et al.
Publicado: (2026)
por: Tanjim, Md Mehrab, et al.
Publicado: (2026)
Tool-Genesis: A Task-Driven Tool Creation Benchmark for Self-Evolving Language Agent
por: Xia, Bowei, et al.
Publicado: (2026)
por: Xia, Bowei, et al.
Publicado: (2026)
Feedback-Normalized Developer Memory for Reinforcement-Learning Coding Agents: A Safety-Gated MCP Architecture
por: Iscan, Mehmet
Publicado: (2026)
por: Iscan, Mehmet
Publicado: (2026)
Natural Language Summarization Enables Multi-Repository Bug Localization by LLMs in Microservice Architectures
por: Oskooei, Amirkia Rafiei, et al.
Publicado: (2025)
por: Oskooei, Amirkia Rafiei, et al.
Publicado: (2025)
ELMTEX: Fine-Tuning Large Language Models for Structured Clinical Information Extraction. A Case Study on Clinical Reports
por: Guluzade, Aynur, et al.
Publicado: (2025)
por: Guluzade, Aynur, et al.
Publicado: (2025)
Repetition Without Exclusivity: Scale Sensitivity of Referential Mechanisms in Child-Scale Language Models
por: Cacioli, Jon-Paul
Publicado: (2026)
por: Cacioli, Jon-Paul
Publicado: (2026)
Evaluating Large Language Models for IUCN Red List Species Information
por: Uryu, Shinya
Publicado: (2025)
por: Uryu, Shinya
Publicado: (2025)
Shallow Robustness, Deep Vulnerabilities: Multi-Turn Evaluation of Medical LLMs
por: Manczak, Blazej, et al.
Publicado: (2025)
por: Manczak, Blazej, et al.
Publicado: (2025)
Categorical Perception in Large Language Model Hidden States: Structural Warping at Digit-Count Boundaries
por: Cacioli, Jon-Paul
Publicado: (2026)
por: Cacioli, Jon-Paul
Publicado: (2026)
Ejemplares similares
-
Automated Bug Triaging using Instruction-Tuned Large Language Models
por: Kiashemshaki, Kiana, et al.
Publicado: (2025) -
SiliconMind-V1: Multi-Agent Distillation and Debug-Reasoning Workflows for Verilog Code Generation
por: Chen, Mu-Chi, et al.
Publicado: (2026) -
CIDR: A Large-Scale Industrial Source Code Dataset for Software Engineering Research
por: Savenkov, Vladislav
Publicado: (2026) -
ContractBench: Can LLM Agents Preserve Observation Contracts?
por: Wang, Jicheng, et al.
Publicado: (2026) -
PARNESS: A Paper Harness for End-to-End Automated Scientific Research with Dynamic Workflows, Full-Text Indexing, and Cross-Run Knowledge Accumulation
por: Wang, Yuchen, et al.
Publicado: (2026)