Next Reply Prediction X Dataset: Linguistic Discrepancies in Naively Generated Content
Fuente:
arXiv
Guardado en:
| Autores principales: | Münker, Simon, Schwager, Nils, Kugler, Kai, Heseltine, Michael, Rettinger, Achim |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Zero-shot prompt-based classification: topic labeling in times of foundation models in German Tweets
por: Münker, Simon, et al.
Publicado: (2024)
por: Münker, Simon, et al.
Publicado: (2024)
Towards Simulating Social Media Users with LLMs: Evaluating the Operational Validity of Conditioned Comment Prediction
por: Schwager, Nils, et al.
Publicado: (2026)
por: Schwager, Nils, et al.
Publicado: (2026)
Don't Trust Generative Agents to Mimic Communication on Social Networks Unless You Benchmarked their Empirical Realism
por: Münker, Simon, et al.
Publicado: (2025)
por: Münker, Simon, et al.
Publicado: (2025)
InvBERT: Reconstructing Text from Contextualized Word Embeddings by inverting the BERT pipeline
por: Kugler, Kai, et al.
Publicado: (2021)
por: Kugler, Kai, et al.
Publicado: (2021)
Cultural Bias in Large Language Models: Evaluating AI Agents through Moral Questionnaires
por: Münker, Simon
Publicado: (2025)
por: Münker, Simon
Publicado: (2025)
Political Bias in LLMs: Unaligned Moral Values in Agent-centric Simulations
por: Münker, Simon
Publicado: (2024)
por: Münker, Simon
Publicado: (2024)
Agent-Based Simulations of Online Political Discussions: A Case Study on Elections in Germany
por: Sittar, Abdul, et al.
Publicado: (2025)
por: Sittar, Abdul, et al.
Publicado: (2025)
LLM Agents Predict Social Media Reactions but Do Not Outperform Text Classifiers: Benchmarking Simulation Accuracy Using 120K+ Personas of 1511 Humans
por: Bojic, Ljubisa, et al.
Publicado: (2026)
por: Bojic, Ljubisa, et al.
Publicado: (2026)
Analysis and Visualization of Linguistic Structures in Large Language Models: Neural Representations of Verb-Particle Constructions in BERT
por: Kissane, Hassane, et al.
Publicado: (2024)
por: Kissane, Hassane, et al.
Publicado: (2024)
Incentivizing News Consumption on Social Media Platforms Using Large Language Models and Realistic Bot Accounts
por: Askari, Hadi, et al.
Publicado: (2024)
por: Askari, Hadi, et al.
Publicado: (2024)
AdParaphrase: Paraphrase Dataset for Analyzing Linguistic Features toward Generating Attractive Ad Texts
por: Murakami, Soichiro, et al.
Publicado: (2025)
por: Murakami, Soichiro, et al.
Publicado: (2025)
Alternatives To Next Token Prediction In Text Generation -- A Survey
por: Wyatt, Charlie, et al.
Publicado: (2025)
por: Wyatt, Charlie, et al.
Publicado: (2025)
Linguistic Patterns in Pandemic-Related Content: A Comparative Analysis of COVID-19, Constraint, and Monkeypox Datasets
por: Sikosana, Mkululi, et al.
Publicado: (2025)
por: Sikosana, Mkululi, et al.
Publicado: (2025)
Enhancing AI-Driven Education: Integrating Cognitive Frameworks, Linguistic Feedback Analysis, and Ethical Considerations for Improved Content Generation
por: Yaacoub, Antoun, et al.
Publicado: (2025)
por: Yaacoub, Antoun, et al.
Publicado: (2025)
Limited Linguistic Diversity in Embodied AI Datasets
por: Wanna, Selma, et al.
Publicado: (2026)
por: Wanna, Selma, et al.
Publicado: (2026)
Palm: A Culturally Inclusive and Linguistically Diverse Dataset for Arabic LLMs
por: Alwajih, Fakhraddin, et al.
Publicado: (2025)
por: Alwajih, Fakhraddin, et al.
Publicado: (2025)
BLUCK: A Benchmark Dataset for Bengali Linguistic Understanding and Cultural Knowledge
por: Kabir, Daeen, et al.
Publicado: (2025)
por: Kabir, Daeen, et al.
Publicado: (2025)
Evaluating Reward Model Generalization via Pairwise Maximum Discrepancy Competitions
por: Luo, Shunyang, et al.
Publicado: (2026)
por: Luo, Shunyang, et al.
Publicado: (2026)
Moving Beyond Next-Token Prediction: Transformers are Context-Sensitive Language Generators
por: Rhee, Phill Kyu
Publicado: (2025)
por: Rhee, Phill Kyu
Publicado: (2025)
Next Concept Prediction in Discrete Latent Space Leads to Stronger Language Models
por: Liu, Yuliang, et al.
Publicado: (2026)
por: Liu, Yuliang, et al.
Publicado: (2026)
Training LLMs Beyond Next Token Prediction -- Filling the Mutual Information Gap
por: Yang, Chun-Hao, et al.
Publicado: (2025)
por: Yang, Chun-Hao, et al.
Publicado: (2025)
Convergent Representations of Linguistic Constructions in Human and Artificial Neural Systems
por: Ramezani, Pegah, et al.
Publicado: (2026)
por: Ramezani, Pegah, et al.
Publicado: (2026)
A Linguistics-Aware LLM Watermarking via Syntactic Predictability
por: Park, Shinwoo, et al.
Publicado: (2025)
por: Park, Shinwoo, et al.
Publicado: (2025)
Linguistically Communicating Uncertainty in Patient-Facing Risk Prediction Models
por: Sivaprasad, Adarsa, et al.
Publicado: (2024)
por: Sivaprasad, Adarsa, et al.
Publicado: (2024)
Next-Token Prediction Task Assumes Optimal Data Ordering for LLM Training in Proof Generation
por: An, Chenyang, et al.
Publicado: (2024)
por: An, Chenyang, et al.
Publicado: (2024)
Genie: Achieving Human Parity in Content-Grounded Datasets Generation
por: Yehudai, Asaf, et al.
Publicado: (2024)
por: Yehudai, Asaf, et al.
Publicado: (2024)
Linguistic Characteristics of AI-Generated Text: A Survey
por: Terčon, Luka, et al.
Publicado: (2025)
por: Terčon, Luka, et al.
Publicado: (2025)
Ace-CEFR -- A Dataset for Automated Evaluation of the Linguistic Difficulty of Conversational Texts for LLM Applications
por: Kogan, David, et al.
Publicado: (2025)
por: Kogan, David, et al.
Publicado: (2025)
A comprehensive study of on-device NLP applications -- VQA, automated Form filling, Smart Replies for Linguistic Codeswitching
por: Goyal, Naman
Publicado: (2024)
por: Goyal, Naman
Publicado: (2024)
Interpretable Next-token Prediction via the Generalized Induction Head
por: Kim, Eunji, et al.
Publicado: (2024)
por: Kim, Eunji, et al.
Publicado: (2024)
Bridging Linguistic Gaps: Cross-Lingual Mapping in Pre-Training and Dataset for Enhanced Multilingual LLM Performance
por: Zheng, Weihua, et al.
Publicado: (2026)
por: Zheng, Weihua, et al.
Publicado: (2026)
Understanding Emotion in Discourse: Recognition Insights and Linguistic Patterns for Generation
por: Jeong, Cheonkam, et al.
Publicado: (2026)
por: Jeong, Cheonkam, et al.
Publicado: (2026)
Fractal Patterns May Illuminate the Success of Next-Token Prediction
por: Alabdulmohsin, Ibrahim, et al.
Publicado: (2024)
por: Alabdulmohsin, Ibrahim, et al.
Publicado: (2024)
NDP: Next Distribution Prediction as a More Broad Target
por: Ruan, Junhao, et al.
Publicado: (2024)
por: Ruan, Junhao, et al.
Publicado: (2024)
NeoBERT: A Next-Generation BERT
por: Breton, Lola Le, et al.
Publicado: (2025)
por: Breton, Lola Le, et al.
Publicado: (2025)
Cautious Next Token Prediction
por: Wang, Yizhou, et al.
Publicado: (2025)
por: Wang, Yizhou, et al.
Publicado: (2025)
From Next-Token to Next-Block: A Principled Adaptation Path for Diffusion LLMs
por: Tian, Yuchuan, et al.
Publicado: (2025)
por: Tian, Yuchuan, et al.
Publicado: (2025)
Harnessing Linguistic Dissimilarity for Language Generalization on Unseen Low-Resource Varieties
por: Kim, Jinju, et al.
Publicado: (2026)
por: Kim, Jinju, et al.
Publicado: (2026)
Transition-Matrix Regularization for Next Dialogue Act Prediction in Counselling Conversations
por: Rudolph, Eric, et al.
Publicado: (2026)
por: Rudolph, Eric, et al.
Publicado: (2026)
PILOT: Steering Synthetic Data Generation with Psychological & Linguistic Output Targeting
por: Cisar, Caitlin, et al.
Publicado: (2025)
por: Cisar, Caitlin, et al.
Publicado: (2025)
Ejemplares similares
-
Zero-shot prompt-based classification: topic labeling in times of foundation models in German Tweets
por: Münker, Simon, et al.
Publicado: (2024) -
Towards Simulating Social Media Users with LLMs: Evaluating the Operational Validity of Conditioned Comment Prediction
por: Schwager, Nils, et al.
Publicado: (2026) -
Don't Trust Generative Agents to Mimic Communication on Social Networks Unless You Benchmarked their Empirical Realism
por: Münker, Simon, et al.
Publicado: (2025) -
InvBERT: Reconstructing Text from Contextualized Word Embeddings by inverting the BERT pipeline
por: Kugler, Kai, et al.
Publicado: (2021) -
Cultural Bias in Large Language Models: Evaluating AI Agents through Moral Questionnaires
por: Münker, Simon
Publicado: (2025)