From Plausible to Causal: Counterfactual Semantics for Policy Evaluation in Simulated Online Communities
Fuente:
arXiv
Saved in:
| Main Authors: | Goyal, Agam, Wang, Yian, Chandrasekharan, Eshwar, Sundaram, Hari |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Social Simulacra in the Wild: AI Agent Communities on Moltbook
by: Goyal, Agam, et al.
Published: (2026)
by: Goyal, Agam, et al.
Published: (2026)
CausalDetox: Causal Head Selection and Intervention for Language Model Detoxification
by: Wang, Yian, et al.
Published: (2026)
by: Wang, Yian, et al.
Published: (2026)
ArgCMV: An Argument Summarization Benchmark for the LLM-era
by: Gurjar, Omkar, et al.
Published: (2025)
by: Gurjar, Omkar, et al.
Published: (2025)
The Hidden Toll of Social Media News: Causal Effects on Psychosocial Wellbeing
by: Pal, Olivia, et al.
Published: (2026)
by: Pal, Olivia, et al.
Published: (2026)
MoMoE: Mixture of Moderation Experts Framework for AI-Assisted Online Governance
by: Goyal, Agam, et al.
Published: (2025)
by: Goyal, Agam, et al.
Published: (2025)
Toxic HallucinAItions: Perturbing Prompts and Tracing LLM Circuits
by: Shimgekar, Soorya Ram, et al.
Published: (2026)
by: Shimgekar, Soorya Ram, et al.
Published: (2026)
Answer Bubbles: Information Exposure in AI-Mediated Search
by: Huang, Michelle, et al.
Published: (2026)
by: Huang, Michelle, et al.
Published: (2026)
The Language of Approval: Identifying the Drivers of Positive Feedback Online
by: Goyal, Agam, et al.
Published: (2025)
by: Goyal, Agam, et al.
Published: (2025)
Algorithmic Cultivation: How Social Media Feeds Shape User Language
by: Pal, Olivia, et al.
Published: (2026)
by: Pal, Olivia, et al.
Published: (2026)
SLM-Mod: Small Language Models Surpass LLMs at Content Moderation
by: Zhan, Xianyang, et al.
Published: (2024)
by: Zhan, Xianyang, et al.
Published: (2024)
The Chilling: Identifying Strategic Antisocial Behavior Online and Examining the Impact on Journalists
by: Wang, Yian, et al.
Published: (2025)
by: Wang, Yian, et al.
Published: (2025)
Breaking Bad Tokens: Detoxification of LLMs Using Sparse Autoencoders
by: Goyal, Agam, et al.
Published: (2025)
by: Goyal, Agam, et al.
Published: (2025)
VASTU: Value-Aligned Social Toolkit for Online Content Curation
by: Goyal, Agam, et al.
Published: (2026)
by: Goyal, Agam, et al.
Published: (2026)
Mind Your Ps and Qs: Supporting Positive Reinforcement in Moderation Through a Positive Queue
by: Lambert, Charlotte, et al.
Published: (2025)
by: Lambert, Charlotte, et al.
Published: (2025)
Masking or Mitigating? Deconstructing the Impact of Query Rewriting on Retriever Biases in RAG
by: Goyal, Agam, et al.
Published: (2026)
by: Goyal, Agam, et al.
Published: (2026)
Uncovering the Internet's Hidden Values: An Empirical Study of Desirable Behavior Using Highly-Upvoted Content on Reddit
by: Goyal, Agam, et al.
Published: (2024)
by: Goyal, Agam, et al.
Published: (2024)
State Contamination in Memory-Augmented LLM Agents
by: Wang, Yian, et al.
Published: (2026)
by: Wang, Yian, et al.
Published: (2026)
Venire: A Machine Learning-Guided Panel Review System for Community Content Moderation
by: Koshy, Vinay, et al.
Published: (2024)
by: Koshy, Vinay, et al.
Published: (2024)
LLUMI: Improving LLM Writing Assistance for Mental Health Support with Online Community Feedback
by: Kim, Jiwon, et al.
Published: (2026)
by: Kim, Jiwon, et al.
Published: (2026)
Needling Through the Threads: A Visualization Tool for Navigating Threaded Online Discussions
by: Liu, Yijun, et al.
Published: (2025)
by: Liu, Yijun, et al.
Published: (2025)
Detecting Early and Implicit Suicidal Ideation via Longitudinal and Information Environment Signals on Social Media
by: Shimgekar, Soorya Ram, et al.
Published: (2025)
by: Shimgekar, Soorya Ram, et al.
Published: (2025)
Towards a Better Modqueue: Designing for Diversity Across Moderator Objectives and Workflows
by: Bajpai, Tanvi, et al.
Published: (2024)
by: Bajpai, Tanvi, et al.
Published: (2024)
"Think about it like you're a firefighter": Understanding How Reddit Moderators Use the Modqueue
by: Bajpai, Tanvi, et al.
Published: (2025)
by: Bajpai, Tanvi, et al.
Published: (2025)
Designing Usable Controls for Customizable Social Media Feeds
by: Choi, Frederick, et al.
Published: (2025)
by: Choi, Frederick, et al.
Published: (2025)
PAIR-SAFE: A Paired-Agent Approach for Runtime Auditing and Refining AI-Mediated Mental Health Support
by: Kim, Jiwon, et al.
Published: (2026)
by: Kim, Jiwon, et al.
Published: (2026)
Plausible-Parrots @ MSP2023: Enhancing Semantic Plausibility Modeling using Entity and Event Knowledge
by: Shen, Chong, et al.
Published: (2024)
by: Shen, Chong, et al.
Published: (2024)
Modelling Adjectival Modification Effects on Semantic Plausibility
by: Golub, Anna, et al.
Published: (2025)
by: Golub, Anna, et al.
Published: (2025)
Learning to Simulate Human Dialogue
by: Gandhi, Kanishk, et al.
Published: (2026)
by: Gandhi, Kanishk, et al.
Published: (2026)
CEV-LM: Controlled Edit Vector Language Model for Shaping Natural Language Generations
by: Moorjani, Samraj, et al.
Published: (2024)
by: Moorjani, Samraj, et al.
Published: (2024)
Causal-Counterfactual RAG: The Integration of Causal-Counterfactual Reasoning into RAG
by: Khadilkar, Harshad, et al.
Published: (2025)
by: Khadilkar, Harshad, et al.
Published: (2025)
Causal Discovery and Counterfactual Reasoning to Optimize Persuasive Dialogue Policies
by: Zeng, Donghuo, et al.
Published: (2025)
by: Zeng, Donghuo, et al.
Published: (2025)
Estimating Commonsense Plausibility through Semantic Shifts
by: Cui, Wanqing, et al.
Published: (2025)
by: Cui, Wanqing, et al.
Published: (2025)
Spontaneous Speech Variables for Evaluating LLMs Cognitive Plausibility
by: Wang, Sheng-Fu, et al.
Published: (2025)
by: Wang, Sheng-Fu, et al.
Published: (2025)
BERTScoreVisualizer: A Web Tool for Understanding Simplified Text Evaluation with BERTScore
by: Jaskowski, Sebastian, et al.
Published: (2024)
by: Jaskowski, Sebastian, et al.
Published: (2024)
Does Positive Reinforcement Work?: A Quasi-Experimental Study of the Effects of Positive Feedback on Reddit
by: Lambert, Charlotte, et al.
Published: (2024)
by: Lambert, Charlotte, et al.
Published: (2024)
PRobELM: Plausibility Ranking Evaluation for Language Models
by: Yuan, Zhangdie, et al.
Published: (2024)
by: Yuan, Zhangdie, et al.
Published: (2024)
Topic-aware Causal Intervention for Counterfactual Detection
by: Nguyen, Thong, et al.
Published: (2024)
by: Nguyen, Thong, et al.
Published: (2024)
Simulating Opinion Dynamics with Networks of LLM-based Agents
by: Chuang, Yun-Shiuan, et al.
Published: (2023)
by: Chuang, Yun-Shiuan, et al.
Published: (2023)
Estimating Online Influence Needs Causal Modeling! Counterfactual Analysis of Social Media Engagement
by: Tian, Lin, et al.
Published: (2025)
by: Tian, Lin, et al.
Published: (2025)
Plausibility Vaccine: Injecting LLM Knowledge for Event Plausibility
by: Chmura, Jacob, et al.
Published: (2025)
by: Chmura, Jacob, et al.
Published: (2025)
Similar Items
-
Social Simulacra in the Wild: AI Agent Communities on Moltbook
by: Goyal, Agam, et al.
Published: (2026) -
CausalDetox: Causal Head Selection and Intervention for Language Model Detoxification
by: Wang, Yian, et al.
Published: (2026) -
ArgCMV: An Argument Summarization Benchmark for the LLM-era
by: Gurjar, Omkar, et al.
Published: (2025) -
The Hidden Toll of Social Media News: Causal Effects on Psychosocial Wellbeing
by: Pal, Olivia, et al.
Published: (2026) -
MoMoE: Mixture of Moderation Experts Framework for AI-Assisted Online Governance
by: Goyal, Agam, et al.
Published: (2025)