INTIMA: A Benchmark for Human-AI Companionship Behavior
Fuente:
arXiv
Saved in:
| Main Authors: | Kaffee, Lucie-Aimée, Pistilli, Giada, Jernite, Yacine |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CIVICS: Building a Dataset for Examining Culturally-Informed Values in Large Language Models
by: Pistilli, Giada, et al.
Published: (2024)
by: Pistilli, Giada, et al.
Published: (2024)
Presumed Cultural Identity: How Names Shape LLM Responses
by: Pawar, Siddhesh, et al.
Published: (2025)
by: Pawar, Siddhesh, et al.
Published: (2025)
Economies of Open Intelligence: Tracing Power & Participation in the Model Ecosystem
by: Longpre, Shayne, et al.
Published: (2025)
by: Longpre, Shayne, et al.
Published: (2025)
Coordinated Flaw Disclosure for AI: Beyond Security Vulnerabilities
by: Cattell, Sven, et al.
Published: (2024)
by: Cattell, Sven, et al.
Published: (2024)
Investigating Wit, Creativity, and Detectability of Large Language Models in Domain-Specific Writing Style Adaptation of Reddit's Showerthoughts
by: Buz, Tolga, et al.
Published: (2024)
by: Buz, Tolga, et al.
Published: (2024)
Can AI be Consentful?
by: Pistilli, Giada, et al.
Published: (2025)
by: Pistilli, Giada, et al.
Published: (2025)
Local Differences, Global Lessons: Insights from Organisation Policies for International Legislation
by: Kaffee, Lucie-Aimée, et al.
Published: (2025)
by: Kaffee, Lucie-Aimée, et al.
Published: (2025)
AutoPal: Autonomous Adaptation to Users for Personal AI Companionship
by: Cheng, Yi, et al.
Published: (2024)
by: Cheng, Yi, et al.
Published: (2024)
Probing Pre-Trained Language Models for Cross-Cultural Differences in Values
by: Arora, Arnav, et al.
Published: (2022)
by: Arora, Arnav, et al.
Published: (2022)
Why Should This Article Be Deleted? Transparent Stance Detection in Multilingual Wikipedia Editor Discussions
by: Kaffee, Lucie-Aimée, et al.
Published: (2023)
by: Kaffee, Lucie-Aimée, et al.
Published: (2023)
Introducing ELLIPS: An Ethics-Centered Approach to Research on LLM-Based Inference of Psychiatric Conditions
by: Rocca, Roberta, et al.
Published: (2024)
by: Rocca, Roberta, et al.
Published: (2024)
Beyond Release: Access Considerations for Generative AI Systems
by: Solaiman, Irene, et al.
Published: (2025)
by: Solaiman, Irene, et al.
Published: (2025)
Can LLMs Truly Embody Human Personality? Analyzing AI and Human Behavior Alignment in Dispute Resolution
by: Kwon, Deuksin, et al.
Published: (2026)
by: Kwon, Deuksin, et al.
Published: (2026)
How Different AI Chatbots Behave? Benchmarking Large Language Models in Behavioral Economics Games
by: Xie, Yutong, et al.
Published: (2024)
by: Xie, Yutong, et al.
Published: (2024)
Fully Autonomous AI Agents Should Not be Developed
by: Mitchell, Margaret, et al.
Published: (2025)
by: Mitchell, Margaret, et al.
Published: (2025)
Wikimedia data for AI: a review of Wikimedia datasets for NLP tasks and AI-assisted editing
by: Johnson, Isaac, et al.
Published: (2024)
by: Johnson, Isaac, et al.
Published: (2024)
PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm
by: Li, Jing-Jing, et al.
Published: (2026)
by: Li, Jing-Jing, et al.
Published: (2026)
Inner Speech as Behavior Guides: Steerable Imitation of Diverse Behaviors for Human-AI coordination
by: Trivedi, Rakshit, et al.
Published: (2026)
by: Trivedi, Rakshit, et al.
Published: (2026)
DeceptionBench: A Comprehensive Benchmark for AI Deception Behaviors in Real-world Scenarios
by: Huang, Yao, et al.
Published: (2025)
by: Huang, Yao, et al.
Published: (2025)
A Human Behavioral Baseline for Collective Governance in Software Projects
by: Noori, Mobina, et al.
Published: (2025)
by: Noori, Mobina, et al.
Published: (2025)
Digital Companionship: Overlapping Uses of AI Companions and AI Assistants
by: Manoli, Aikaterina, et al.
Published: (2025)
by: Manoli, Aikaterina, et al.
Published: (2025)
Human Psychometric Questionnaires Mischaracterize LLM Behavior
by: Song, Woojung, et al.
Published: (2025)
by: Song, Woojung, et al.
Published: (2025)
Let the Results Speak: A Replication-First Paradigm for LLM Behavioral Benchmarking
by: Yuming, et al.
Published: (2026)
by: Yuming, et al.
Published: (2026)
Tailored Conversations beyond LLMs: A RL-Based Dialogue Manager
by: Galland, Lucie, et al.
Published: (2025)
by: Galland, Lucie, et al.
Published: (2025)
Are Large Language Models Really Bias-Free? Jailbreak Prompts for Assessing Adversarial Robustness to Bias Elicitation
by: Cantini, Riccardo, et al.
Published: (2024)
by: Cantini, Riccardo, et al.
Published: (2024)
SimBench: Benchmarking the Ability of Large Language Models to Simulate Human Behaviors
by: Hu, Tiancheng, et al.
Published: (2025)
by: Hu, Tiancheng, et al.
Published: (2025)
Are You Human? An Adversarial Benchmark to Expose LLMs
by: Gressel, Gilad, et al.
Published: (2024)
by: Gressel, Gilad, et al.
Published: (2024)
Can AI Be as Creative as Humans?
by: Wang, Haonan, et al.
Published: (2024)
by: Wang, Haonan, et al.
Published: (2024)
Towards Real-world Human Behavior Simulation: Benchmarking Large Language Models on Long-horizon, Cross-scenario, Heterogeneous Behavior Traces
by: Chen, Jiawei, et al.
Published: (2026)
by: Chen, Jiawei, et al.
Published: (2026)
ReplicatorBench: Benchmarking LLM Agents for Replicability in Social and Behavioral Sciences
by: Nguyen, Bang, et al.
Published: (2026)
by: Nguyen, Bang, et al.
Published: (2026)
AI Idea Bench 2025: AI Research Idea Generation Benchmark
by: Qiu, Yansheng, et al.
Published: (2025)
by: Qiu, Yansheng, et al.
Published: (2025)
ChatBench: From Static Benchmarks to Human-AI Evaluation
by: Chang, Serina, et al.
Published: (2025)
by: Chang, Serina, et al.
Published: (2025)
MTI: A Behavior-Based Temperament Profiling System for AI Agents
by: Jeong, Jihoon
Published: (2026)
by: Jeong, Jihoon
Published: (2026)
CulturalBench: A Robust, Diverse, and Challenging Cultural Benchmark by Human-AI CulturalTeaming
by: Chiu, Yu Ying, et al.
Published: (2024)
by: Chiu, Yu Ying, et al.
Published: (2024)
Aligning Large Language Model Behavior with Human Citation Preferences
by: Ando, Kenichiro, et al.
Published: (2026)
by: Ando, Kenichiro, et al.
Published: (2026)
Teaching Values to Machines: Simulating Human-Like Behavior in LLMs
by: Yehudai, Asaf, et al.
Published: (2026)
by: Yehudai, Asaf, et al.
Published: (2026)
Is Human-Like Text Liked by Humans? Multilingual Human Detection and Preference Against AI
by: Wang, Yuxia, et al.
Published: (2025)
by: Wang, Yuxia, et al.
Published: (2025)
RECODE-H: A Benchmark for Research Code Development with Interactive Human Feedback
by: Miao, Chunyu, et al.
Published: (2025)
by: Miao, Chunyu, et al.
Published: (2025)
HEART: A Unified Benchmark for Assessing Humans and LLMs in Emotional Support Dialogue
by: Iyer, Laya, et al.
Published: (2026)
by: Iyer, Laya, et al.
Published: (2026)
IKnow: Instruction-Knowledge-Aware Continual Pretraining for Effective Domain Adaptation
by: Zhang, Tianyi, et al.
Published: (2025)
by: Zhang, Tianyi, et al.
Published: (2025)
Similar Items
-
CIVICS: Building a Dataset for Examining Culturally-Informed Values in Large Language Models
by: Pistilli, Giada, et al.
Published: (2024) -
Presumed Cultural Identity: How Names Shape LLM Responses
by: Pawar, Siddhesh, et al.
Published: (2025) -
Economies of Open Intelligence: Tracing Power & Participation in the Model Ecosystem
by: Longpre, Shayne, et al.
Published: (2025) -
Coordinated Flaw Disclosure for AI: Beyond Security Vulnerabilities
by: Cattell, Sven, et al.
Published: (2024) -
Investigating Wit, Creativity, and Detectability of Large Language Models in Domain-Specific Writing Style Adaptation of Reddit's Showerthoughts
by: Buz, Tolga, et al.
Published: (2024)