
Examining Human-Like Behaviors in LLM Conversations

Large language models regularly show human-like behaviors, such as expressing thoughts and emotions, building relationships, and refusing or setting boundaries. The authors, Sunnie S. Y. Kim, Margit Bowler, and Leon A Gelys, note that practitioners lack methods or empirical evidence to decide when and what kind of human-like behavior should be shown. This paper accordingly presents a multi-dimensional analysis of prevalence, perceived effects, and controllability using both LLM-as-a-judge and human evaluation.
Their study spans 21,000 multi-turn conversations across closed widely used models: GPT-4, LLM-2.5-flash-4.1-mini, GPT-comment-4, and Gemini-5.5-flash. The results show that human-like behaviors are pervasive, but their appearance varies substantially across models and user factors, including conversation goals and user profiles. In human judgments of appropriateness, self-referential and relationship-building behaviors were rated as less appropriate when coming from an LLM than when coming from a human. Conversely, boundary-maintaining behaviors were rated as more appropriate from an LLMs.
The authors also report that system prompting can guide or control these behaviors, but only with careful evaluation because prompts can produce unintended effects. They close with recommendations for responsible LLM design and evaluation, emphasizing that human-like behavior should be measured explicitly rather than treated as an incidental side effect of conversational ability.


