Immerse & befriend: The role of synthetic relationship perception and narrative transportation in a mental health app
Vol.20,No.4(2026)
The global mental health crisis has driven the rise of apps using conversational agents (CAs) to complement or replace human professional support. Offline, the success of such interventions often depends on the relationship between the help-seeker and the support provider. Online, psychological processes elicited by specific app features relating to the CA and its unique story design might play a significant role. However, research on these processes and their consequences on user engagement and well-being is scarce. To shed light on this, we conducted a cross-sectional study with 348 active users of the mental health app Betwixt. The app integrates a scripted CA as a mentor, guiding users through levels of self-development within a dream-like fantasy narrative. We tested how perceived synthetic relationships with the CA (based on the Relational Models Theory, Fiske, 1992) and narrative transportation (i.e., the extent to which users become absorbed in the story) relate to app engagement and perceived stress. Our findings indicate that a strong sense of narrative transportation and perceiving the CA as friend-like positively predict hedonic benefits and future usage intentions. However, these factors were not significantly associated with perceived stress. Our results contribute to the emerging field of synthetic relationships and highlight the potential of narrative transportation for developers of apps using similar designs. We also discuss the ethical concerns of fostering friend-like relationships with CAs and their practical implications for developers and mental health professionals.
mental health; conversational agent; narrative transportation; human-agent relationship; usage intention; stress; synthetic relationship
Marisa Tschopp
Institute of Human Behaviour, Society and Technology, ZHAW School of Applied Psychology, Zurich, Switzerland & scip AG, Zurich, Switzerland
Dr. Marisa Tschopp is the head of the Humans & AI Subject Area at the ZHAW School of Applied Psychology and a senior corporate researcher at the cybersecurity company scip AG. As an AI psychologist, her work focuses on human–AI relationships across various contexts, ranging from AI companionship to human–AI teaming and AI in mental health.
Stefanie H. Klein
Everyday Media Lab, Leibniz-Institut für Wissensmedien, Tübingen, Germany
Dr. Stefanie H. Klein is a postdoctoral researcher in the Everyday Media Lab at Leibniz-Institut für Wissensmedien in Tübingen, Germany. Her research focuses on human-machine communication, with an emphasis on people’s perceptions of communicative artificial intelligence in information search and health contexts, and on how interactions with AI affect interpersonal relationships.
Christine Anderl
Everyday Media Lab, Leibniz-Institut für Wissensmedien, Tübingen, Germany
Dr. Christine Anderl is head of science at Endo Health GmbH, Germany. Previously, she was a postdoc in the Everyday Media Lab at Leibniz-Institut für Wissensmedien in Tübingen, Germany. Her research focuses on digital health, digital communication, effects of social and mobile media use, and digital transformation.
Henrik Sætra
Department of Informatics, University of Oslo, Oslo, Norway
Dr. Henrik Skaug Sætra is an associate professor at the Institute of Informatics (University of Oslo) and leads the research group Technology and Sustainable Futures. He has a background in political philosophy and adopts a broad and interdisciplinary approach to the study of the political, ethical, and social implications of emerging technologies.
Sonja Utz
Everyday Media Lab, Leibniz-Institut für Wissensmedien, Tübingen, Germany & Department of Psychology, Eberhard Karls Universität Tübingen, Tübingen, Germany
Prof. Dr. Sonja Utz is the head of the Everyday Media lab at Leibniz-Institut für Wissensmedien in Tübingen and a full professor for communication via social media at the University of Tübingen. Her research focuses on the effects of social and mobile media use, especially in knowledge-related contexts, and on human-machine interaction.
Abd-Alrazaq, A. A., Rababeh, A., Alajlani, M., Bewick, B. M., & Househ, M. (2020). Effectiveness and safety of using chatbots to improve mental health: Systematic review and meta-analysis. Journal of Medical Internet Research, 22(7), Article 16021. https://doi.org/10.2196/16021
Altman, I., & Taylor, D. A. (1973). Social penetration: The development of interpersonal relationships. Holt, Rinehart & Winston.
Appel, M., Gnambs, T., Richter, T., & Green, M. C. (2015). The Transportation Scale–Short Form (TS–SF). Media Psychology, 18(2), 243–266. https://doi.org/10.1080/15213269.2014.987400
Berscheid, E. (1994). Interpersonal relationships. Annual Review of Psychology, 45, 79–129. https://doi.org/10.1146/annurev.ps.45.020194.000455
Betwixt (2023a). Betwixt. https://www.betwixt.life/
Betwixt (2023b). Betwixt – Research. https://www.Betwixt.life/research
Bowlby, J. (1979). The Bowlby-Ainsworth attachment theory. Behavioral and Brain Sciences, 2(4), 637–638. https://doi.org/10.1017/S0140525X00064955
Brandtzaeg, P. B., Skjuve, M., & Følstad, A. (2022). My AI friend: How users of a social chatbot understand their human–AI friendship. Human Communication Research, 48(3), 404–429. https://doi.org/10.1093/hcr/hqac008
Bucci, S., Schwannauer, M., & Berry, N. (2019). The digital revolution and its impact on mental health care. Psychology and Psychotherapy, 92(2), 277–297. https://doi.org/10.1111/papt.12222
Chandrashekar, P. (2018). Do mental health mobile apps work: Evidence and recommendations for designing high-efficacy mental health mobile apps. mHealth, 4(3), Article 3. https://doi.org/10.21037/mhealth.2018.03.02
Chen, C., Di Russo, C., Yang, H., Shao, R., Krieger, M., & Sundar, S. S. (2019, August 6). Alexa, Netflix, and Siri: User perceptions of AI-driven technologies [Paper presentation]. The 102nd Annual Conference of Association for Education in Journalism and Mass Communication (AEJMC), Toronto.
Cohen, S., Kamarck, T., & Mermelstein, R. (1983). A global measure of perceived stress. Journal of Health and Social Behavior, 24(4), 385–396. https://doi.org/10.2307/2136404
Darcy, A., Daniels, J., Salinger, D., Wicks, P., & Robinson, A. (2021). Evidence of human-level bonds established with a digital conversational agent: Cross-sectional, retrospective observational study. JMIR Formative Research, 5(5), Article e27868. https://doi.org/10.2196/27868
Davis, M. (2023, March 31). Man talks with AI ChatBot about climate change fears, ends up killing self while AI assures him they’ll be “together as one in heaven.” Science Times. https://www.sciencetimes.com/articles/43072/20230331/man-talks-ai-chatbot-climate-change-fears-ends-up-killing.htm
De Graaf, A., Sanders, J., & Hoeken, H. (2016). Characteristics of narrative interventions and health effects: A review of the content, form, and context of narratives in health-related narrative persuasion research. Review of Communication Research, 4, 88–131. https://doi.org/10.12840/issn.2255-4165.2016.04.01.011
Epley, N., Waytz, A., & Cacioppo, J. T. (2007). On seeing human: A three-factor theory of anthropomorphism. Psychological Review, 114(4), 864–886. https://doi.org/10.1037/0033-295X.114.4.864
Faul, F., Erdfelder, E., Lang, A.-G., & Buchner, A. (2007). G*Power 3: A flexible statistical power analysis program for the social, behavioral, and biomedical sciences. Behavior Research Methods, 39(2), 175–191. https://doi.org/10.3758/bf03193146
Fiske, A. P. (1992). The four elementary forms of sociality: Framework for a unified theory of social relations. Psychological Review, 99(4), 689–723. https://doi.org/10.1037/0033-295X.99.4.689
Gabriel, S., Valenti, J., & Young, A. F. (2016). Social surrogates, social motivations, and everyday activities: The case for a strong, subtle, and sneaky social self. In J. M. Olson & M. P. Zanna (Eds.), Advances in experimental social psychology (Vol. 53, pp. 189–243). Academic Press. https://doi.org/10.1016/bs.aesp.2015.09.003
Green, M. C., & Appel, M. (2024). Narrative transportation: How stories shape how we see ourselves and the world. Advances in Experimental Social Psychology, 70, 1–82. https://doi.org/10.1016/bs.aesp.2024.03.002
Green, M. C., & Brock, T. C. (2000). The role of transportation in the persuasiveness of public narratives. Journal of Personality and Social Psychology, 79(5), 701–721. https://doi.org/10.1037/0022-3514.79.5.701
Green, M. C., Brock, T. C., & Kaufman, G. F. (2004). Understanding media enjoyment: The role of transportation into narrative worlds. Communication Theory, 14(4), 311–327. https://doi.org/10.1111/j.1468-2885.2004.tb00317.x
Green, M. C., & Clark, J. L. (2013). Transportation into narrative worlds: Implications for entertainment media influences on tobacco use. Addiction, 108(3), 477–484. https://doi.org/10.1111/j.1360-0443.2012.04088.x
Guzman, A. L., & Lewis, S. C. (2020). Artificial intelligence and communication: A human-machine communication research agenda. New Media & Society, 22(1), 70–86. https://doi.org/10.1177/1461444819858691
Harmon, S. (2021). Master of two worlds: Narrative intelligence as the next step for mental health chatbots. In Proceedings of the Realizing AI in Healthcare: Challenges Appearing in the Wild Workshop (CHI '21). Association for Computing Machinery. https://francisconunes.me/RealizingAIinHealthcareWS/papers/Harmon2021.pdf
Haslam, N., & Fiske, A. P. (1999). Relational models theory: A confirmatory factor analysis. Personal Relationships, 6(2), 241–250. https://doi.org/10.1111/j.1475-6811.1999.tb00190.x
He, Y., Yang, L., Qian, C., Li, T., Su, Z., Zhang, Q., & Hou, X. (2023). Conversational agent interventions for mental health problems: Systematic review and meta-analysis of randomized controlled trials. Journal of Medical Internet Research, 25(1), Article e43862. https://doi.org/10.2196/43862
Hickey, B. A., Chalmers, T., Newton, P., Lin, C.-T., Sibbritt, D., McLachlan, C. S., Clifton-Bligh, R., Morley, J., & Lal, S. (2021). Smart devices and wearable technologies to detect and monitor mental health conditions and stress: A systematic review. Sensors, 21(10), Article 10. https://doi.org/10.3390/s21103461
Hofmann, S. G., Asnaani, A., Vonk, I. J. J., Sawyer, A. T., & Fang, A. (2012). The efficacy of cognitive behavioral therapy: A review of meta-analyses. Cognitive Therapy and Research, 36(5), 427–440. https://doi.org/10.1007/s10608-012-9476-1
Horton, D., & Wohl, R. (1956). Mass communication and para-social interaction: Observations on intimacy at a distance. Psychiatry, 19(3), 215–229. https://doi.org/10.1080/00332747.1956.11023049
IBM Corp. (2021). IBM SPSS Statistics for Windows (Version 28.0) [Computer software]. IBM Corp.
Inkster, B., Sarda, S., & Subramanian, V. (2018). An empathy-driven, conversational artificial intelligence agent (Wysa) for digital mental well-being: Real-world data evaluation mixed-methods study. JMIR mHealth and uHealth, 6(11), Article e12106. https://doi.org/10.2196/12106
Isberner, M.-B., Richter, T., Schreiner, C., Eisenbach, Y., Sommer, C., & Appel, M. (2018). Empowering stories: Transportation into narratives with strong protagonists increases self-related control beliefs. Discourse Processes, 56(8), 575–598. https://doi.org/10.1080/0163853X.2018.1526032
King, D. R., Emerson, M. R., Tartaglia, J., Nanda, G., & Tatro, N. A. (2023). Methods for navigating the mobile mental health app landscape for clinical use. Current Treatment Options in Psychiatry, 10(2), 72–86. https://doi.org/10.1007/s40501-023-00288-4
Knapp, M. L. (1978). Social intercourse: From greeting to goodbye. Allyn and Bacon. http://archive.org/details/socialintercours00knap
Liang, J.-C., & Hwang, G.-J. (2023). A robot-based digital storytelling approach to enhancing EFL learners’ multimodal storytelling ability and narrative engagement. Computers & Education, 201, Article 104827. https://doi.org/10.1016/j.compedu.2023.104827
Limpanopparat, S., Gibson, E., & Harris, D. A. (2024). User engagement, attitudes, and the effectiveness of chatbots as a mental health intervention: A systematic review. Computers in Human Behavior: Artificial Humans, 2(2), Article 100081. https://doi.org/10.1016/j.chbah.2024.100081
Lucas, G. M., Rizzo, A., Gratch, J., Scherer, S., Stratou, G., Boberg, J., & Morency, L.-P. (2017). Reporting mental health symptoms: Breaking down barriers to care with virtual human interviewers. Frontiers in Robotics and AI, 4, Article 51. https://doi.org/10.3389/frobt.2017.00051
Malfacini, K. (2025). The impacts of companion AI on human relationships: Risks, benefits, and design considerations. AI & SOCIETY, 40(7), 5527–5540. https://doi.org/10.1007/s00146-025-02318-6
Maples, B., Cerit, M., Vishwanath, A., & Pea, R. (2024). Loneliness and suicide mitigation for students using GPT3-enabled chatbots. npj Mental Health Research, 3(1), Article 4. https://doi.org/10.1038/s44184-023-00047-6
NIMH. (2024, August). Technology and the future of mental health treatment. National Institute of Mental Health. https://www.nimh.nih.gov/health/topics/technology-and-the-future-of-mental-health-treatment
Nass, C., & Moon, Y. (2000). Machines and mindlessness: Social responses to computers. Journal of Social Issues, 56(1), 81–103. https://doi.org/10.1111/0022-4537.00153
Oh, J., Lim, H. S., & Hwang, A. H.-C. (2020). How interactive storytelling persuades: The mediating role of website contingency and narrative transportation. Journal of Broadcasting & Electronic Media, 64(5), 714–735. https://doi.org/10.1080/08838151.2020.1848180
Olawade, D. B., Wada, O. Z., Odetayo, A., David-Olawade, A. C., Asaolu, F., & Eberhardt, J. (2024). Enhancing mental health with Artificial Intelligence: Current trends and future prospects. Journal of Medicine, Surgery, and Public Health, 3, Article 100099. https://doi.org/10.1016/j.glmedi.2024.100099
Pentina, I., Xie, T., Hancock, T., & Bailey, A. (2023). Consumer–machine relationships in the age of artificial intelligence: Systematic literature review and research directions. Psychology & Marketing, 40(8), 1593–1614. https://doi.org/10.1002/mar.21853
Roose, K. (2024, October 23). Can A.I. be blamed for a teen’s suicide? The New York Times. https://www.nytimes.com/2024/10/23/technology/characterai-lawsuit-teen-suicide.html
Sætra, H. S. (2020). The parasitic nature of social AI: Sharing minds with the mindless. Integrative Psychological and Behavioral Science, 54(2), 308–326. https://doi.org/10.1007/s12124-020-09523-6
Sætra, H. S. (2021). Social robot deception and the culture of trust. Paladyn, Journal of Behavioral Robotics, 12(1), 276–286. https://doi.org/10.1515/pjbr-2021-0021
Sætra, H. S., & Selinger, E. (2023). The siren song of technological remedies for social problems: Defining, demarcating, and evaluating techno-fixes and techno-solutionism. SSRN Electronic Journal. https://doi.org/10.2139/ssrn.4576687
Schmidt-Kraepelin, M., Thiebes, S., Warsinsky, S. L., Petter, S., & Sunyaev, A. (2023). Narrative transportation in gamified information systems: The role of narrative-task congruence. In A. Schmidt, K. Väänänen, T. Goyal, P. O. Kristensson, & A. Peters (Eds.), Extended abstracts of the 2023 CHI Conference on Human Factors in Computing Systems (Article 215). Association for Computing Machinery. https://doi.org/10.1145/3544549.3585595
Schroeder, J., Wilkes, C., Rowan, K., Toledo, A., Paradiso, A., Czerwinski, M., Mark, G., & Linehan, M. M. (2018). Pocket Skills: A conversational mobile web app to support dialectical behavioral therapy. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems (Article 398). Association for Computing Machinery. https://doi.org/10.1145/3173574.3173972
Sedlakova, J., & Trachsel, M. (2023). Conversational artificial intelligence in psychotherapy: A new therapeutic tool or agent? The American Journal of Bioethics, 23(5), 4–13. https://doi.org/10.1080/15265161.2022.2048739
Seymour, W., & Van Kleek, M. (2021). Exploring interactions between trust, anthropomorphism, and relationship development in voice assistants. Proceedings of the ACM on Human-Computer Interaction, 5(CSCW2), Article 371. https://doi.org/10.1145/3479515
Sharma, M., & Rush, S. E. (2014). Mindfulness-based stress reduction as a stress management intervention for healthy individuals: A systematic review. Journal of Evidence-Based Complementary & Alternative Medicine, 19(4), 271–286. https://doi.org/10.1177/2156587214543143
Shen, F., Sheer, V. C., & Li, R. (2015). Impact of narratives on persuasion in health communication: A meta-analysis. Journal of Advertising, 44(2), 105–113. https://doi.org/10.1080/00913367.2015.1018467
Sherrick, B. (2018). The role of engagement in facilitating games-based persuasion. In N. D. Bowman (Ed.), Video games: A medium that demands our attention (pp. 44–59). Routledge.
https://doi.org/10.4324/9781351235266-3
Skjuve, M., Følstad, A., Fostervold, K. I., & Brandtzaeg, P. B. (2021). My chatbot companion-A study of human-chatbot relationships. International Journal of Human-Computer Studies, 149, Article 102601. https://doi.org/10.1016/j.ijhcs.2021.102601
Stringer, H. (2024). Mental health care is in high demand. Psychologists are leveraging tech and peers to meet the need. Monitor on Psychology, 55(1). https://www.apa.org/monitor/2024/01/trends-pathways-access-mental-health-care
Starke, C., Ventura, A., Bersch, C., Cha, M., de Vreese, C., Doebler, P., Dong, M., Krämer, N., Leib, M., Peter, J., Schäfer, L., Soraperra, I., Szczuka, J., Tuchtfeld, E., Wald, R., & Köbis, N. (2024). Risks and protective measures for synthetic relationships. Nature Human Behaviour, 8(10), 1834–1836. https://doi.org/10.1038/s41562-024-02005-4
Sundar, S. S., & Kim, J. (2019). Machine heuristic: When we trust computers more than humans with our personal information. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems (Article 538). Association for Computing Machinery. https://doi.org/10.1145/3290605.3300768
Sundar, S. S. (2008). The MAIN model: A heuristic approach to understanding technology effects on credibility. In M. Metzger & A. Flanagin (Eds.), Digital media, youth, and credibility (pp. 73–100). The MIT Press.
Tassiello, V., Tillotson, J. S., & Rome, A. S. (2021). “Alexa, order me a pizza!”: The mediating role of psychological power in the consumer–voice assistant interaction. Psychology & Marketing, 38(7), 1069–1080. https://doi.org/10.1002/mar.21488
Thomas, V. L., & Grigsby, J. L. (2024). Narrative transportation: A systematic literature review and future research agenda. Psychology & Marketing, 41(8), 1805–1819. https://doi.org/10.1002/mar.22011
Tong, F., Lederman, R., D’Alfonso, S., Berry, K., & Bucci, S. (2022). Digital therapeutic alliance with fully automated mental health smartphone apps: A narrative review. Frontiers in Psychiatry, 13, Article 819623. https://doi.org/10.3389/fpsyt.2022.819623
Torous, J., Nicholas, J., Larsen, M. E., Firth, J., & Christensen, H. (2018). Clinical review of user engagement with mental health smartphone apps: Evidence, theory and improvements. Evidence-Based Mental Health, 21(3), 116–119. https://doi.org/10.1136/eb-2018-102891
Tschopp, M., Gieselmann, M., & Sassenberg, K. (2023). Servant by default? How humans perceive their relationship with conversational AI. Cyberpsychology: Journal of Psychosocial Research on Cyberspace, 17(3), Article 3. https://doi.org/10.5817/CP2023-3-9
Tschopp, M., & Sassenberg, K. (2024). The impact of human-AI relationship perception on voice shopping intentions. Human-Machine Communication, 8(1), 101–117. https://doi.org/10.30658/hmc.8.5
Van Laer, T., de Ruyter, K., & Wetzels, M. (2012). Effects of narrative transportation on persuasion: A meta-analysis. NA – Advances in Consumer Research, 40, 579–581. https://openaccess.city.ac.uk/id/eprint/16984/
Venkatesh, V., Thong, J. Y. L., & Xu, X. (2012). Consumer acceptance and use of information technology: Extending the unified theory of acceptance and use of technology. MIS Quarterly, 36(1), 157–178. https://doi.org/10.2307/41410412
Ventura, A., Starke, C., Righetti, F., & Köbis, N. (2025). Relationships in the age of AI: A review on the opportunities and risks of synthetic relationships to reduce loneliness. OSF. https://doi.org/10.31234/osf.io/w7nmz_v1
Watts, J. (2023). A journey through communication research on transportation: The future of narrative transportation on emerging forms of media. Review of Communication, 23(4), 367–384. https://doi.org/10.1080/15358593.2023.2239321
Weizenbaum, J. (1966). ELIZA—a computer program for the study of natural language communication between man and machine. Communications of the ACM, 9(1), 36–45. https://doi.org/10.1145/365153.365168
World Health Organization. (2022, June 17). Mental health. World Health Organization. https://www.who.int/news-room/fact-sheets/detail/mental-health-strengthening-our-response
Xie, T., & Pentina, I. (2022). Attachment theory as a framework to understand relationships with social chatbots: A case study of Replika. In Proceedings of the 55th Hawaii International Conference on System Sciences (pp. 2046–2055). https://aisel.aisnet.org/hicss-55/da/social_robots/4
Authors' Contribution
Marisa Tschopp: conceptualization, data curation, investigation, methodology, project administration, resources, validation, visualization, writing—original draft, writing—review & editing. Stefanie H. Klein: conceptualization, methodology, resources, data curation, validation, formal analysis, visualization, writing—original draft, writing—review & editing. Christine Anderl: conceptualization, methodology, writing—review & editing. Henrik Skaug Sætra: conceptualization, writing—original draft, writing—review & editing. Sonja Utz: conceptualization, methodology, resources, data curation, formal analysis, writing—review & editing, supervision, project administration.
Editorial Record
First submission received:
March 28, 2025
Revisions received:
February 24, 2026
June 12, 2026
Accepted for publication:
June 18, 2026
Editor in charge:
Lenka Dedkova
Introduction
We are currently facing a global mental health crisis, with a significant shortage of human mental health services (Stringer, 2024). While systemic solutions are needed, technology may help (Sætra & Selinger, 2023). Artificial intelligence (AI)-driven mental health apps have grown significantly (Olawade et al., 2024), with conversational agents (CAs) promoted as therapy supplements, well-being enhancers, or crisis support (Sedlakova & Trachsel, 2023). Evidence suggests that human-CA interactions can benefit mental health (CAs with and without AI-technology). For instance, users of AI-chatbot Replika1 reported reduced suicidal thoughts (Maples et al., 2024); other studies found decreases in depression, anxiety (Schroeder et al., 2018), and improved mood with Wysa2 (Inkster et al., 2018). However, interactions with CAs have raised stark concerns lately (Starke et al., 2024), including unhealthy attachments or increased isolation (Sætra, 2020, 2021), reportedly even associated with at least two fatal consequences (Davis, 2023).
Despite these pressing issues, empirical research on CAs in mental health is still in its infancy. The field is characterized by a variety of applications with a wide range of approaches to how user-CA interactions are designed. While many apps use established therapeutic approaches (e.g., cognitive behavioral therapy, CBT, Hofmann et al., 2012), the psychological processes underlying the relationship between CA design, user engagement, and well-being are not yet fully understood. We want to help close this gap by exploring the psychological mechanisms triggered by design features on user engagement, operationalized as perceived hedonic benefits and usage intention, and perceived stress as a common indicator of mental well-being.
If and how the user relates to the CA they interact with might be crucial for success in terms of user engagement and well-being. Initial research has shown that users can form a “bond” with a CA, similar to a relationship with a human therapist (Darcy et al., 2021). Evidence is strengthening across CAs that users perceive or develop “synthetic relationships” (a novel term coined by Starke et al., 2024) with CAs, for instance, with digital assistants like Alexa (Tschopp et al., 2023). Yet, in mental health, there is limited understanding of whether and how users perceive these synthetic relationships and how they influence positive, sustained, and effective interactions.
Against this background, our primary goal was to investigate authentic user responses to a mental health app, with actual users who are intrinsically motivated. We were able to use the app Betwixt (Betwixt, 2023a) thanks to the developers. Alongside a scripted, non-AI-based CA as a guiding mentor, this app’s unique feature is the game-like narrative design. The app immerses the user in a dream-like fantasy world; there, said CA guides the user through the levels.
Guided by the app’s defining design features, our study seeks to explore two interconnected psychological processes potentially elicited by these features. First, we look at users’ synthetic relationship perception of the CA, applying the human-AI relationship perception framework recently validated by Tschopp et al. (2023) and Tschopp and Sassenberg (2024; based on the Relational Models Theory, RMT, Fiske, 1992). To our knowledge, our study is the first where this novel relational approach is used in a mental health context. Second, we investigate perceptions of narrative transportation: the user’s experience of being fully absorbed into the narrative (Green & Brock, 2000). Narrative transportation has been shown to influence a wide range of attitudes, intentions, and behaviors positively (Thomas & Grigsby, 2024; Van Laer et al., 2012), including in health-related narratives (De Graaf et al., 2016), and across various media technologies (Liang & Hwang, 2023; Schmidt-Kraepelin et al., 2023). However, how it impacts human-CA interaction has not been investigated yet.
Our study aims to contribute to the field of technology and mental health by answering the following research question: How do user-CA relationship perception and narrative transportation influence users’ perceived hedonic benefits, usage intention, and perceived stress?
Background
CAs in Mental Health
According to the World Health Organization (WHO), mental health “is a state of mental well-being that enables people to cope with the stresses of life, realize their own abilities, learn well and work well, and contribute to their community” (World Health Organization, 2022). Thus, mental health is “more than the absence of mental disorders” (World Health Organization, 2022), like depression, anxiety, or eating disorders. The term ‘mental health’ includes the promotion of mental well-being on the one hand as well as the treatment of mental health conditions on the other hand.
Unfortunately, mental health services globally are under-resourced, which contributes to the increasing development and deployment of automated digital tools to address the mental health crisis (Bucci et al., 2019; NIMH, 2024). Mental health apps have gained importance in everyday life, offering users instantly accessible and often cost-effective means to manage and monitor their mental well-being (Bucci et al., 2019; Chandrashekar, 2018). These apps often focus on various strategies for stress reduction as stress is considered one of the primary antecedents of mental ill-being (also physical illness), being associated with depression or anxiety (Hickey et al., 2021).
Various mental health apps are currently available (according to King et al., 2023, major app stores released over 90,000 digital health apps in 2020). They encompass a range of functionalities (King et al., 2023), from guided meditation (e.g., Headspace3) or mood trackers (e.g., Calm4) to interactive systems, often including a CA, based on Cognitive Behavior Therapy exercises (e.g., Woebot5). While usually not designed to replace traditional therapy, they act as supplemental resources for promoting mental well-being. They offer immediate assistance and resources for the self-management of minor mental health problems (Bucci et al., 2019). For example, in a quasi-experimental study, Inkster et al. (2018) showed that users with self-reported depression symptoms who frequently interacted with Wysa, a CA for mental well-being, reported a greater improvement in mood compared to the low usage group. There is also generalized evidence for the effectiveness of CAs in improving mental health outcomes. In a meta-analysis of randomized controlled trials, He et al. (2023) arrived at small to moderate positive short-term effects of CA interventions on depressive and anxiety symptoms and stress, thereby confirming earlier meta-analytic results by Abd-Alrazaq et al. (2020).
While recent integrative work on human-CA relationships emphasizes that such systems can offer meaningful benefits, such as perceived support or companionship (Malfacini, 2025), scholars have raised serious concerns related to human-CA relational interaction. Risks like emotional dependence, displacement of human relationships, loss of autonomy, and social alienation underscore the importance of rigorous risk considerations in almost any context (Malfacini, 2025; Sætra, 2020; Starke et al., 2024).
Use Case: Betwixt, a Chat- and Story-Based Mental Health App
Especially in such a delicate field as mental health, it is crucial for user data to be authentic. Compared to participants in classical surveys who are asked to rate pre-recorded materials of or interact with an app they are not familiar with, active app users may have stronger personal motives to use it and are highly invested in the app’s effectiveness for their own benefit. However, several issues, including reputation concerns of app providers and resource constraints or lack of personal connections, challenge researchers in bridging the gap from research to practice. Thus, we leveraged our own professional network by reaching out to the developer team of the Betwixt app, who granted us permission to survey their user base, which consisted of approximately 150,000 users in 2024.
Betwixt is a mental health app based on established emotion regulation strategies: cognitive reappraisal and self-compassion, both central to CBT and mindfulness-based interventions (Betwixt, 2023b; Hofmann et al., 2012; Sharma & Rush, 2014). It is best described as an interactive, game-based mobile app where users progress in their mental health journey. The app has 11 levels called dreams, each lasting about 30 minutes. In Figure 1, we present a typical interaction for one dream from beginning (preparing the environment) to the end (reflecting on the experience in the built-in journal).
Betwixt incorporates two primary design elements: 1) chatting with a scripted CA and 2) reading a fantasy-like story through which the user is “wandering”. First, users are guided by a mentor-like CA called “the Voice”, which poses questions related to personal growth, fears, and stressors (in written format). To progress, users have to chat with “the Voice” via programmed response buttons or open-ended questions (see Figure 2 and Supplement for screenshot transcripts).
Second, the Betwixt app follows a game-like progression, where users advance from level to level (dreams) in a fantasy world called the “In-Between”, see Figure 3. Each dream is a creative story, exploring immersive fantasy worlds accompanied by level-specific music. Users read the description of the world, how it unfolds, what it looks like, and what the fantasy creatures they encounter look like. The music fits the theme of the current narrative. For example, when the story reads that users are on the top of a cliff looking towards the horizon, wind sounds, and high tones are played. Using headphones intensifies the experience by increasing immersion and reducing distractions.
Figure 1. User Journey Through One Level (= Dream) in the In-Between (= Dream World).

Figure 2. CA in Action: “The Voice” Talks With the User About Fear.

Figure 3. Narrative Transportation in Action: Setting the Stage for Immersion in a
Fantasy World With Mythical Creatures (e.g., Chimeras, See Below).

Betwixt offers an immersive and interactive alternative to traditional mindfulness practices and journaling. Through storytelling, sound, and chat with a CA, the app supports self-exploration and emotional resilience (Betwixt, 2023a, 2023b). Combining immersive, relaxing elements and opportunities for self-distanced reflection aims to promote sustained, active engagement with the narrative (Harmon, 2021).
In summary, Betwixt focuses on reducing stress by utilizing story and chat-based design elements to create a therapeutic environment that promotes mental well-being, which seems unique (examples of popular apps involving CAs for well-being are provided in the Supplement, Table S1). Initial experimental research indicates that the app can reduce stress (Betwixt, 2023b). Our study focuses on the psychological processes potentially triggered by the design and their effects on usage- and wellbeing-related outcomes.
Related Work and Hypothesis Development
Considering the app’s defining design features (i.e., a CA and immersive stories), we aim to explore the psychological processes these design elements might elicit, namely, synthetic relationship perceptions and narrative transportation. Moreover, we investigate if synthetic relationship perception and narrative transportation drive positive, sustained, and effective interactions with the app. We assess how synthetic relationship perceptions and narrative transportation associate with a) perceived hedonic benefits (i.e., how much users like using the app), b) a more future-oriented outcome, namely the intention to use the app further, and finally, c) we assessed perceived stress to evaluate the app’s effectiveness in terms of well-being. These variables are important desirable outcomes in mental health applications (Abd-Alrazaq et al., 2020; Limpanopparat et al., 2024). Enjoyment has been shown to foster intrinsic motivation and sustained interaction with digital systems, including health-related applications, while continued usage intention is considered a key prerequisite for the effectiveness of interactive mental health tools, where benefits typically accrue through repeated engagement over time (Venkatesh et al., 2012; Torous et al., 2018). Positive and sustained interactions lay the foundation for a mental health app to be effective, that is, to improve mental well-being (Torous et al., 2018).
In the following section, we provide an overview of the literature on the perception of synthetic relationships and narrative transportation, from which we derive our hypotheses and research questions.
The Role of Synthetic Relationship Perception
The CA personified in the app by “the Voice”, does “more” than deliver information. Instead, it functions as a conversational partner, simulating reciprocal interactions that mimic human-to-human exchanges. This may lead users to experience a “relationship” with this non-human entity. A novel phenomenon that comes in different forms (e.g., Brandtzaeg et al., 2022; Pentina et al., 2023; Skjuve et al., 2021; Tschopp et al., 2023) but has recently been captured under the concept of synthetic relationships (Starke et al., 2024). The term is defined as “continuing associations between humans and AI tools that interact with one another wherein the AI tool(s) influence(s) humans’ thoughts, feelings and/or actions [emphasis added by authors]” (p. 1834). We argue that this definition can be plausibly extended to our CA (non-AI), as it can fulfill the above criteria (in bold) under the condition that the system is deemed “worthy” of users perceiving it as a relational entity (and this can be independent of the underlying technology). While other studies have used different terms (e.g., human-AI relationships, Tschopp et al., 2023, consumer-machine relationship, Pentina et al., 2023), we found synthetic relationships the most suitable term because in our example, “The Voice”’s inherent goal is to influence humans’ thoughts, feelings and/or actions–arguably in a positive direction.
While social relationships were long considered an exclusively human domain (Guzman & Lewis, 2020), prior research on parasocial relationships with media figures (Horton & Wohl, 1956) and on social surrogates (Gabriel et al., 2016) has already challenged this assumption. Contemporary research on synthetic relationships builds on and overlaps with this work, but extends it by examining interactive, seemingly reciprocal, one-on-one relationships between humans and technological agents (Pentina et al., 2023; Starke et al., 2024). A growing body of research shows that users increasingly perceive and engage with intelligent systems in relational terms, ranging from instrumental assistance to socially and emotionally meaningful bonds (Malfacini, 2025; Skjuve et al., 2021; Ventura et al., 2025). Recent systematic reviews highlight that such consumer-machine relationships are shaped by perceived sociality, responsiveness, and role framing, and have implications for trust, engagement, and well-being (Pentina et al., 2023). This is of particular interest in mental health settings where the synthetic relationship may resemble a therapeutic alliance, where user and chatbot work together towards a mental health-related goal (Tong et al., 2022).
By examining these novel interactions through the lens of established psychological theories and recent empirical findings, we argue that the quality of the bond between users and CAs might play a critical role in shaping user experiences. Research over the past decades has revealed that humans are predisposed to ascribe human-like qualities to technology (Epley et al., 2007; Nass & Moon, 2000). Even simple machines could elicit polite behavior or feelings of familiarity from users (Weizenbaum, 1966). With the advent of natural language-capable CAs, the potential for establishing stronger, reciprocal social relationships with machines (at least in the users’ perception) has grown significantly (Pentina et al., 2023; Starke et al., 2024). Although these relationships lack the full reciprocity of human connections, they can nonetheless invoke similar feelings of trust, empathy, and engagement and even deep attachment, such as friendship and love (Brandtzaeg et al., 2022; Pentina et al., 2023; Skjuve et al., 2021).
To better understand the nature of synthetic relationships, drawing upon frameworks originally developed for human interactions has been shown to be useful. In human-CA interaction, for instance, Social Penetration Theory (Altman & Taylor, 1987 in Skjuve et al., 2021), Bowlby’s Attachment Model (Bowlby, 1979 in Xie & Pentina, 2022), or Knapp’s Staircase model (Knapp, 1978 as cited in Seymour & Van Kleek, 2021) were used to explain development or drivers of synthetic relationship perception. While these perspectives provide valuable insights, they do not capture the multidimensional nature of human-machine relationship perception, a limitation that has long been discussed in research on interpersonal relationships (e.g., Berscheid, 1994; Haslam & Fiske, 1999).
Therefore, we relied on a social cognition perspective to derive our hypothesis: Taking a different approach towards “the other” rather than the process of development over time, studies by Tschopp et al. (2023) and Tschopp and Sassenberg (2024), based on Fiske’s Relational Models Theory (RMT; Fiske, 1992), have recently proven to be insightful. RMT provides insight into how people think about social interactions on multiple dimensions, independent of the partner’s actual social role, like friend or husband. Recent work by Tschopp and colleagues (Tschopp et al., 2023; Tschopp & Sassenberg, 2024) extended these ideas into human–AI interaction. They found that human-CA relationships come in three modes: friend-like (peer bonding), master-servant (authority ranking), and rational-equal (market pricing). Peer bonding can be best described as users perceiving the CA as a companion, engaging in interactions characterized by mutual support. This mode reflects a social bond where the CA is almost like a friend. A user’s core thought would be: "I feel supported and cared for by the CA, regardless of how much I type”.
Authority ranking is prevalent when users perceive a hierarchy. Traditionally, many view the CA as a subordinate tool designed to execute commands. The interaction is, therefore, often characterized as users being the master and the tool the servant. However, in the context of mental health apps, it could also be the other way round, where user’s core thought would be: “This bot is an expert 'Guide,' and I should follow its instructions”.
In market pricing, the relationship is transactional, akin to a business exchange based on currencies. In contrast to the rather emotional perceptions in peer bonding and the hierarchical perceptions in authority ranking, users make sense of their relationship with the CA through notions of efficiency and utility, the interaction being best described as a rational exchange on “eye-level”. Even though no physical money is changing hands during the conversation, the user treats their emotional data, time, and effort as a currency that should buy a specific quality of service. A user’s core thought would be, for instance: “Am I getting a fair 'deal' for the time and data I’m giving this bot?”.
In mental health contexts, the socio-emotional peer bonding mode might be particularly relevant. This mode positions the CA as “someone” who listens, understands, and offers empathetic feedback. Such an interaction mode might be instrumental in fostering an environment where users feel comfortable disclosing personal thoughts and emotions, which is essential for therapeutic engagement and effectiveness.
For example, studies on the friendly, Wall-E-like mental health CA Woebot have demonstrated that users can form meaningful, supportive bonds with these systems, even if the relationship is unidirectional in its emotional reciprocity (Darcy et al., 2021). Such findings support the idea that a perceived peer-like bond can bridge the gap between mere functionality and genuine engagement.
The multidimensional framework provided by Tschopp et al. (2023) underscores that relationship perceptions of CAs do not depend on the CA’s intended design. Thus, while all perceptions might be prevalent, we argue that the friend-like or peer bonding mode might be the most relevant. We argue that high values in peer bonding can influence our suggested outcomes:
a) Higher Perceived Hedonic Benefits:
When users view the CA as a supportive peer, the interaction might be more enjoyable and intrinsically rewarding. The hedonic benefits of using an app, encompassing pleasure and satisfaction, might be amplified by the sociality of a peer-like bond. As users see the CA as friendly and personal, the experience itself might become a source of positive affect. After all, users might be more likely to value an app that brings joy and a sense of connection.
b) Higher Usage Intentions:
The establishment of a friend-like synthetic relationship could also increase usage intentions. Socio-emotional elements might motivate users to return to the app repeatedly. In mental health, where continuous engagement is key to achieving therapeutic outcomes (Torous et al., 2018), a strong socio-emotional bond could catalyze sustained usage. Users who feel the CA is like a caring friend might be more inclined to integrate the app into their daily routines, reinforcing ongoing engagement.
c) Lower Perceived Stress:
Perhaps one of the most compelling arguments for the importance of peer bonding lies in its potential to alleviate stress. The emotional support from a friend-like relationship could provide users with a safe space to process their emotions and feel understood. When the CA is perceived as an empathetic peer, users might have a reduced level of perceived stress. The sense of being cared for and listened to, even by a non-human agent, might have a calming effect on mental well-being.
The literature indicates that the emotional depth of user–CA interactions can influence both the subjective experience of the app and behavioral intentions (Tong et al., 2022). When users feel a strong bond with the CA, they are more likely to enjoy the app (hedonic benefits), continue using it (usage intentions), and experience lower stress due to supportive interaction. Therefore, we propose the following hypothesis:
H1: Higher peer bonding predicts (a) higher perceived hedonic benefits from using the app, (b) higher usage intentions, and (c) lower perceived stress.6
However, it is also plausible that users perceive a hierarchical relationship (i.e., authority ranking), viewing the CA either as akin to a “guru”, a figure of authority capable of providing psychological guidance, or merely as a servant that executes instructions. In such scenarios, users might appreciate a hierarchical, distanced approach, where the CA either provides structured support similar to established therapeutic approaches like CBT (Hofmann et al., 2012) or users can “order” support from the CA as if it were their personal assistant.
Ultimately, some individuals may perceive the CA as an equal rational companion (i.e., market pricing). This lens is probably akin to users who approach their self-help journey more analytically, perceiving the CA as agentic but on an eye level devoid of hierarchies. A relational approach that offers a differentiated view from a social cognition perspective may be crucial for designing mental health apps that effectively address user needs and preferences. Since the framework for understanding relationship perception is still novel, we explore the research question of how the other relationship dimensions and outcome variables associate.
The Role of Narrative Transportation
Whereas the role of synthetic relationships presents a novel issue with scattered knowledge across disciplines to cover, the role of narrative transportation is rather straightforward, presenting consistent empirical evidence over the past two decades (Shen et al., 2015; Thomas & Grigsby, 2024; Van Laer et al., 2012). Narrative transportation is “an experiential state of immersion in which all mental processes are concentrated on the events occurring in the narrative” (Green & Appel, 2024, p. 1). It refers to the cognitive and emotional absorption individuals experience when immersed in a spoken, written, or audiovisual story (Green & Appel, 2024; Green & Brock, 2000). Narrative transportation plays a significant role in how narratives influence persuasion (Green et al., 2004). For instance, it has been shown to lead to changes in attitudes and behaviors in a multitude of contexts, including marketing (Thomas & Grigsby, 2024; Van Laer et al., 2012), health communication (De Graaf et al., 2016), and entertainment games (Sherrick, 2018).
With regard to media technologies, this phenomenon is leveraged to create interactive environments where users become deeply involved with the narrative, enhancing user engagement. This is particularly relevant in applications such as educational tools (Liang & Hwang, 2023; Oh et al., 2020), gamified information systems (Schmidt-Kraepelin et al., 2023), entertainment technology (Green & Clark, 2013), and, more recently, extended reality environments (Watts, 2023). Overall, there is strong empirical support that the user’s immersion can foster a stronger connection to the content, increase enjoyment, and drive meaningful outcomes, such as behavioral change or emotional relief (Green et al., 2004). In comparison to written stories, audiovisual elements provide additional sensory stimuli that may enrich the user experience and, in turn, elicit attitude or behavior change, as previously shown for health narratives (Shen et al., 2015).
Furthermore, as Green and Brock (2000) suggest, characters are central to fiction, and attachment to characters can significantly influence belief changes in narrative-driven stories. For instance, one study showed that the role of the empowering protagonist influenced a positive change in self-related control beliefs (Isberner et al., 2018). Thus, in narratives, a reader’s attachment to a protagonist can shape how persuasive the story is. Research argues that as readers become transported into the story, they may develop a greater liking for sympathetic protagonists, becoming deeply engaged not just with the narrative world but also with the characters in it (Green et al., 2004). In Betwixt, users become the protagonists as they “watch themselves” wander through these dreams, which may positively affect the intended goals, like lowering perceived stress. Put simply, as they imagine themselves calming down in the story, they may also calm down in real life.
Since the gamified story design of Betwixt is a major design element, narrative transportation may play a decisive role. It could enhance the user’s focus on the task or therapy at hand and provide a better user experience. Narrative transportation not only increases engagement but also facilitates a sense of presence, allowing users to momentarily escape their immediate concerns and fully engage with the mental well-being content (Green & Brock, 2000). Thus, we expect the immersive process to influence several outcomes:
a) Higher Perceived Hedonic Benefits
Narrative transportation can transform app interactions into engaging, immersive experiences that are both pleasurable and intrinsically rewarding. As users become absorbed in the narrative, the emotional richness of the experience may enhance their overall enjoyment and satisfaction with the app.
b) Higher Usage Intentions
The immersive quality of narrative transportation fosters a deep connection with the app’s content, which may lead to increased user engagement. By creating memorable and personally joyful experiences, narrative transportation may be a powerful motivator for long-term behavioral intentions.
c) Lower Perceived Stress
Engagement in a compelling narrative may reduce perceived stress through general immersion that temporarily distracts from stressors and by enhancing engagement with therapeutic content. In the context of Betwixt, both processes may plausibly contribute to lower perceived stress. We therefore formulate a broad prediction for perceived stress without distinguishing between these mechanisms at the hypothesis level.
Based on these arguments, we hypothesize:
H2: Higher narrative transportation predicts a) higher perceived hedonic benefits from using the app, b) higher usage intentions, and c) lower perceived stress.
Figure 4 shows the study’s conceptual framework. Furthermore, as the topic and assessment of synthetic relationships is still in its infancy, we conducted further analyses to explore synthetic relationship perception and its associations with narrative transportation:
RQ1: How do the other relationship dimensions, narrative transportation and outcome variables associate?
Figure 4. Hypotheses and Research Question.

Methods
Participants and Procedure
The study has been approved by the ethics committee of the Leibniz-Institut für Wissensmedien (LEK 2024/017), and hypotheses, data collection procedure, measures, and analysis methods were preregistered at https://aspredicted.org/4FS_SNH. Data, analysis code, additional analyses, and the Supplement are accessible on Researchbox (https://researchbox.org/2669). An a-priori power analysis for a small effect (f2 ≥ .02) in a multiple linear regression with two predictors, a power of 1 – β ≥ .80, and α = .05, conducted with G*Power (Faul et al., 2007), yielded a target sample size of N = 311. Participants had to be at least 18 years old. As we aimed for participants who were experienced in using the app, participants had to be at level two or higher to be eligible for the study. Additionally, participants had to pass an attention check placed at the end of the questionnaire.
We recruited Betwixt users via the Betwixt app. Users were invited to click on a link in their app leading to an English Qualtrics questionnaire. After providing information about the study and asking for their consent to participate, they were supposed to answer questions on how they perceive their relationship with “the Voice”, about outcomes related to app usage and well-being, and demographic information. Participants could then provide feedback on the app and our study, and provided consent to the use of their data for scientific purposes. Data were collected from 5 June to 14 August 2024.
Of 441 participants who started the questionnaire, 354 finished it and consented to scientific data use. After excluding five participants who indicated they had not tried their best to follow our instructions and answer the questions (i.e., failed the preregistered attention check), as well as one participant who exhibited zero within-person variance across all measures, consistent with a satisficing response pattern7, we arrived at a final sample size of N = 348. All participants indicated themselves to be 18 or older; the average age was 38.11 years (SD = 12.80, range = 18–76). Of the final sample, 70% identified as female, 18% as male, 10% as non-binary/third gender, and 2% preferred not to say. The majority had a college degree (58%). Participants were from a total of 35 countries. Most participants came from English-speaking countries (US: 47%, UK: 12%, Canada: 12%, Australia: 8%). All participants were at least in level three of the Betwixt app, with most participants currently in levels four or five out of eleven (82%).
Measures
Independent Variables
To measure synthetic relationship perception, we used the Human-AI Relationship Questionnaire (HAIRQ) developed by Tschopp et al. (2023), with minor changes to the wording. The original scale consisted of nine items, e.g., You are like peers or fellow co-partners, measured on seven-point scales ranging from 1 (not at all true for this relationship) to 7 (very true for this relationship). The two other relationship dimensions were originally measured with four items each, i.e., authority ranking (e.g., It feels as though the relationship is hierarchical and one of you is above the other) and market pricing (e.g., When interacting with the Voice, you expect a fair rate of return for the effort you put in). To test the validity of the scale, we conducted a confirmatory factor analysis (CFA), which yielded a poor fit (χ2(116) = 444.551, p < .001, CFI = .759,
RMSEA = .090, 95% CIRMSEA [.081, .099], SRMR = .092, see Supplement for full results). We thus re-conceptualized the scale by removing one market pricing item with a very small factor loading, one item of the authority ranking subscale, and four items of the peer bonding subscale that showed cross-loadings on other factors. The adapted scales yielded improved fit (χ2(41) = 91.777, p < .001, CFI = .937, RMSEA = .060, 95% CIRMSEA [.043, .76], SRMR = .051). All standardized factor loadings of the HAIRQ subscales were positive and significant. The loadings for peer bonding and authority ranking were all > .50; the loadings for market pricing were between .37 and .45. Internal consistency improved for peer bonding (α = .81, decreased slightly for authority ranking (α = .67), and stayed the same for market pricing (α = .40). Due to the low internal consistency of the market pricing subscale, we conducted item-level diagnostics, including item total correlations and analyses of Cronbach’s α if items were deleted. These checks indicated that removing individual items would not meaningfully improve reliability. Following this detailed evaluation of the market pricing subscale, we determined that it lacks sufficient reliability for direct interpretation. However, to maintain conceptual consistency, we have retained it as a control variable in our linear regression model. Sensitivity analyses without market pricing are reported in the Supplement (Tables S4–S6). We address this issue as a limitation in the discussion section.
We had preregistered to measure narrative transportation using an adaptation of the Transportation scale by Green and Brock (2000), comprising eleven items on seven-point Likert-type rating scales. However, a CFA yielded poor fit (χ2(44) = 291.284, p < .001, CFI = .671, RMSEA = .127, 95% CIRMSEA [.113, .141], SRMR = .099). The original transportation scale had been validated in the context of a short fictional story about a murder at a shopping mall (Green & Brock, 2000). In contrast, the present study was conducted in the context of a mental health app featuring an ongoing narrative involving a mentor character (“the Voice”), with narrative content dynamically unfolding in response to users’ inputs rather than following a fixed plot. We thus adapted the scale by focusing on the three items that were most compatible with our application context, i.e., While I am reading the narrative, I can easily picture the events in it taking place; I can picture myself in the scene of the events described in the narrative, and I am mentally involved in the narrative while reading it. We decided to remove the reverse-coded items and items like I find myself thinking of ways the narrative could have turned out differently and I want to learn how the narrative ends due to small factor loadings and limited conceptual fit. The corresponding CFA model was just-identified (df = 0), with all factor loadings statistically significant and > .50. The resulting three-item scale showed improved internal consistency (α = .77).
Dependent Variables
App-related outcomes included perceived hedonic benefits and usage intention, which we measured with established scales. Each was measured using three items from Venkatesh et al. (2012) on seven-point Likert-type rating scales. Perceived hedonic benefits included items like: Using the Betwixt app is fun, and usage intention consisted of items like: I intend to continue using Betwixt in the future. Separate CFA models for perceived hedonic benefits and usage intention were just-identified (df = 0). Standardized factor loadings were all positive (> .5) and significant. Internal consistency was good (hedonic benefits: α = .82, usage intention: α = .69). Perceived stress was assessed using the Perceived Stress Scale (Cohen et al., 1983), with ten items, such as: In the last month, how often have you felt nervous and ‘stressed’?, measured on five-point scales from 1 (never) to 5 (very often). The CFA for perceived stress yielded moderate fit (χ2(35) = 170.293, p < .001, CFI = .912, RMSEA = .105, 95% CIRMSEA [.090, .121], SRMR = .057). We decided to remove two of the reverse-coded items with small factor loadings, leading to an improved fit (χ2(20) = 73.373, p < .001, CFI = .957, RMSEA = .088, 95% CIRMSEA [.067, .109], SRMR = .043). Standardized factor loadings of the 8-item CFA model were all positive (> .5) and significant. Internal consistency was good (α = .89).
A complete list of the original and final items can be found in Appendix A. Univariate distributions are provided in Appendix B.
Results
We conducted all analyses in SPSS 28 (IBM Corp, 2021). Descriptive statistics and bivariate correlations among all study variables are given in Table 1.
Table 1. Descriptive Statistics and Correlations for Study Variables and Demographics.
|
Variable |
M/% |
SD |
1 |
2 |
3 |
4 |
5 |
6 |
7 |
8 |
9 |
10 |
11 |
|
Peer bonding |
4.68 |
1.28 |
— |
|
|
|
|
|
|
|
|
|
|
|
Authority ranking |
3.13 |
1.38 |
−.29*** |
— |
|
|
|
|
|
|
|
|
|
|
Market pricing |
5.30 |
1.06 |
.26*** |
.01 |
— |
|
|
|
|
|
|
|
|
|
Narrative transportation |
6.13 |
0.81 |
.38*** |
−.14** |
.16** |
— |
|
|
|
|
|
|
|
|
Perceived hedonic benefits |
6.02 |
0.86 |
.41*** |
−.17** |
.19*** |
.41*** |
— |
|
|
|
|
|
|
|
Usage intention |
5.91 |
0.79 |
.42*** |
−.27*** |
.15** |
.37*** |
.52*** |
— |
|
|
|
|
|
|
Perceived stress |
3.48 |
0.75 |
.03 |
−.08 |
−.03 |
.02 |
.07 |
.17*** |
— |
|
|
|
|
|
Age |
38.11 |
12.80 |
−.07 |
0.07 |
.14** |
−.05 |
−.21*** |
−.03 |
−.21*** |
— |
|
|
|
|
Gender |
70.00 |
— |
.003 |
.004 |
.02* |
.03 |
.001 |
.004 |
.01 |
.04*** |
— |
|
|
|
Education |
4.16 |
1.59 |
−.16** |
.04 |
.02 |
−.12* |
−.12* |
−.14** |
−.15** |
.28*** |
.03* |
— |
|
|
App level |
4.87 |
1.31 |
−.02 |
.12* |
.07 |
.08 |
.02 |
−.15** |
−.02 |
.02 |
.01 |
−.06 |
— |
|
Note. Pearson correlations. N = 348. All scales used 7-point Likert scales, except perceived stress (5-point scale). For gender, included as a categorical variable: male (1), female (2), non-binary/third gender (3), the percentage of female participants and η² values from one-way ANOVAs are reported. Education was included as a quasi-continuous variable: less than high school (1), high school graduate (2), some college (3), 2-year degree (4), 4-year degree (5), professional degree (6), and doctorate or above (7). App level ranges from 3–11. *p < .05, **p < .01, *** p < .001. |
|||||||||||||
The mean values indicated that users perceived the relationship more in terms of peer bonding and less in terms of authority ranking. The average level of narrative transportation was high. The mean values of the outcome variables indicated that the users seemed to like the experience, had high future use intentions, and were, on average, sometimes to fairly often stressed in the last month.
Peer bonding and narrative transportation were moderately positively correlated with each other, and both constructs were likewise positively correlated with the app-related outcomes. Hedonic benefits and usage intention correlated highly with each other. Authority ranking was negatively correlated with peer bonding and the app-related outcomes. Perceived stress was not significantly correlated with any variable, except weakly positively with usage intention.
The user-CA relationship dimensions are relatively new concepts, and prior research suggested associations with demographic variables (Tschopp et al., 2023). We thus looked at potential boundary conditions by including age, gender, education, and users’ current app level in the correlation table. Like in prior studies, peer bonding was negatively correlated with education (Tschopp et al., 2023), and so was narrative transportation. Whether people intend to use the app in the future was negatively associated with education and app level. Interestingly, perceived hedonic benefits and perceived stress were negatively correlated with age and education. One-way analyses of variance (ANOVA) with gender as independent and our study variables as dependent variables only yielded small significant effects on age (F(2, 338) = 7.09, p < .001) and education (F(3, 338) = 4.09, p = .018). Post-hoc tests only indicated that female participants were older than non-binary/third-gender participants (p < .001). Post-hoc tests for education only suggested that female participants were more highly educated than male participants (p = .033) (see Supplement, Table S3, for full ANOVA results).
Table 2. Regression Coefficients of Relationship Dimensions and Narrative Transportation on Perceived Hedonic Benefits.
|
Variable |
Model 1 |
Model 2 |
Model 3 |
|||||||||
|
B |
SE |
β |
p |
B |
SE |
β |
p |
B |
SE |
β |
p |
|
|
Constant |
3.17 |
.31 |
|
< .001 |
3.10 |
.40 |
|
< .001 |
3.59 |
.40 |
|
< .001 |
|
Peer bonding |
0.20 |
.03 |
.30 |
< .001 |
0.18 |
.04 |
.27 |
< .001 |
0.17 |
.04 |
.26 |
< .001 |
|
Narrative transportation |
0.31 |
.05 |
.29 |
< .001 |
0.30 |
.05 |
.29 |
< .001 |
0.29 |
.05 |
.28 |
< .001 |
|
Authority ranking |
|
|
|
|
−0.03 |
.03 |
−.05 |
.323 |
−0.03 |
.03 |
−.05 |
.303 |
|
Market pricing |
|
|
|
|
0.06 |
.04 |
.07 |
.148 |
0.09 |
.04 |
.11 |
.034 |
|
Age |
|
|
|
|
|
|
|
|
−0.01 |
.003 |
−.20 |
< .001 |
|
Gender |
|
|
|
|
|
|
|
|
|
|
|
|
|
Female |
|
|
|
|
|
|
|
|
−0.01 |
.11 |
−.01 |
.903 |
|
Non-binary / third |
|
|
|
|
|
|
|
|
−0.18 |
.16 |
−.06 |
.272 |
|
Education |
|
|
|
|
|
|
|
|
0.000 |
.03 |
.000 |
.994 |
|
Adjusted R2 |
.24 |
|
|
|
.24 |
|
|
|
.27 |
|
|
|
|
Note. N = 348. Reference category for gender = male. |
||||||||||||
Table 3. Regression Coefficients of Relationship Dimensions and Narrative Transportation on Usage Intention.
|
Variable |
Model 1 |
Model 2 |
Model 3 |
|||||||||
|
B |
SE |
β |
p |
B |
SE |
β |
p |
B |
SE |
β |
p |
|
|
Constant |
3.52 |
.29 |
|
< .001 |
3.85 |
.34 |
|
< .001 |
3.97 |
.37 |
|
< .001 |
|
Peer bonding |
0.20 |
.03 |
.33 |
< .001 |
0.17 |
.03 |
.28 |
< .001 |
0.17 |
.03 |
.27 |
< .001 |
|
Narrative transportation |
0.23 |
.05 |
.24 |
< .001 |
0.23 |
.05 |
.23 |
< .001 |
0.22 |
.05 |
.23 |
< .001 |
|
Authority ranking |
|
|
|
|
−0.09 |
.03 |
−.16 |
.001 |
−0.09 |
.03 |
−.16 |
.001 |
|
Market pricing |
|
|
|
|
0.03 |
.04 |
.04 |
.382 |
0.09 |
.04 |
.11 |
.482 |
|
Age |
|
|
|
|
|
|
|
|
0.002 |
.003 |
.03 |
.514 |
|
Gender |
|
|
|
|
|
|
|
|
|
|
|
|
|
Female |
|
|
|
|
|
|
|
|
0.08 |
.10 |
.05 |
.401 |
|
Non-binary / third |
|
|
|
|
|
|
|
|
−0.01 |
.15 |
−.01 |
.927 |
|
Education |
|
|
|
|
|
|
|
|
−0.04 |
.03 |
−.08 |
.109 |
|
Adjusted R2 |
.22 |
|
|
|
.24 |
|
|
|
.24 |
|
|
|
|
Note. N = 348. Reference category for gender = male. |
||||||||||||
Table 4. Regression Coefficients of Relationship Dimensions and Narrative Transportation on Perceived Stress.
|
Variable |
Model 1 |
Model 2 |
Model 3 |
|||||||||
|
B |
SE |
β |
p |
B |
SE |
β |
p |
B |
SE |
β |
p |
|
|
Constant |
3.36 |
.31 |
|
< .001 |
3.65 |
.37 |
|
< .001 |
4.19 |
.40 |
|
< .001 |
|
Peer bonding |
0.02 |
.03 |
.03 |
.623 |
0.01 |
.04 |
.02 |
.808 |
−0.01 |
.04 |
−.01 |
.858 |
|
Narrative transportation |
0.01 |
.05 |
.01 |
.890 |
0.01 |
.05 |
.01 |
.902 |
−0.01 |
.05 |
−.01 |
.846 |
|
Authority ranking |
|
|
|
|
−0.04 |
.03 |
−.08 |
.171 |
−0.03 |
.03 |
−.06 |
.279 |
|
Market pricing |
|
|
|
|
−0.02 |
.04 |
−.03 |
.566 |
−0.002 |
.04 |
−.003 |
.963 |
|
Age |
|
|
|
|
|
|
|
|
−0.01 |
.003 |
−.17 |
.003 |
|
Gender |
|
|
|
|
|
|
|
|
|
|
|
|
|
Female |
|
|
|
|
|
|
|
|
0.06 |
.11 |
.04 |
.573 |
|
Non-binary / third |
|
|
|
|
|
|
|
|
0.18 |
.16 |
.07 |
.264 |
|
Education |
|
|
|
|
|
|
|
|
−0.04 |
.03 |
−.09 |
.111 |
|
Adjusted R2 |
−.01 |
|
|
|
−.004 |
|
|
|
.03 |
|
|
|
|
Note. N = 348. Reference category for gender = male. |
||||||||||||
To test H1 and H2 and to answer RQ1, we ran three multiple linear regression models per outcome variable. The first model included peer bonding and narrative transportation as predictors; the second model additionally included authority ranking, and market pricing as a control variable. Because we found age, gender, and education to be correlated with some of the study variables, we included them as demographic control variables in the third model. The full regression results can be found in Tables 2–4. Regression diagnostics as well as sensitivity analyses using the preregistered scales and all participants are available in the Supplement (Tables S10–S12). The pattern of results in the sensitivity analyses did not differ from the one reported here.
In the first models, the predictors explained 24% of the variance in a) hedonic benefits and 22% in b) usage intention. This remained unchanged for hedonic benefits but increased to 24% for usage intention in Model 2, while controlling for demographic variables in Model 3 only improved explained variance for usage intention. Models 1 and 2 did not meaningfully explain perceived stress, as reflected by negative adjusted R2 values, while explained variance increased slightly with the inclusion of demographic controls (Model 3).
The analyses yielded significant positive effects of peer bonding (H1) and narrative transportation (H2) on a) perceived hedonic benefits and b) usage intention. The effects held when controlling for market pricing (Model 2), age, gender, and education (Model 3). We could thus confirm H1a, b and H2a, b. No significant effects on perceived stress emerged. Thus, H1c and H2c were rejected.
Authority ranking (RQ1), did not significantly affect a) perceived hedonic benefits and c) perceived stress. However, as suggested by the correlation analysis, we found a very small significant negative effect of authority ranking on b) usage intention. The control variables market pricing, gender, and education did not have significant effects on the outcomes; only age had small significant negative effects on perceived hedonic benefits and perceived stress, as already indicated by the correlation results (Table 1).
Figure 5. Linear Regression Results.

Note. b = unstandardized regression coefficients, CI = confidence interval. Results based on Model 3, controlling for market pricing, age, gender, and education.
We considered it plausible that stress levels would decrease if users had spent more time and gained more experience. In an exploratory analysis, we used app level as a proxy variable for app experience, where lower levels signified shorter usage, while higher levels represented longer experience of use and thus potentially less perceived stress. However, as shown in Table 1, the bivariate correlation between app level and perceived stress was not significant (r = −.02, p = .718), which ruled out further analyses.
In summary, peer bonding with “the Voice” and narrative transportation predicted users’ app-related outcomes (i.e., perceived hedonic benefits and usage intention). While Table 1 shows that authority ranking was significantly negatively correlated with the app-related outcomes, this significant relationship became smaller after controlling for peer bonding in the linear regression model (Table 3). Perceived hedonic benefits and perceived stress were not influenced by authority ranking (Table 4).
Discussion
We conducted a study with N = 348 users of the CA and story-based mental health app Betwixt. In this study, we explored the association between synthetic relationship perception with the CA they are chatting with and narrative transportation (i.e., how immersed they are in the story), and usage-relevant outcomes as well as perceived stress.
The CA as a Friend Matters
We found that a strong perception of a friend-like relationship with the CA guiding users through the levels positively influenced perceived hedonic benefits and usage intention, partially supporting H1. This finding suggests that users who perceived “the Voice” as friend-like (i.e., high in peer bonding) perceived the app as more enjoyable and more beneficial. It is possible that through increased socio-emotional elements in their attachment to the CA, users might feel safer and are more emotionally engaged. Our findings extend previous research demonstrating positive effects of synthetic relationships characterized as peer bonding in commercial contexts (Tschopp & Sassenberg, 2024), now to a mental health context.
As suggested by the descriptive data, the perception of “the Voice” as a friend was rather high, in contrast to prior studies with digital assistants like Alexa and Siri (cf. Tschopp & Sassenberg, 2024). This finding makes sense and lends further empirical support to the recently developed instrument, which was previously tested on rather functional assistants (Tschopp et al., 2023; Tschopp & Sassenberg, 2024). While these possess some human-like social cues (e.g., human-like voices and speech styles), they are ultimately rather task-oriented with short command and response dialogues. “The Voice” as the CA in the app remains personal, characterized by socio-emotional elements displaying messages of empathy and interest. In addition, compared to voice assistants like Alexa and Siri, communication with the CA in the app resembles a back-and-forth conversation, allowing for longer and deeper interaction.
While the hypothesis testing provided valuable insights, the exploratory results for RQ1, particularly the role of market pricing, remain inconclusive. Several attempts to improve its psychometric properties were unsuccessful. Prior research has noted minor issues with this scale, such as the omission of Item 17 in Tschopp and Sassenberg (2024), but the significant lack of reliability in the current study, hindering a meaningful interpretation of the results, warrants careful, future investigation. In the mental health context, users may lack a clear "price tag" for such digital interactions (typical for market pricing), rendering the transactional logic of this subscale less applicable to their experience with a mental health CA in contrast to the experience with functional digital assistants like Alexa. Consequently, future research must carefully evaluate the dimensional structure of the proposed human-AI relationship framework, i.e., the properties of the items and subscales when applying them to other technological contexts than digital assistants.
Authority ranking was low in prevalence and predictive power, indicating that power differences are likely not involved. This is interesting because some users may prefer a rather non-emotional, more distant relationship with the CA. In such cases, users may even desire a more authoritative relationship, treating the CA as a “Guru-like” helper that users not only talk to but whose lead they follow. Servant- and tool-like CA perceptions were shown to be relevant in commercial contexts (Chen et al., 2019; Tassiello et al., 2021; Tschopp & Sassenberg, 2024). It is up to future research with different CAs in different app designs to determine whether a hierarchical distance in the synthetic relationship perception would be more relevant—and—whether it is more design-dependent or user-dependent. Thus, a peer bonding perception might facilitate the creation of an environment characterized by rather positive emotions. Instead of the general notion of machines being impersonal and cold, peer bonding may be particularly helpful for individuals who lack people they are close with or with whom they can discuss their problems openly without fear of judgment (see the literature on the machine heuristic, Sundar & Kim, 2019). Notably, in more clinical mental health contexts, research suggests that individuals are even inclined to disclose more mental health symptoms to a virtual interviewer than to a human (Lucas et al., 2017).
However, while the data suggest users enjoy the app, the ‘therapeutic’ effects remain unclear. We expected that the more users develop a friend-like synthetic relationship with “the Voice”, the lower their perceived stress. In other words, the more users look at CA from an analytical, distanced point of view, the lower their well-being. However, no significant effects of these psychological processes on perceived stress emerged. Contrary to our hypotheses, friend-like relationship perception was not significantly related to perceived stress. Betwixt’s primary purpose is to support mental well-being through the general assumption that using the app is associated with feeling better overall. The null effects observed for perceived stress may have several explanations. First, the sample consisted of active users, which likely introduces selection and timing effects: users who benefited may have already experienced changes earlier in their usage trajectory, whereas users who did not benefit may have discontinued use, reducing variance in cross-sectional analyses. Second, peer bonding with the CA and narrative transportation may primarily influence momentary relief rather than perceptions of stress over the past month. In this sense, hedonic benefits and friend-like CA perceptions may function as buffering processes that support coping in the moment without necessarily reducing overall stress levels driven by external life demands. More broadly, perceived stress does not reflect broader conceptions of well-being, which may help explain the null findings.
‘Losing’ Oneself in the Story Matters
Narrative transportation positively influenced users’ perceived hedonic benefits and usage intention, partially supporting H2. The result suggests that a higher perceived transportation of the app creates a sense of presence that allows the user to focus solely on the task at hand and immerse themselves in the story. Envision users fully immersed in a narrative, headphones in place and music enveloping them as they navigate a virtual landscape, oblivious to their actual surroundings. They focus entirely on the app, the music, and the task at hand, leaving minimal room for intrusive thoughts of work, relationships, or other stressors. In contrast, when utilizing a meditation app that relies solely on music, users may find their minds more susceptible to wandering back to these pressures. Our results indicate the important role of narrative transportation as it offers a meaningful escape from everyday challenges, allowing users to momentarily release worries and engage with the experience, resulting in higher enjoyment and an increased intention to use the app.
The positive effects align with previous research on the beneficial effects of narrative transportation on outcomes in digital contexts, for example, in gamified information systems (Schmidt-Kraepelin et al., 2023). We showed that narrative transportation is also effective in a mental health context when it comes to two variables highly relevant for evaluating the success of such apps.
However, as with synthetic relationship perception, the data showed no association with perceived stress. Therefore, we conducted additional exploratory analyses to uncover the role of app usage in perceived stress. Due to the exploratory, cross-sectional design, we used users’ app progress as a proxy for app usage. Results did not reveal significant correlations with perceived stress levels. Moreover, the observed positive association between perceived stress and usage intention suggests a potential bidirectional effect. On the one hand, higher stress could drive a greater intention to use the app as users seek support; on the other hand, higher usage intention might lead to more frequent engagement, which in some cases could reduce stress. These opposing effects may have canceled each other out.
It will be crucial to better understand the relationship between the predictor variables and usage patterns and what is arguably most important: positive mental health outcomes. As we have measured positive contributions from CA relationship and narrative transportation on engagement, yet without any significant effects on stress, we need a better understanding of when and how immersion and friend-like CA perceptions can be conducive to improved mental health. A key risk related to apps emphasizing immersion and fostering friend-like synthetic relationships is that they are turned into enjoyable games more than effective therapeutic interventions. If transportation is achieved without any positive mental health effects, the app risks becoming a distraction from effective interventions and could even end up being detrimental to mental health if it leads to users spending an undue amount of time in the app due to its immersive nature. The key question will be how to generate engagement and immersion for those who experience benefits from the app while guiding others, particularly those who might experience negative effects, towards other activities and interventions.
Theoretical Implications
Our findings advance theory on synthetic relationship perception by extending the RMT to a mental health context involving a non-generative, scripted CA. While the Human–AI Relationship Questionnaire (Tschopp et al., 2023) has been validated across various traditional conversational AI settings, our results indicate that the instrument cannot be fully adapted to specialized therapeutic contexts without further refinement.
Prior work applying the scale in utilitarian voice shopping contexts with Alexa has emphasized authority ranking and market pricing as dominant relational modes, whereas peer bonding, despite lower prevalence, emerged as a strong predictor of shopping behavior (Tschopp & Sassenberg, 2024). In the present mental health context, peer bonding likewise emerged as a consistent predictor of engagement-related outcomes, while authority-based (and transactional) relationship perceptions played a limited role.
While our study yields inconclusive results regarding a broad multidimensional view of synthetic relationship perception, it provides valuable insights into how socio-emotional modes of perception become particularly consequential in this context. These findings suggest that while RMT may lay the conceptual groundwork, its psychometric properties remain unclear; thus, further research is warranted to investigate these relationships across different settings.
This study contributes to research on immersive applications by successfully applying the concept of narrative transportation in a mental health technology. While prior work has predominantly conceptualized narrative transportation as a pathway to attitude or belief change in persuasive contexts, our findings demonstrate that narrative immersion in a reflective, non-persuasive context predicts enjoyment and sustained engagement. At the same time, our research underscores the importance of careful operationalization when applying narrative transportation measures in novel contexts. Transportation scales have been adapted across diverse domains and formats since their original development for story-based and immersive media, often with modifications to items, factor structures, and measurement models (Appel et al., 2015). Future research should pay close attention to scale selection, item wording, and construct validity when studying narrative transportation in different applied settings.8
Finally, this study contributes to mental health app research by conceptually disentangling experience-related outcomes from mental well-being outcomes. Our findings show that friend-like CA perceptions and narrative transportation predict enjoyment and continued use, yet do not necessarily translate into changes in global stress perceptions. Engagement should be understood as a distinct and theoretically relevant outcome in its own right, as well as a potential enabling condition for downstream well-being effects, rather than a direct indicator of mental health improvement.
Practical Implications for Responsible Design
Our findings highlight that scholars and practitioners must pay close attention to both design elements and the psychological processes that these elements elicit. Our results are particularly pertinent for those seeking to implement interventions that combine storytelling with chat functionalities. We contend that design features and psychological processes are interdependent and should be selected in relation to one another. Based on our findings, we discuss key aspects for the ethically responsible design of mental health apps, particularly those based on using CAs and game- or story-like elements for greater immersion.
Any use of a CA taking on a role in a mental health app will raise questions about how the users perceive and what kind of synthetic relationship they perceive. Particularly in an immersive game marketed as a mental health app, there is a greater risk of having vulnerable users who might seek a deep and therapeutic relationship with mentor-like entities in the app. While this could be good for encouraging a working alliance-like relationship (Tong et al., 2022), it also makes the user highly vulnerable to various forms of deception, dependence, and influence from the CA. If the CA is designed in a way that supports or encourages such a relationship, the user will, for example, be likely to disclose highly sensitive information, leading to greater vulnerabilities related to data gathering, and subsequent sale or theft, or personal data gathered by the app. Furthermore, if the user relies on the CA, it will be crucial for the developers to have full control over the conversations and what kind of advice the CA gives, including how the data is used. As already mentioned, there have already been cases of suicides linked to unfortunate interactions with human-like CAs (Davis, 2023; Roose, 2024).
In light of the proliferation of AI, specifically LLM-based chatbots and mental health tools, using a scripted CA represents a wise security choice, which however, may also involve clear trade-offs. Using non-AI or simple AI offers high consistency and more control over therapeutic content, as the risk of “hallucinations” and generating harmful output by LLM-based CAs is not solved (Starke et al., 2024). However, simple CAs may risk generating overly generic responses to open-ended user input, limiting personalization and adaptivity. Such responses allow users to interpret them broadly, which, of course, developers cannot fully manage. This also may lead to unexpected interpretations, especially from vulnerable users. Accordingly, the present findings primarily generalize to controlled, story-based mental health applications using scripted mentor-like CAs and may not directly extend to highly adaptive, AI-based systems.
From an affordance perspective, the implications are not limited to Betwixt as a specific app but extend to platforms that share similar affordances, particularly narrative immersion and conversational guidance. Consistent with the MAIN model (Sundar, 2008), such affordances can activate heuristic processing, shaping engagement and relational perceptions rather than direct persuasive or clinical outcomes. In this light, our findings suggest that platforms affording sustained narrative immersion and socio-emotional interaction with a conversational guide may benefit from design choices that support peer-like relationship perceptions rather than hierarchical or purely instrumental ones. Moreover, affordances that enable focused experiential engagement, such as structured story progression or audio-based environments, may foster enjoyment and continued use, which is still a desirable outcome even in the absence of measurable reductions in global stress.
Strengths, Limitations, and Future Research
Given the limited research on how perceptions of the user-CA relationship influence user experience and well-being in mental health apps, we conducted an exploratory cross-sectional study that also examined the role of narrative transportation.
Having access to an authentic user base is a major strength of the study. We also provided more empirical knowledge on the novel topic of synthetic relationships and validated this framework in a different context. The multidimensional approach to synthetic relationships provides differentiated results: in other words, how—not that—the user relates to the CA matters. Finally, we provide evidence for the positive effect of narrative transportation that practitioners of similar apps can build upon and leverage.
However, the cross-sectional design only captures a moment in time, while perceived stress and relationship perception are potentially time-sensitive variables. A longitudinal design would serve various issues: It would allow for tracking intrapersonal developments regarding the relationship development between the user and the CA, as well as the temporal alignment between the intervention and the stress measure. Betwixt’s narrative “dreams” are short and episodic, and users may engage with them intermittently, whereas the perceived stress scale assesses stress retrospectively over the past month. Accordingly, the absence of associations with perceived stress should not be interpreted as evidence that the app has no stress-related effects. Future research would benefit from longitudinal designs that allow tracking intrapersonal developments over time and from including additional well-being-related outcomes that more comprehensively capture mental well-being and allow for momentary and broader effects of app use. Additionally, qualitative investigations, such as in-depth interviews, could offer richer insights into the processes underlying users’ experiences of stress and overall well‐being.
This study was based on data from active Betwixt users, providing high internal validity. While we aimed for authentic responses from real app users, we are aware of potential biases introduced through selection effects (e.g., only highly motivated users would decide to participate in the study). Our findings are not representative of the general population, and it is unclear whether they can be generalized to other mental health apps.
The CA in the app we studied did not rely on generative AI and large language models (LLM). Despite the proclaimed potential of generative AI in the mental health context, strong evidence is missing. A move in this direction would entail that developers must relinquish much of the desirable control over the potential risks related to human-CA interaction, where synthetic relationships evolve. Thorough testing before being introduced in a context where vulnerable users seek guidance is critical. However, even if the app is thoroughly tested or verified, e.g., by a health institution, using chatbots for emotional purposes like coping with mental ill-being or worse, may introduce risks in ways not intended and hard to predict (Starke et al., 2024).
Conclusion
We investigated how users interact with a CA in a story-based mental health app, Betwixt. This study addresses the mixed evidence surrounding CAs in mental health apps amid the global mental health crisis. Our findings indicate that perceiving the CA as friend-like and high narrative transportation were significant predictors of key usage behaviors but not of perceived stress. The impact of these factors on mental well-being outcomes remains unclear and offers exciting avenues for future research. By reflecting on the benefits and challenges, we hope to contribute to promoting safer, more effective digital mental health interventions and improving access to support.
Footnotes
1 Replika AI—the AI companion who cares: https://replika.com/
2 Wysa—Everyday Mental Health: https://www.wysa.com/
3 App to practice mindfulness to improve stress or sleep. https://www.headspace.com/de/about-us
4 App to track mood and emotions to take care of mental health. www.calm.com
5 App that provides chat-based mental health support. www.woebothealth.com
6 Minor change to the preregistration, see Supplement (order of hypotheses).
7 Sensitivity analyses including this participant yielded slightly poorer fit for the confirmatory factor analysis (CFA) models (see https://researchbox.org/2669) but did not alter the regression results (see Supplement Tables S7–S8).
8 We thank the editor and the reviewers for their valuable suggestions on this topic.
Conflict of Interest
Christine Anderl is employed by Endo Health GmbH, a company specializing in digital treatment programs for women’s health. Her contributions to this manuscript were completed prior to this employment, and she has no other conflicts of interest to disclose.
Use of AI Services
The authors declare they have used AI services, specifically Grammarly, DeepL, and ChatGPT, for grammar correction and style refinements. They carefully reviewed all suggestions from these services to ensure the original meaning and factual accuracy were preserved.
Data Availability Statement
All data are available on ResearchBox: https://researchbox.org/2669.
Acknowledgement
We thank Hazel Gale and Ellie Dee (Co-CEOs, Betwixt) for allowing us to survey Betwixt users.
Appendices
Appendix A. List of Items
Narrative Transportation, Green and Brock (2000)
Please rate how much these statements are true when entering the Betwixt app. The narrative refers to the storyline, the environment created or displayed for you in the Betwixt app including your written interactions.
- While I am reading the narrative, I can easily picture the events taking place.*
- While I am reading the narrative, activity going on in the room around me is on my mind.
- I can picture myself in the scene of the events described in the narrative.*
- I am mentally involved in the narrative while reading it.*
- After finishing the narrative, I find it easy to put it out of my mind.
- I want to learn how the narrative ends.
- The narrative affects me emotionally.
- I find myself thinking of ways the narrative could turn out differently.
- I find my mind wandering while reading the narrative.
- The events in the narrative are relevant to my everyday life.
- The events in the narrative have changed my life.
(measured on Likert-type rating scales from (1) Strongly disagree to (7) Strongly agree)
*final items
Human-AI Relationship Questionnaire (HAIRQ), Tschopp et al. (2023)
In the Betwixt app you are chatting with the “Voice”, guiding you through these dreams. Think about the Voice and your conversations with it in your Betwixt app. Please rate how much these statements describe how you see your relationship with the Voice.
- Within your relationship with the Voice, it feels as though there is a moral obligation to act kindly toward one another.
- Decisions are made together.*
- You tend to adopt similar attitudes and behaviours.*
- It feels as though you and the Voice have something special in common.*
- You two are like a unit: you belong together.*
- Your relationship with the Voice is reciprocal: when you do something, you expect something similar in return.
- Both you and the Voice have an equal say when a decision is made.
- You take turns doing what the other wants.
- You are like peers or fellow co-partners.*
- It feels like one of you is entitled to more than the other.*
- One directs the work, while the other pretty much follows.*
- Together, you are like leader and follower.
- It feels as though the relationship is hierarchical and one of you is above the other.*
- What you get from your interactions with the Voice is directly proportional to how much you give.*
- When interacting with the Voice, you expect a fair rate of return for the effort you put in.*
- You expect the Voice to treat you the same as it would anyone else.*
- Your relationship with the Voice is rather rational, where you are weighing benefits against costs (or other investments, like time spent on the app).
(measured on 7-point Likert-type rating scales from (1) Not at all true for this relationship to (7) Very true for this relationship)
*final items
Perceived Hedonic Benefits, Venkatesh et al. (2012)
Think about your past experiences with the Betwixt app. Please rate how much these statements describe your experiences while using the app.
- Using the Betwixt app is fun.
- Using the Betwixt app is enjoyable.
- Using the Betwixt app is very entertaining.
(measured on 7-point Likert-type rating scales from (1) Strongly disagree to (7) Strongly agree)
Usage Intention, Venkatesh et al. (2012)
Please rate how much you agree with these statements.
- I intend to continue using Betwixt in the future.
- I will always try to use Betwixt in my daily life.
- I plan to continue to use Betwixt frequently.
- I would recommend Betwixt to others.
- I have already recommended Betwixt to others.
(measured on 7-point Likert-type rating scales from (1) Strongly disagree to (7) Strongly agree)
Perceived Stress, Cohen et al. (1983)
The questions in this scale ask you about your feelings and thoughts during the last month. In each case, you will be asked to indicate how often you felt or thought a certain way.
- In the last month, how often have you been upset because of something that happened unexpectedly?*
- In the last month, how often have you felt that you were unable to control the important things in your life?*
- In the last month, how often have you felt nervous and “stressed”?*
- In the last month, how often have you felt confident about your ability to handle your personal problems?*
- In the last month, how often have you felt that things were going your way?
- In the last month, how often have you found that you could not cope with all the things that you had to do?*
- In the last month, how often have you been able to control irritations in your life?
- In the last month, how often have you felt that you were on top of things?*
- In the last month, how often have you been angered because of things that were outside of your control?*
- In the last month, how often have you felt difficulties were piling up so high that you could not overcome them?*
(measured on five-point scales from (1) Never to (5) Very often)
*final items
Appendix B
Figure B1. Univariate Distributions of Study Variables.

Note. Violin plots with overlaid boxplots and mean indicators illustrating the distributions of the study variables. The violin displays the kernel density, the boxplot represents the median and interquartile range, and the orange square denotes the mean value. All variables were measured on scales ranging from 1 to 7, except perceived stress (1 to 5).

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
Copyright © 2026 Marisa Tschopp, Stefanie H. Klein, Christine Anderl, Henrik Sætra , Sonja Utz
