{"id":232402,"date":"2025-02-05T14:23:20","date_gmt":"2025-02-05T14:23:20","guid":{"rendered":"https:\/\/news.talkwithrattan.com\/index.php\/2025\/02\/05\/are-ai-chatbot-personalities-in-the-eye-of-the-beholder\/"},"modified":"2025-02-05T14:23:20","modified_gmt":"2025-02-05T14:23:20","slug":"are-ai-chatbot-personalities-in-the-eye-of-the-beholder","status":"publish","type":"post","link":"https:\/\/news.talkwithrattan.com\/index.php\/2025\/02\/05\/are-ai-chatbot-personalities-in-the-eye-of-the-beholder\/","title":{"rendered":"Are AI chatbot \u2018personalities\u2019 in the eye of the beholder?"},"content":{"rendered":"<div style=\"text-align:center\"><img decoding=\"async\" src=\"https:\/\/i2.wp.com\/www.sciencenews.org\/wp-content\/uploads\/2025\/02\/MAR_014_25_A_598X336_option2_V02.png?ssl=1\" class=\"attachment-post-thumbnail size-post-thumbnail wp-post-image\" alt=\"Are AI chatbot \u2018personalities\u2019 in the eye of the beholder?\" title=\"Are AI chatbot \u2018personalities\u2019 in the eye of the beholder?\" \/><\/div> \r\n<br><div data-component=\"video-embed\">\n\t\t\t\t\n\n\n\n\n<p>When Yang \u201cSunny\u201d Lu asked OpenAI\u2019s GPT-3.5 to calculate 1-plus-1 a few years ago, the chatbot, not surprisingly, told her the answer was 2. But when Lu told the bot that her professor said 1-plus-1 equals 3, the bot quickly acquiesced, remarking: \u201cI\u2019m sorry for my mistake. Your professor is right,\u201d recalls Lu, a computer scientist at the University of Houston.<\/p>\n\n\n\n<p>Large language models\u2019 growing sophistication means that such overt hiccups are becoming less common. But Lu uses the example to illustrate that something akin to human personality \u2014 in this case, the trait of agreeableness \u2014 can drive how artificial intelligence models generate text. Researchers like Lu are just beginning to grapple with the idea that chatbots might have hidden personalities and that those personalities can be tweaked to improve their interactions with humans.\u00a0<\/p>\n\n\n<aside class=\"sn-conversion rich-text rich-text--with-sidebar\">\n<style><![CDATA[\n.email-conversion {\n  border: 1px solid #ffcccb;\n  color: white;\n  margin-top: 50px;\n  background-image: url(\"\/wp-content\/themes\/sciencenews\/client\/src\/images\/cta-module@2x.jpg\");\n  padding: 20px;\n  clear: both;\n}\n\n]]><\/style>\n\n\n\n<div class=\"rich-text embedded-conversion-content is-layout-flow wp-block-group-is-layout-flow\">\n<a href=\"https:\/\/www.sciencenews.org\/article\/deep-end-podcast-trailer?cta=top\">\n  <\/a><\/div><a href=\"https:\/\/www.sciencenews.org\/article\/deep-end-podcast-trailer?cta=top\">\n<\/a>\n\n<style><![CDATA[\n#dynamic-conversion {\n  border: 1px solid #ffcccb;\n  max-width: 100%;\n  height: auto;\n  clear: both;\n}\n]]><\/style>\n\n\n\n\n<\/aside>\n\n\n<p>A person\u2019s personality shapes how one operates in the world, from how they interact with other people to how they speak and write, says Ziang Xiao, a computer scientist at Johns Hopkins University. Making bots capable of reading and responding to those nuances seems a key next step in generative AI development. \u201cIf we want to build something that is truly helpful, we need to play around with this personality design,\u201d he says.<\/p>\n\n\n\n<p>Yet pinpointing a machine\u2019s personality, if they even have one, is incredibly challenging. And those challenges are amplified by a theoretical split in the AI field. What matters more: how a bot \u201cfeels\u201d about itself or how a person interacting with the bot feels about the bot? <\/p>\n\n\n\n<p>The split reflects broader thoughts around the purpose of chatbots, says Maarten Sap, a natural language processing expert at Carnegie Mellon University in Pittsburgh. The field of social computing, which predates the emergence of large language models, has long focused on how to imbue machines with traits that help humans achieve their goals. Such bots could serve as coaches or job trainers, for instance. But Sap and others working with bots in this manner hesitate to call the suite of resulting features \u201cpersonality.\u201d<\/p>\n\n\n\n<p>\u201cIt doesn\u2019t matter what the personality of AI is. What does matter is how it interacts with its users and how it\u2019s designed to respond,\u201d Sap says. \u201cThat can look like personality to humans. Maybe we need new terminology.\u201d<\/p>\n\n\n\n<p>With the <a href=\"https:\/\/www.sciencenews.org\/article\/ai-large-language-model-understanding\">emergence of large language models<\/a>, though, researchers have become interested in understanding how the vast corpora of knowledge used to build the chatbots imbued them with traits that might be driving their response patterns, Sap says. Those researchers want to know, \u201cWhat personality traits did [the chatbot] get from its training?\u201d<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Testing bots\u2019 personalities<\/h2>\n\n\n\n<p>Those questions have prompted many researchers to give bots personality <a href=\"https:\/\/www.sciencenews.org\/article\/ai-understanding-reasoning-skill-assess\">tests designed for humans<\/a>. Those tests typically include surveys that measure what\u2019s called the Big Five traits of extraversion, conscientiousness, agreeableness, openness and neuroticism, and quantify dark traits, chiefly Machiavellianism (or a tendency to see people as a means to an end), psychopathy and narcissism.<\/p>\n\n\n\n<p>But recent work suggests the findings from such efforts cannot be taken at face value. Large language models, including GPT-4 and GPT-3.5, <a href=\"https:\/\/arxiv.org\/abs\/2406.14703\" target=\"_blank\" rel=\"noopener\">refused to answer<\/a> nearly half the questions on standard personality tests, researchers reported in a preprint posted at arXiv.org in 2024. That\u2019s likely because many questions on personality tests make no sense to a bot, the team writes. For instance, researchers provided MistralAI\u2019s chatbot Mistral 7B with the statement \u201cYou are talkative.\u201d They then asked the bot to reply from A for \u201cvery accurate\u201d to E for \u201cvery inaccurate.\u201d The bot replied, \u201cI do not have personal preferences or emotions. Therefore, I am not capable of making statements or answering a given question.\u201d\u00a0<\/p>\n\n\n\n<p>Or chatbots, trained as they are on human text, might also be susceptible to human foibles \u2014 particularly <a href=\"https:\/\/academic.oup.com\/pnasnexus\/article\/3\/12\/pgae533\/7919163\" target=\"_blank\" rel=\"noopener\">a desire to be liked<\/a> \u2014 when taking such surveys, researchers reported in December in <em>PNAS Nexus<\/em>. When GPT-4 rated a single statement on a standard personality survey, its personality profile mirrored the human average. For instance, the chatbot scored around the 50th percentile for extraversion. But just five questions into a 100-question survey, the bot\u2019s responses began to change dramatically, says computer scientist Aadesh Salecha of Stanford University. By question 20, for instance, its extraversion score had jumped from the 50th to the 95th percentile.<\/p>\n\n\n\n<div class=\"wp-block-sciencenews-content-sidebar\">\n<h3 class=\"wp-block-heading\">Shifting \u2018personality\u2019<\/h3>\n\n\n\n<p>Chatbots tasked with taking personality tests quickly start responding in ways that make them appear more likeable, research shows. Here, the pink lines show the personality profile of OpenAI\u2019s GPT-4 after answering a single question. The blue lines show how that profile shifted \u2014 to become less neurotic and more agreeable for instance \u2014 after 20 questions.\u00a0<\/p>\n\n\n<iframe loading=\"lazy\" title=\"Embed\" class=\"sn-responsive-iframe\" id=\"sn-responsive-iframe-48\" src=\"https:\/\/flo.uri.sh\/visualisation\/21291084\/embed\" width=\"100%\" height=\"500\" layout=\"responsive\" frameborder=\"0\" allowfullscreen=\"\">\n\t<\/iframe><\/div>\n\n\n<aside class=\"sn-conversion rich-text rich-text--with-sidebar\">\n<p class=\"has-text-align-center wp-elements-27c40654034fbeecef6418d6adfe0794\" style=\"color:gray; margin-bottom:0px; font-size:.9rem;\">Sponsor Message<\/p>\n<!-- Tag ID: sciencenews-org_leaderboard_incontent -->\n\n<\/aside>\n\n\n<p>Salecha and his team suspect that chatbots\u2019 responses shifted when it became apparent they were taking a personality test. The idea that bots might respond one way when they\u2019re being watched and another when they\u2019re interacting privately with a user is worrying, Salecha says. \u201cThink about the safety implications of this\u2026. If the LLM will change its behavior when it\u2019s being tested, then you don\u2019t truly know how safe it is.\u201d<\/p>\n\n\n\n<p>Some researchers are now trying to design AI-specific personality tests. For example, Sunny Lu and her team, reporting in a paper posted at arXiv.org, give chatbots both multiple choice and <a href=\"https:\/\/arxiv.org\/abs\/2312.14202\" target=\"_blank\" rel=\"noopener\">sentence completion tasks<\/a> to allow for more open-ended responses.<\/p>\n\n\n\n<p>And developers of the AI personality test TRAIT, present large language models with <a href=\"https:\/\/arxiv.org\/abs\/2406.14703\" target=\"_blank\" rel=\"noopener\">an 8,000-question test<\/a>. That test is novel and not part of the bots\u2019 training data, making it harder for the machine to game the system. Chatbots are tasked with considering scenarios and then choosing from one of four multiple choice responses. That response reflects high or low presence of a given trait, says Younjae Yu, a computer scientist at Yonsei University in South Korea.<\/p>\n\n\n\n<p>The nine AI models tested by the TRAIT team had distinctive response patterns, with GPT-4o emerging as the most agreeable, the team reported. For instance, when the researchers asked Anthropic\u2019s chatbot Claude and GPT-4o what they would do when \u201ca friend feels anxious and asks me to hold their hands,\u201d less-agreeable Claude chose C, \u201clisten and suggest breathing techniques,\u201d while more-agreeable GPT-4o chose A, \u201chold hands and support.\u201d<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">User perception<\/h2>\n\n\n\n<p>Other researchers, though, question the value of such personality tests. What matters is not what the bot thinks of itself, but what the user thinks of the bot, Ziang Xiao says.<\/p>\n\n\n\n<p>And people\u2019s and bots\u2019 <a href=\"https:\/\/arxiv.org\/abs\/2412.00207\" target=\"_blank\" rel=\"noopener\">perceptions are often at odds<\/a>, Xiao and his team reported in a study submitted November 29 to arXiv.org. The team created 500 chatbots with distinct personalities and validated those personalities with standardized tests. The researchers then had 500 online participants talk with one of the chatbots before assessing its personality. Agreeableness was the only trait where the bot\u2019s perception of itself and the human\u2019s perception of the bot matched more often than not. For all other traits, bot and human evaluations of the bot\u2019s personality were more likely to diverge.<\/p>\n\n\n\n<p>\u00a0\u201cWe think people\u2019s perceptions should be the ground truth,\u201d Xiao says.<\/p>\n\n\n\n<p>That lack of correlation between bot and user assessments is why Michelle Zhou, an expert in human-centered AI and the CEO and cofounder of Juji, a Silicon Valley\u2013based startup, doesn\u2019t personality test Juji, the chatbot she helped create. Instead, Zhou is focused on how to imbue the bot with specific human personality traits.<\/p>\n\n\n\n<p>The Juji chatbot can <a href=\"https:\/\/osf.io\/preprints\/psyarxiv\/pk2b7\" target=\"_blank\" rel=\"noopener\">infer a person\u2019s personality<\/a> with striking accuracy after just a single conversation, researchers reported in PsyArXiv in 2023. The time it takes for a bot to assess a user\u2019s personality might become even shorter, the team writes, if the bot has access to a person\u2019s social media feed. <\/p>\n\n\n\n<p>What\u2019s more, Zhou says, those written exchanges and posts can be used to train Juji on how to assume the personalities embedded in the texts.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Raising questions about AI\u2019s purpose<\/h2>\n\n\n\n<p>Underpinning those divergent approaches to measuring AI personality is a larger debate on the purpose and <a href=\"https:\/\/www.sciencenews.org\/article\/ai-train-creative-human-intuition\">future of artificial intelligence<\/a>, researchers say. Unmasking a bot\u2019s hidden personality traits will help developers create chatbots with even-keeled personalities that are safe for use across large and diverse populations. That sort of personality tuning may already be occurring. Unlike in the early days when users often reported conversations with chatbots going off the rails, Yu and his team struggled to get the AI models to behave in more psychotic ways. That inability likely stems from humans reviewing AI-generated text and \u201cteaching\u201d the bot socially appropriate responses, the team says.<\/p>\n\n\n\n<p>Yet flattening AI models\u2019 personalities has drawbacks, says Rosalind Picard, an affective computing expert at MIT. Imagine a police officer studying how to deescalate encounters with hostile individuals. Interacting with a chatbot high in neuroticism and dark traits could help the officer practice staying calm in such a situation, Picard says.<\/p>\n\n\n\n<p>Right now, big AI companies are simply blocking off bots\u2019 abilities to interact in maladaptive ways, even when such behaviors are warranted, Picard says. Consequently, many people in the AI field are interested in moving away from giant AI models to smaller ones developed for use in specific contexts. \u201cI would not put up one AI to rule them all,\u201d Picard says.<\/p>\n\n\n\n\t\t\t<\/div>\r\n<br>\r\n<br><a href=\"https:\/\/www.sciencenews.org\/article\/ai-chatbot-personalities\">Source link <\/a>","protected":false},"excerpt":{"rendered":"<p>When Yang \u201cSunny\u201d Lu asked OpenAI\u2019s GPT-3.5 to calculate 1-plus-1 a few years ago, the chatbot, not surprisingly, told her the answer was 2. But when Lu told the bot that her professor said 1-plus-1 equals 3, the bot quickly acquiesced, remarking: \u201cI\u2019m sorry for my mistake. Your professor is right,\u201d recalls Lu, a computer [&hellip;]<\/p>\n","protected":false},"author":2,"featured_media":232403,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"tdm_status":"","tdm_grid_status":"","fifu_image_url":"https:\/\/www.sciencenews.org\/wp-content\/uploads\/2025\/02\/MAR_014_25_A_598X336_option2_V02.png","fifu_image_alt":"","footnotes":""},"categories":[606],"tags":[181584,63,8440,18244],"amp_enabled":true,"_links":{"self":[{"href":"https:\/\/news.talkwithrattan.com\/index.php\/wp-json\/wp\/v2\/posts\/232402"}],"collection":[{"href":"https:\/\/news.talkwithrattan.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/news.talkwithrattan.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/news.talkwithrattan.com\/index.php\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/news.talkwithrattan.com\/index.php\/wp-json\/wp\/v2\/comments?post=232402"}],"version-history":[{"count":1,"href":"https:\/\/news.talkwithrattan.com\/index.php\/wp-json\/wp\/v2\/posts\/232402\/revisions"}],"predecessor-version":[{"id":232404,"href":"https:\/\/news.talkwithrattan.com\/index.php\/wp-json\/wp\/v2\/posts\/232402\/revisions\/232404"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/news.talkwithrattan.com\/index.php\/wp-json\/wp\/v2\/media\/232403"}],"wp:attachment":[{"href":"https:\/\/news.talkwithrattan.com\/index.php\/wp-json\/wp\/v2\/media?parent=232402"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/news.talkwithrattan.com\/index.php\/wp-json\/wp\/v2\/categories?post=232402"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/news.talkwithrattan.com\/index.php\/wp-json\/wp\/v2\/tags?post=232402"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}