{"id":15346,"date":"2026-08-22T12:43:14","date_gmt":"2026-08-22T12:43:14","guid":{"rendered":"https:\/\/koshalsambada.in\/?p=15346"},"modified":"2026-08-22T12:43:14","modified_gmt":"2026-08-22T12:43:14","slug":"meet-zeynep-demirbas-the-new-york-eighth-grader-who-tested-whether-ai-can-recognise-stress-a-basic-machine-learning-model-beat-chatgpt-4o","status":"publish","type":"post","link":"https:\/\/koshalsambada.in\/?p=15346","title":{"rendered":"Meet Zeynep Demirbas, the New York eighth-grader who tested whether AI can recognise stress; a basic machine-learning model beat ChatGPT-4o"},"content":{"rendered":"<p><br \/>\n<\/p>\n<div>\n<div class=\"e9jwa\">\n<div class=\"vdo_embedd\">\n<div class=\"GfdvZ\">\n<section class=\"_bIDB  clearfix id-r-component leadmedia undefined undefined  E9tg9 \" style=\"top:0px\">\n<div class=\"_bIDB\" data-ua-type=\"1\" onclick=\"stpPgtnAndPrvntDefault(event)\">\n<div class=\"ypVvZ\">\n<div class=\"WGttI\"><img src=\"https:\/\/static.toiimg.com\/thumb\/msid-133400178,imgsize-101722,width-400,height-225,resizemode-4\/zeyneps-work-earned-her-a-finalist-place-in-the-2025-thermo-fisher-scientific-junior-innovators-challenge.jpg\" alt=\"Meet Zeynep Demirbas, the New York eighth-grader who tested whether AI can recognise stress; a basic machine-learning model beat ChatGPT-4o\" title=\"Zeynep's work earned her a finalist place in the 2025 Thermo Fisher Scientific Junior Innovators Challenge\" decoding=\"async\" fetchpriority=\"high\"\/><\/div>\n<\/div>\n<\/div>\n<div class=\"Ta7d_ img_cptn\"><span title=\"Zeynep's work earned her a finalist place in the 2025 Thermo Fisher Scientific Junior Innovators Challenge\">Zeynep&#8217;s work earned her a finalist place in the 2025 Thermo Fisher Scientific Junior Innovators Challenge<\/span><\/div>\n<\/section>\n<\/div><\/div>\n<\/div>\n<p>Fourteen-year-old Zeynep Demirbas has found that ChatGPT-4o is less accurate than a mental health-focused AI model and a simpler machine-learning system at detecting stress in human written text.<span class=\"id-r-component br\" data-pos=\"2\"\/>Zeynep, an eighth-grade student at Transit Middle School in East Amherst, New York, tested four different models using more than 3,500 Reddit posts. According to Society for Science, the posts had already been labelled by humans as showing stress or no stress.<span class=\"id-r-component br\" data-pos=\"4\"\/>Her project, titled \u201c<a href=\"https:\/\/sspcdn.blob.core.windows.net\/files\/Documents\/SEP\/JIC\/2025\/poster\/2025JIC_Demirbas.Zeynep.Poster.pdf\" rel=\"noopener nofollow noreferrer\" styleobj=\"[object Object]\" class=\"\" target=\"_blank\" commonstate=\"[object Object]\" frmappuse=\"1\">Evaluating the reliability of Large Language Models for stress detection<\/a>\u201d, has earned her a place among the finalists in the 2025 Thermo Fisher Scientific Junior Innovators Challenge. The competition recognises young students working on science-based projects.<span class=\"id-r-component br\" data-pos=\"10\"\/>Zeynep became interested in the project after speaking with a family friend who is a psychologist. The psychologist told her that some health insurance companies were exploring large language models (LLMs) as cheaper, 24\/7 alternatives to human therapists. Zeynep wondered whether AI systems could actually be trusted to identify stress.<span class=\"id-r-component br\" data-pos=\"12\"\/><\/p>\n<p><h2>Testing AI models <\/h2>\n<\/p>\n<p><span class=\"id-r-component br\" data-pos=\"14\"\/>To test the models, Zeynep used a dataset called Dreaddit. It contains 3,553 Reddit posts that human raters had labelled according to whether they contained signs of stress.<span class=\"id-r-component br\" data-pos=\"17\"\/>She gave the data to four different models including Bidirectional Encoder Representations from Transformers (BERT), MentalBERT, Random Forest and ChatGPT-4o. MentalBERT is a version of BERT designed for mental health-related language, while Random Forest is a basic machine-learning technique that uses multiple decision trees to make predictions.<span class=\"id-r-component br\" data-pos=\"19\"\/>Zeynep asked each model to identify which posts showed stress. <!-- -->She then used a measure called an F1-score to compare their performance. The score considers both how accurately a model identifies stress and how often it misses stress or wrongly labels a post as showing stress.<span class=\"id-r-component br\" data-pos=\"23\"\/>MentalBERT performed the best in her testing, with a score of about 82 percent. BERT followed with about 79 percent. ChatGPT-4o scored about 74 percent. It also performed worse than the Random Forest model, which was included as a simpler baseline for comparison.<span class=\"id-r-component br\" data-pos=\"26\"\/>The result surprised Zeynep because Random Forest is a much simpler machine-learning method and does not understand language and context in the same way an LLM does. \u201cChatGPT performing badly was \u2018really surprising,\u2019\u201d Zeynep said.<span class=\"id-r-component br\" data-pos=\"28\"\/>She found it particularly interesting that a simpler model could outperform an LLM with millions of parameters. \u201cRandom-forest is \u2018supposed to be a very simple and old technique. So I just put it in as a baseline,\u2019\u201d Zeynep said, as quoted by Science News Explores. <!-- -->\u201cThat was very interesting; how something so small and simple was able to beat an LLM like ChatGPT that used millions of parameters and had so much coding go into it,\u201d she added.<span class=\"id-r-component br\" data-pos=\"32\"\/><span class=\"id-r-component br\" data-pos=\"34\"\/><\/p>\n<p><h2>What results mean<\/h2>\n<\/p>\n<p><span class=\"id-r-component br\" data-pos=\"36\"\/>Zeynep&#8217;s findings made her question whether general-purpose LLMs are reliable enough to be used for mental health assessment. A large language model is a type of machine-learning system trained on very large amounts of text. It learns patterns in language and uses them to produce responses.<span class=\"id-r-component br\" data-pos=\"39\"\/>Her results led her to conclude that LLMs should not replace human therapists.\u201cMy project shows that LLMs are currently unreliable and unsafe to deploy as diagnostic tools,\u201d Zeynep said.<span class=\"id-r-component br\" data-pos=\"41\"\/>She added that the findings did not mean that LLMs are bad or cannot be useful. But, they show that general-purpose AI systems may not be suitable for a task as sensitive as assessing mental health.<span class=\"id-r-component br\" data-pos=\"43\"\/>\u201cWe should be mindful with AI, because it doesn\u2019t really have an acceptable grade in mental health,\u201d Zeynep said. <!-- -->\u201cThat doesn\u2019t mean that LLMs are bad, because they\u2019re for general use. They\u2019re not necessarily meant for mental health,\u201d she added.<span class=\"id-r-component br\" data-pos=\"47\"\/>Zeynep also suggested that LLMs could potentially have a different role. Instead of replacing mental health professionals, they might help identify people who are struggling and refer them to a mental health professional.<span class=\"id-r-component br\" data-pos=\"49\"\/><\/p>\n<p><h2>Zeynep wants to study AI bias<\/h2>\n<\/p>\n<p><span class=\"id-r-component br\" data-pos=\"51\"\/>The project also made Zeynep interested in whether LLMs might show different results depending on a person&#8217;s gender. <!-- -->\u201cOne way I feel I could expand it is seeing whether LLMs carry biases toward different genders,\u201d Zeynep said.<span class=\"id-r-component br\" data-pos=\"55\"\/>She said she had read about cases where doctors dismiss symptoms reported by female patients because they believe women are exaggerating. She sees this as an example of personal bias and wants to know whether AI systems could show similar patterns.<span class=\"id-r-component br\" data-pos=\"58\"\/>Since LLMs are trained using large amounts of text created by people, Zeynep believes they can also pick up human biases.<span class=\"id-r-component br\" data-pos=\"61\"\/>Zeynep&#8217;s work earned her a finalist place in the 2025 Thermo Fisher Scientific Junior Innovators Challenge. She hopes to become a computer scientist.<span class=\"id-r-component br\" data-pos=\"63\"\/>She said she enjoys programming but is particularly interested in how computer science can be applied to real-world problems.<\/div>\n<p><br \/>\n<br \/><a href=\"https:\/\/timesofindia.indiatimes.com\/world\/us\/meet-zeynep-demirbas-the-new-york-eighth-grader-who-tested-whether-ai-can-recognise-stress-a-basic-machine-learning-model-beat-chatgpt-4o\/articleshow\/133399352.cms\" target=\"_blank\" rel=\"noopener\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Zeynep&#8217;s work earned her a finalist place in the 2025 Thermo Fisher Scientific Junior Innovators Challenge Fourteen-year-old Zeynep Demirbas has found that ChatGPT-4o is less accurate than a mental health-focused AI model and a simpler machine-learning system at detecting stress in human written text.Zeynep, an eighth-grade student at Transit Middle School in East Amherst, New [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":15347,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[31],"tags":[],"class_list":["post-15346","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-31"],"magazineBlocksPostFeaturedMedia":{"thumbnail":"https:\/\/koshalsambada.in\/wp-content\/uploads\/2026\/08\/zeyneps-work-earned-her-a-finalist-place-in-the-2025-thermo-fisher-scientific-junior-innovators-chal-150x150.avif","medium":"https:\/\/koshalsambada.in\/wp-content\/uploads\/2026\/08\/zeyneps-work-earned-her-a-finalist-place-in-the-2025-thermo-fisher-scientific-junior-innovators-chal-300x169.avif","medium_large":"https:\/\/koshalsambada.in\/wp-content\/uploads\/2026\/08\/zeyneps-work-earned-her-a-finalist-place-in-the-2025-thermo-fisher-scientific-junior-innovators-chal.avif","large":"https:\/\/koshalsambada.in\/wp-content\/uploads\/2026\/08\/zeyneps-work-earned-her-a-finalist-place-in-the-2025-thermo-fisher-scientific-junior-innovators-chal.avif","1536x1536":"https:\/\/koshalsambada.in\/wp-content\/uploads\/2026\/08\/zeyneps-work-earned-her-a-finalist-place-in-the-2025-thermo-fisher-scientific-junior-innovators-chal.avif","2048x2048":"https:\/\/koshalsambada.in\/wp-content\/uploads\/2026\/08\/zeyneps-work-earned-her-a-finalist-place-in-the-2025-thermo-fisher-scientific-junior-innovators-chal.avif","blogsy-small":"https:\/\/koshalsambada.in\/wp-content\/uploads\/2026\/08\/zeyneps-work-earned-her-a-finalist-place-in-the-2025-thermo-fisher-scientific-junior-innovators-chal.avif","blogsy-small-tall":"https:\/\/koshalsambada.in\/wp-content\/uploads\/2026\/08\/zeyneps-work-earned-her-a-finalist-place-in-the-2025-thermo-fisher-scientific-junior-innovators-chal.avif","blogsy-small-square":"https:\/\/koshalsambada.in\/wp-content\/uploads\/2026\/08\/zeyneps-work-earned-her-a-finalist-place-in-the-2025-thermo-fisher-scientific-junior-innovators-chal.avif","blogsy-small-masonry":"https:\/\/koshalsambada.in\/wp-content\/uploads\/2026\/08\/zeyneps-work-earned-her-a-finalist-place-in-the-2025-thermo-fisher-scientific-junior-innovators-chal.avif","blogsy-medium":"https:\/\/koshalsambada.in\/wp-content\/uploads\/2026\/08\/zeyneps-work-earned-her-a-finalist-place-in-the-2025-thermo-fisher-scientific-junior-innovators-chal.avif","blogsy-medium-masonry":"https:\/\/koshalsambada.in\/wp-content\/uploads\/2026\/08\/zeyneps-work-earned-her-a-finalist-place-in-the-2025-thermo-fisher-scientific-junior-innovators-chal.avif","blogsy-large":"https:\/\/koshalsambada.in\/wp-content\/uploads\/2026\/08\/zeyneps-work-earned-her-a-finalist-place-in-the-2025-thermo-fisher-scientific-junior-innovators-chal.avif","blogsy-wide":"https:\/\/koshalsambada.in\/wp-content\/uploads\/2026\/08\/zeyneps-work-earned-her-a-finalist-place-in-the-2025-thermo-fisher-scientific-junior-innovators-chal.avif"},"magazineBlocksPostAuthor":{"name":"admin","avatar":"https:\/\/secure.gravatar.com\/avatar\/8709732a479614e7a8aa24d3eb1b239f30dc6d90c61464ed495001e7a469d856?s=96&d=mm&r=g"},"magazineBlocksPostCommentsNumber":"0","magazineBlocksPostExcerpt":"Zeynep&#8217;s work earned her a finalist place in the 2025 Thermo Fisher Scientific Junior Innovators Challenge Fourteen-year-old Zeynep Demirbas has found that ChatGPT-4o is less accurate than a mental health-focused AI model and a simpler machine-learning system at detecting stress in human written text.Zeynep, an eighth-grade student at Transit Middle School in East Amherst, New [&hellip;]","magazineBlocksPostCategories":["\u0b26\u0b47\u0b36 \u0b2c\u0b3f\u0b26\u0b47\u0b36"],"magazineBlocksPostViewCount":2,"magazineBlocksPostReadTime":4,"magazine_blocks_featured_image_url":{"full":["https:\/\/koshalsambada.in\/wp-content\/uploads\/2026\/08\/zeyneps-work-earned-her-a-finalist-place-in-the-2025-thermo-fisher-scientific-junior-innovators-chal.avif",400,225,false],"medium":["https:\/\/koshalsambada.in\/wp-content\/uploads\/2026\/08\/zeyneps-work-earned-her-a-finalist-place-in-the-2025-thermo-fisher-scientific-junior-innovators-chal-300x169.avif",300,169,true],"thumbnail":["https:\/\/koshalsambada.in\/wp-content\/uploads\/2026\/08\/zeyneps-work-earned-her-a-finalist-place-in-the-2025-thermo-fisher-scientific-junior-innovators-chal-150x150.avif",150,150,true]},"magazine_blocks_author":{"display_name":"admin","author_link":"https:\/\/koshalsambada.in\/author\/admin"},"magazine_blocks_comment":0,"magazine_blocks_author_image":"https:\/\/secure.gravatar.com\/avatar\/8709732a479614e7a8aa24d3eb1b239f30dc6d90c61464ed495001e7a469d856?s=96&d=mm&r=g","magazine_blocks_category":"<a href=\"#\" class=\"category-link category-link-31\">\u0b26\u0b47\u0b36 \u0b2c\u0b3f\u0b26\u0b47\u0b36<\/a>","_links":{"self":[{"href":"https:\/\/koshalsambada.in\/index.php?rest_route=\/wp\/v2\/posts\/15346","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/koshalsambada.in\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/koshalsambada.in\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/koshalsambada.in\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/koshalsambada.in\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=15346"}],"version-history":[{"count":0,"href":"https:\/\/koshalsambada.in\/index.php?rest_route=\/wp\/v2\/posts\/15346\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/koshalsambada.in\/index.php?rest_route=\/wp\/v2\/media\/15347"}],"wp:attachment":[{"href":"https:\/\/koshalsambada.in\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=15346"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/koshalsambada.in\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=15346"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/koshalsambada.in\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=15346"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}