Putting Digital Technologies to Work in Improving Mental Health

Posted: August 31, 2026
Putting Digital Technologies to Work in Improving Mental Health

PAIRING SMARTPHONE SENSORS & ChatGPT TO TRACK ANHEDONIA IN DEPRESSED ADOLESCENTS

The ubiquity of smartphones and the rapid advance and widespread public adoption of AI tools like ChatGPT has raised the question of how such technologies can be harnessed to improve psychiatric assessment and treatment.

A recent pilot study directed by Christian A. Webb, Ph.D. a 2018 and 2015 BBRF Young Investigator at McLean Hospital and Harvard Medical School, tested one way in which mobile phones and the large language model (LLM)-based ChatGPT tool can together aid in the assessment of behavioral changes that are important for therapeutic improvement in a behavioral therapy for adolescents with depression.

Innovations of this kind could open doors to therapy adjustments made in real time based on indicators gleaned directly from patients. Also involved in the research were Diego A. Pizzagalli, Ph.D., BBRF Scientific Council, 2017 BBRF Distinguished Investigator, 2008 Independent Investigator; and Erika E. Forbes, Ph.D., 2014 BBRF Independent Investigator, 2006 Young Investigator. Smartphones can passively collect rich behavioral data—in research like this, with users’ consent—by tracking patterns of movement and activity. LLMs like ChatGPT can analyze text and language with impressive accuracy. Still, it has remained unclear whether these technologies can provide insight into the emotional or motivational aspects of an individual user’s behavior

This question helped to drive the team’s study, which examined the two technologies in the context of a talk therapy called behavioral activation (BA), which aims to reduce the classic depression symptom of anhedonia (the loss of interest in pursuing rewards or pleasure) by targeting patterns of avoidance and withdrawal and increasing patients’ engagement with rewarding activities.

A core assumption of BA is that increased activity in daily life leads to symptom change. “Tools that assess activation between [therapy] sessions could help therapists monitor treatment progress and make timely adjustments, an objective consistent with calls for data-informed psychotherapy,” noted Dr. Webb and his team.

Their pilot study made use of smartphone-based passive sensing of patients’ mobility patterns and LLM-derived ratings of daily text entries they provided during the 3-month therapy period. These data were examined in relation to daily assessments of positive and negative emotions, as well as weekly assessments of anhedonia, depressive symptoms, and a traditional self-report of activation.

The team recruited 38 adolescents, ages 13–18, to receive BA treatment targeting their anhedonia. Each was offered 12 weekly hour-long individual BA therapy sessions. Before each session, participants rated their anhedonia, depression, and “activation” (in essence, their engagement in daily activities and avoidance patterns), using self-report measures. A subset of 13 participants also contributed passive smartphone data—collected continuously via built-in cellphone acclerometers and GPS sensors.

Every other week during treatment all participants completed a “burst” of real-time assessments (“ecological momentary assessments”) over 5 consecutive days via 2–3 surveys delivered each day via a smartphone app. Each time they were prompted by the app, participants were asked to rate their positive and negative affect (6 items, to be rated on a 5-point scale, generating data about being “happy,” “interested,” and “excited”; and “sad,” “nervous,” and “angry”). Next, they were asked to provide responses via unrestricted texting, “about what they were doing the night before, with whom they were interacting, and the most enjoyable and stressful events since the previous prompt.”

The text responses were analyzed using OpenAI’s GPT-4o model, as well as by a human “rater” for one-fourth of responses, to evaluate consistency between human and AI ratings. The team found “substantial agreement” between the two. As for the “silent” real-time monitoring made possible by the smartphone, this data was used to measure activity levels using the intensity of physical activity during the day; the percentage of time a participant spent at home; the distance and “mobility area” traveled each day relative to the home; and places visited each day.

The study produced interesting results. The most important, perhaps, was that “activation” among the depressed and anhedonic adolescents—as deduced from passive sensors in their smartphones—correlated positively with “activation ratings” generated by GPT analysis of their texts and patients’ activation self-report score over the course of therapy.

On the GPS side, greater activation corresponded to more places visited and more time spent away from home, though not distance traveled. In addition, increases in GPT-rated activation were associated with higher daily positive and lower daily negative affect. The team also found that smartphone passive sensing features could be used to forecast weekly improvements in anhedonia and depressive symptoms.

The results led the researchers to suggest that LLMs such as GPT can extract psychologically meaningful information from unstructured text. “Importantly, our study shows that LLM-based assessments can also provide clinically relevant insights based on language generated outside the therapy room, offering scalable and unobtrusive ways to monitor therapeutic processes in patients’ daily lives.”

“Daily text assessments could help clinicians monitor” the response to behavioral activation therapy “in real time, while mobility patterns could indicate whether treatment is gradually translating into symptom improvement.”

In another potential application of mobile technology to treatment of psychiatric disorders, Dr. Webb and colleagues, in a separate study, tested whether a freely available smartphone app featuring “mindfulness” exercises could be useful in helping adolescents ruminate less. Rumination refers to repetitive and negative self-focused thinking, often concerning stressful or negative past events. It is often seen in adolescents who are anxious or depressed, and studies have shown that it is a style of thinking that can predict the onset of both disorders.

The app used in the study, called CARE, was downloaded on each of the participants’ phones and they were taught how to use it. Based on their inputs of sleep and wake times, users were prompted via random notifications within that time window to engage the app. Each time they used the app, they took a survey to assess whether they were ruminating and to what degree, and to indicate their current mood.

90% (72 of 80) of the adolescents completed a 3-week trial with the CARE app, with the typical user completing a total of 29 mindfulness training sessions, an average of 1.5 sessions per day. Reductions in rumination were assessed over two time intervals—”immediate” (i.e., pre- to post-mindfulness exercise) and “cumulative” (i.e., overall change in rumination over the course of the 3-week trial). Use of the app led to better immediate success among girls and older adolescents.

Those with higher levels of rumination at the beginning of the study, and those who suppressed their emotions less, had better cumulative outcomes. Levels of anxiety and depression symptoms prior to the trial did not predict who would most likely be helped.

Figure 1

AN APP TO DETECT LIKELIHOOD OF AUTISM AT ‘WELL-CHILD’ PEDIATRIC VISIT

Researchers have developed an app delivered on digital tablet devices that significantly improves upon currently existing tools to screen toddlers for a possible diagnosis of autism spectrum disorder (ASD). Early detection of ASD is widely regarded as critical, making it more likely a child will have prompt access to therapeutic interventions that can improve outcomes.

The app, called SenseToKnow, was deployed in a multi-clinic setting. In an initial test reported in Nature Medicine, it was administered in 475 cases during a routine “well-child visit,” typically made when a child is 18–24 months of age.

Often, pediatricians give parents a questionnaire at the well-child visit called the Modified Checklist for Autism in Toddlers (M-CHAT). This tool has been shown to have higher accuracy in research settings than in the real world, especially in cases involving girls and children of color. Another important screening tool uses eye-tracking technology to measure children’s attentional preferences for social vs. non-social stimuli. Autism is characterized by reduced spontaneous visual attention to social stimuli, and while machine-learning analysis of eye-tracking data has yielded encouraging results for its ability to distinguish autistic vs. neurotypical children, it does so on the basis of only one of autism’s various behavioral signs, which also include facial expressions, head movements, response to name, and motor behaviors (including restricted or repetitive behaviors).

The SenseToKnow app involves the display on the tablet of brief movies; the child’s behavioral responses are recorded via the tablet’s frontal camera and are designed to elicit a wide range of autism-related behaviors including those just mentioned as well as social attention and eye-blink rate. It was developed by a team led by psychologist Geraldine Dawson, Ph.D., and electrical and computer engineer Guillermo Sapiro, Ph.D., of Duke University. The team that tested the app and analyzed the results included 2015 BBRF Young Investigator Kimberly L. H. Carpenter, Ph.D., and 2001 BBRF Young Investigator Scott N. Compton, Ph.D., both also at Duke.

The data collected by the app are quantified via computer vision analysis (CVA). Machine learning is used to integrate multiple digital data streams into a combined algorithm that classifies each child as either “autistic” or “nonautistic.” It also automatically generates metrics reflecting a statistical “confidence level” associated with the diagnostic assessment made in each case. If the app administration falls below a certain level, it can easily be re-administered to boost the confidence metric.

The app does not determine if a child has ASD. Screening tools like the app are designed to predict the likelihood of the presence or absence of a condition, but are not by themselves able to make a positive diagnosis. Expert clinical examination and various diagnostic methodologies are needed to make a “gold-standard” determination in each case.

In the group of 475 toddlers tested (they ranged in age from 17 to 36 months and included 269 boys and 206 girls), 49 were diagnosed with ASD and 98 were diagnosed with developmental delay without autism; 328 of the children were assessed as “neurotypical.” The key question was: how well did the app identify children in each category?

Overall, its sensitivity was 88%; its specificity was 81%. By comparison, in a prior test of over 25,000 children, the M-CHAT questionnaire had an excellent sensitivity of 95%, but a specificity of only 39%. This means the parental questionnaire didn’t miss very many cases of ASD, but it also generated a high number of false positives. The low false positive rate in the app suggests an important strength.

While the new app has the advantage of making its determination on the basis of entirely objective, computeranalyzed data, and captures features reflecting the variety of autism-related behaviors and early-warning signals, it remains imperfect. As the researchers note, some autistic toddlers will exhibit only a subset of ASD-related features. Conversely, some nonautistic children may exhibit behaviors typically associated with ASD.

Importantly, when results of SenseToKnow assessments were combined with those from the M-CHAT questionnaire, the “positive predictive value” of the result for each child increased substantially. Thus, one of the team’s conclusions was that combining available screening methods, specifically the new app and M-CHAT, which are both highly accessible at no or negligible cost, would significantly enhance the accuracy of autism screening in the real-world setting of the pediatrician’s office at the time of the “well-child” exam.

Efforts to improve SenseToKnow are under way, as are larger clinical tests that assess broader samples of the population with more typical rates of ASD and developmental delay.

Figure 2

WEARABLE SENSORS HELPED TO DISTINGUISH ADOLESCENT BIPOLAR DISORDER FROM ADHD AND OTHER DISORDERS

Diagnosing bipolar disorder (BD) can be difficult. To make a positive diagnosis, an individual must experience at least one major depressive episode and at least one manic or hypomanic episode. (Mania is characterized by significantly elevated and/or irritable mood as well as notably increased energy and activity; hypomania is a less intense version, but with equally important mental health implications). When an initial depressive episode precedes an initial manic/hypomanic one, it is impossible with current diagnostic tools to predict at that point that a given individual will or will not at some future time experience a manic or hypomanic episode. A diagnosis of major depressive disorder can be made, but not a BD diagnosis.

In some depressed people, a subsequent manic episode will occur, making a BD diagnosis possible. But even then, it is possible to confuse symptoms of other psychiatric conditions for changes in energy and activity that also occur in mania/hypomania. In ADHD, for example, hyperactivity can in some cases be mistaken for mania/hypomania. In view of this, a person with depressive symptoms and hyperactivity could conceivably have unipolar depression (e.g., major depressive disorder) and co-occurring ADHD—or, perhaps, bipolar disorder without ADHD.

Distinguishing people with overlapping symptoms involving both depressive mood and significantly elevated levels of energy and activity can result in delays of positive diagnosis (whether for BD or other disorders) that can be measured, in some cases, in years. Misdiagnosis or a long time-lag before a correct diagnosis is made can translate into missed opportunities for matching the patient with the therapies most likely to help them, which in turn can lead to less satisfactory long-term outcomes.

Researchers at the University of Pittsburgh are among those who have been working to discover objective markers—biological and/or behavioral—that might enable doctors to predict at an early stage that an individual has BD or one or more other disorders, or BD and one or more other disorders.

Michele A. Bertocci, Ph.D., a 2019 BBRF Young Investigator, and Rasim S. Diler M.D., led a team that also included Boris Birmaher, M.D., 2022 BBRF Ruane Prize winner and 2013 BBRF Colvin Prize winner, in utilizing actigraphy and artificial intelligence (AI) to test models that might be used in the clinic to help differentiate BD in adolescents from other conditions.

Actigraphy is an objective method of measuring an individual’s rest and activity cycles, via a device worn on the wrist that contains a digital accelerometer akin to those incorporated into smartphones and sleep- and exercise-monitoring devices.Actigraphy has helped other researchers better understand disturbances in daily circadian rhythms in BD patients. The new study uses the technology to examine the relationships between daytime activity and BD, excluding nightly periods of sleep. This focus on daily activity follows in part from past research showing that in adults, a decrease in daily activity has been seen in patients experiencing a depressive episode, relative to activity in manic/hypomanic episodes.

The team studied data from 389 adolescents, average age 15, who had been admitted to the specialized Child and Adolescent Bipolar Spectrum Services unit in Pittsburgh’s Western Psychiatric Hospital. The cohort studied included patients with BD but not ADHD (61); BD with ADHD (86); ADHD without BD (115); and a variety of other psychiatric disorders (127). All of these youths were inpatients at the time of data collection, and each had at least 4 days of actigraphy data.

Two AI programs were used to search for patterns in data that for each participant was broken down into non-overlapping hour-long periods during the day (between 7 am and 8 pm) characterized by comparative levels of activity. This yielded analytical units that registered intervals of minimum and maximum activity throughout the day. These 60-minute intervals were regarded by the team as “objective markers of clinical self-reports of elevated and reduced daily activity levels.” They aligned with DSM diagnostic thresholds such as the 4-hour daily criterion for elevated activity associated with hypo/mania.

A large set of what computer modelers call “features” was built into the algorithms used by the two AI programs to train themselves to parse the data and discover useful patterns. Both AI programs were highly effective in this task. With an accuracy of 90% or more, both programs used data on individual maximum and minimum daily activity along with patient age to classify the four diagnostic subsets within the total cohort.

“Actigraphy data, collected during inpatient stays,” succeeded in “differentiating between difficult-to-distinguish psychopathologies of BD without ADHD, BD with ADHD, ADHD without BD, and other illnesses,” the team reported in Psychiatry Research Communications.

“Each diagnostic group [in the cohort] shows behavioral trends toward elevations and reductions in daily activity differently,” they noted.

The team suggested that its findings, if replicated and extended, “strongly support” using actigraphy to provide objective and quantifiable measures to assess inpatient activity and classify diagnoses. This, they said, “can complement clinical observations, especially by identifying subtle differences in activity patterns that may not be visually apparent.”

Figure 3

FITBIT DATA STREAMS HELP PREDICT MOOD SHIFTS IN PATIENTS WITH BIPOLAR DISORDER

In addition to efforts to use digital technologies to help make accurate diagnoses of bipolar disorder, there are also applications that seek to address the mood shifts that characterize BD—an objective akin to efforts decribed earlier to track anhedonia and rumination symptoms in depressed and anxious young people.

A team led by 2019 BBRF Young Investigator Jessica M. Lipschitz, Ph.D., of Brigham and Women’s Hospital, recently tested whether data from wearable Fitbit devices (worn like a watch) could generate predictions about shifts in mood that are accurate enough to “meaningfully inform treatment” in BD patients. Katherine E. Burdick, Ph.D., winner of the BBRF Colvin Prize in 2021 and a two-time BBRF grantee, was senior member of the team.

In recent years, 2022 BBRF Young Investigator Sarah Sperry, Ph.D., and colleagues, have published evidence calling into question the commonplace assumption in clinical medicine that periods between low and high mood in bipolar disorder are ones of “normal” mood. Their finding of considerable “mood instability” in many patients between episodes of depression and mania/hypomania has the potential to improve care and quality of life by directing attention to these periods “between” major mood episodes.

The new research by Drs. Lipschitz, Burden, and colleagues suggests one way to track patients in real time. Called “digital phenotyping,” the moment-bymoment measurement of data such as sleep patterns, daily activity, and heart rate, collected by wearable digital devices, offers a potential avenue for early detection of major mood episodes in BD patients.

The team selected the Fitbit device because of its commercial availability and low cost, and most importantly because it is entirely passive: users do not input data of any kind. The devices have also proven over the years to be quite accurate and consistent in registering user data.

Participants in the Fitbit study had been diagnosed with Bipolar I (depression and at least one manic episode) or Bipolar II (depression and at least one hypomanic episode). Those included in the statistical analysis had to have completed at least 24 weeks of a 9-month Fitbit data monitoring period.

Although the Fitbit data is acquired passively, those in the final cohort of 54 also completed a self-report questionnaire given every 2 weeks during the 9 months, indicating any depression and/or manic/hypomanic symptoms. Data from these questionnaires were used to determine the accuracy of machine learning algorithms used to predict mood symptoms, not as input for the algorithms. The team identified one algorithm among the several tested that generated superior predictive results.

Data collected on each user by Fitbit devices covered such facets as number of daily steps taken; number of minutes spent in non-active or “sedentary” mode; heart rate (daily average and resting average); total sleep time; amounts of time and number of nightly intervals spent in deep sleep and REM sleep; number and duration of nightly awakenings; and bedtime.

After filtering out data from 11 of the original 65 participants (17%) who either dropped out or were not sufficiently compliant with the protocol, the team was able to use the Fitbit data to accurately predict clinically significant depression with 80.1% accuracy, and clinically significant manic/hypomanic symptoms with 89.1% accuracy. One other past test of Fitbit for this purpose yielded similar accuracy, but only after filtering out 46% of the data, meaning that it served its predictive purpose mainly in participants who were highly compliant with the protocol, and may not work as well in a general sample of BD patients seeking treatment. The accuracy numbers for this trial were calculated after the fact, looking back on all the data.

“Our findings are particularly noteworthy,” the team wrote, “because all input was passively collected; none of the metrics utilized were invasive in terms of privacy; we used mainstream consumer devices; and our methods did not demand high levels of Fitbit compliance.” For these reasons, the team suggests its results and methods “are an important next step toward a digital phenotyping approach that could be feasibly and broadly implemented across BD patients in routine care.”

Written By Peter Tarr, Ph.D.

Click here to read the Brain & Behavior Magazine's September 2026 issue