Can AI reliably assess freezing of gait across Parkinson’s disease cohorts?
- 4 hours ago
- 5 min read
By: Po-Kai Yang, Juha Carlon, Maaike Goris, Emilie Klaver, Jorik Nonnekes, Richard J. A. van Wezel, Lisa Alcock, Alison J. Yarnall, Lynn Rochester, Clint Hansen, Christian Schlenstedt, Walter Maetzler, David Buzaglo, Marina Brozgol, Jeffrey M. Hausdorff, Alice Nieuwboer, Moran Gilat, Pieter Ginis, Bart Vanrumste, Benjamin Filtjens

A deep learning model using wearable movement sensors performed well in the cohort in which it was developed, but did not transfer reliably across six external cohorts of people with Parkinson’s disease. Rather than relying on fully automated analysis, the model can be used through AID-FOG, a proof-of-concept web platform in which experts review and correct AI-generated freezing-of-gait annotations and retain the final say. Their corrections can also be used to adapt the model to the new cohort, with performance in this study reaching its main plateau after approximately 50 minutes of expert-annotated data. VSC computing resources enabled the large multicentre benchmark and extensive fine-tuning experiments.

Why automated freezing-of-gait assessment is difficult
Freezing of gait (FOG) is a debilitating motor symptom of Parkinson's disease that can reduce mobility, independence, and quality of life and increase fall risk. Manual video review by movement-disorder experts remains the reference standard for research assessment, but reviewing long recordings is difficult to scale, and subtle or borderline episodes can lead to differences between raters. Deep learning models based on small wearable movement sensors, known as inertial measurement units (IMUs), offer a faster and more scalable alternative. The key question is whether a model trained in one cohort still performs reliably in another. Without robust generalization, automated FOG assessment cannot be used confidently across clinical studies or care settings.
Testing one model across seven cohorts
In this study, we developed a deep-learning model using a KU Leuven cohort of 85 participants and 2,043 trials. We then validated it in six independent external cohorts comprising 256 participants and 1,058 trials. In total, the study included 341 people with Parkinson's disease across seven cohorts. The cohorts differed in FOG-provoking tasks and protocols, annotation criteria, IMU configurations, and ranges of disease severity.
Agreement was quantified using the intraclass correlation coefficient (ICC), for which higher values indicate closer agreement between the model and expert assessments. We evaluated two clinically relevant outcomes: percentage of time frozen (%TF) and number of freezing episodes. For %TF, agreement was strong in the local cohort (ICC = 0.886) but fell to fair across the external cohorts (ICC = 0.562 ± 0.141). Agreement for the number of freezing episodes was lower overall. The results showed that performance in the development cohort did not directly transfer to other datasets.
“Performance in the development cohort did not directly transfer to other datasets.”
Adapting the model with limited local data
To test how much cohort-specific expert input was needed to refine the model, we fine-tuned the pre-trained model using between 10 and 360 minutes of expert-annotated data from each external cohort. Performance reached its main plateau after about 50 minutes. At that point, average agreement improved to ICC = 0.732 ± 0.138 for percentage of time frozen and ICC = 0.623 ± 0.214 for the number of freezing episodes. Models trained from scratch required more annotated data and reached lower plateau values. Refining the pre-trained model was therefore more data-efficient than building a new model for each cohort.
“Performance reached its main plateau after about 50 minutes of expert-annotated data.”


Why humans remain in the loop
These results support an AI-assisted workflow in which the model provides an initial set of freezing-of-gait annotations and a clinician or researcher reviews and corrects them. The expert retains the final say over the annotations used for analysis. These verified annotations can also provide the cohort-specific data needed to refine the model and improve its reliability for subsequent analyses.
“The expert retains the final say over the annotations used for analysis.”
Even after fine-tuning, agreement remained at the lower end of the range reported between human raters. Expert oversight therefore remains important when applying the model across heterogeneous cohorts. This approach combines the efficiency of automated initial annotations with expert quality control.
The findings also reinforce the broader harmonization effort led by the International Consortium for FOG (ICFOG), which is working to standardize FOG definitions, assessment protocols, and measurement technology. Greater harmonization should improve future model transfer, while human-in-the-loop workflows provide a cautious pathway for using these models across current datasets.
Putting the workflow into practice with AID-FOG
To put this workflow into practice, we developed AID-FOG, a proof-of-concept web platform in which experts can inspect and correct AI-generated freezing-of-gait annotations. The expert-verified annotations remain the final output, while the corrections can also be used to fine-tune the model for that cohort. The platform therefore combines immediate expert quality control with a mechanism for preparing cohort-specific models for clinical studies.
Key findings
The study evaluated one deep-learning model across seven cohorts, including 341 people with Parkinson’s disease and 3,101 trials.
Agreement for percentage of time frozen fell from ICC = 0.886 in the development cohort to ICC = 0.562 ± 0.141 across six external cohorts, demonstrating limited transfer to new datasets.
Fine-tuning with approximately 50 minutes of cohort-specific expert-annotated data improved agreement to ICC = 0.732 ± 0.138 for percentage of time frozen and ICC = 0.623 ± 0.214 for the number of freezing episodes.
The results support an AI-assisted workflow in which the model proposes initial annotations, experts verify and correct them while retaining the final say, and the resulting verified data can be used to refine the model for the cohort.
AID-FOG demonstrates this workflow in a proof-of-concept web platform for expert review and cohort-specific fine-tuning.
VSC contribution
“VSC's high-performance computing resources made this multicentre evaluation computationally feasible. We trained and evaluated hundreds of deep-learning models across seven cohorts and repeated the analyses for many fine-tuning dataset sizes. Running these experiments in parallel on A100 and H100 GPUs in the VSC Tier-2 infrastructure allowed us to complete the external-validation benchmark and fine-tuning sweep within a feasible project timeline.”
Explore the research
🔍 Your Research Matters — Let’s Share It!
Have you used VSC’s computing power in your research? Did our infrastructure support your simulations, data analysis, or workflow?
We’d love to hear about it!
Take part in our #ShareYourSuccess campaign and show how VSC helped move your research forward. Whether it’s a publication, a project highlight, or a visual from your work, your story can inspire others.
🖥️ Be featured on our website and social media. Show the impact of your work. Help grow our research community
📬 Submit your story: https://www.vscentrum.be/sys




