Comparison of Artificial Intelligence Tools With Human Coding for Sentiment, Topic, and Thematic Analysis Tasks of Public Health Datasets During the COVID-19 Pandemic in Australia: Case Study
Online Journal of Public Health Informatics (OJPHI)Research Authors: Danielle Hutchinson, Lauren Lee, Haley Stone, Aye Moa, Holly Seale, C. Raina MacIntyreAIIM Authors: Anisha Ojha, Amanda ZhongApproved by President Reda RiffiPublication Date: 4/7/2026Comprehensive Summary
This study evaluated five AI-based tools -- VADER, SentimentGI, SentimentQDAP, Microsoft Azure, and ChatGPT-4 -- against human coders for sentiment analysis, topic modeling, and thematic analysis of COVID-19-related public health datasets from Google Alerts (news media) and X (formerly Twitter) in Australia (December 2022-February 2023). Human coding revealed mostly neutral sentiment in news media and more polarized views on social media, with vaccines, lockdowns, and mandates attracting more negativity. All AI tools performed poorly on sentiment classification (Cohen κ <0.5 for all tools), with neutral sentiment being especially difficult to detect. However, ChatGPT-4 showed stronger alignment with human raters for topic modeling and thematic analysis, with 13 of 20 LLM-generated themes rated proficient and all rated "very reasonable," though outputs lacked contextual depth compared to human analysis.
Outcomes and Implications
Readily available AI tools are currently unreliable for automated sentiment classification of health-related online text and should not be used as standalone tools for public health decision-making. However, LLMs like ChatGPT-4 show meaningful promise for rapid thematic and topic exploration, potentially serving as a triage tool to surface emerging themes from large datasets during health emergencies -- provided human oversight is maintained. Public health agencies seeking to monitor public opinion on interventions such as vaccines and mask mandates should pair AI tools with human review to ensure accuracy, contextual nuance, and mitigation of bias. Future work should address non-English datasets, domain-specific model fine-tuning, and the challenge of model drift over time.
Connect medicine with AI innovation.
No spam. Only the latest AI breakthroughs, simplified and relevant to your field.