Biases in SDK-based GPS data use for epidemic modelling
Active project
Abstract
During the COVID-19 pandemic, a sizeable share of the literature in data-driven epidemic modeling made use of mobile GPS data passively collected by Software Development Kits (SDKs) and provided by third-party companies. Despite the substantial influence of many of these models’ predictions on high-stakes policy decisions, the methodologies and validation of these datasets have remained private and widely overlooked. In our efforts to build epidemic models with these data, we realized that their use is rife with deleterious assumptions and that results based on these data may be misleading. Consequently, we began analyzing and independently validating three different sources of SDK-based GPS data: the Kochava Collective, Veraset, and GroundTruth. Our objective is to test the robustness of specific claims relevant to COVID-19 that are justified with these datasets by independently reproducing analyses with different datasets and alternative methodologies. To further this goal, we propose to validate these data by testing for biases, comparing aggregation methodologies, and analyzing their sparsity. Additionally, we present a transparent framework for addressing biases in these datasets during their collection or when used in research projects. We are further using this data to assist the Philadelphia Department of Public Health (PDPH) in its vaccination campaign by providing human mobility insights and dashboards
Results (0)
PI
Duncan Watts; University of Pennsylvania