A statistical approach for studying the spatio-temporal distribution of geolocated tweets in urban environments

Fernando Santa, Roberto Henriques, Joaquín Torres-Sospedra, Edzer Pebesma

Research output: Contribution to journalArticlepeer-review

4 Citations (Scopus)
220 Downloads (Pure)


An in-depth descriptive approach to the dynamics of the urban population is fundamental as a first step towards promoting effective planning and designing processes in cities. Understanding the behavioral aspects of human activities can contribute to their effective management and control. We present a framework, based on statistical methods, for studying the spatio-temporal distribution of geolocated tweets as a proxy for where and when people carry out their activities. We have evaluated our proposal by analyzing the distribution of collected geolocated tweets over a two-week period in the summer of 2017 in Lisbon, London, and Manhattan. Our proposal considers a negative binomial regression analysis for the time series of counts of tweets as a first step. We further estimate a functional principal component analysis of second-order summary statistics of the hourly spatial point patterns formed by the locations of the tweets. Finally, we find groups of hours with a similar spatial arrangement of places where humans develop their activities through hierarchical clustering over the principal scores. Social media events are found to show strong temporal trends such as seasonal variation due to the hour of the day and the day of the week in addition to autoregressive schemas. We have also identified spatio-temporal patterns of clustering, i.e., groups of hours of the day that present a similar spatial distribution of human activities.

Original languageEnglish
Article number595
Issue number3
Publication statusPublished - 23 Jan 2019


  • Functional principal component analysis
  • Human activity
  • Multitype spatial point patterns
  • Negative binomial regression
  • Spatio-temporal statistics


Dive into the research topics of 'A statistical approach for studying the spatio-temporal distribution of geolocated tweets in urban environments'. Together they form a unique fingerprint.

Cite this