Geospatial Data Management in Apache Spark

Jia Yu, Mohamed Sarwat

2019-09-09 10:00 #geospatial

The volume of spatial data increases at a staggering rate. This talk comprehensively studies how existing works, such as GeoSpark, extend Apache Spark to uphold massive-scale spatial data. During this talk, we first provide a background introduction of the characteristics of spatial data and the history of distributed data management systems. A follow-up section presents the common approaches used by the practitioners to extend Spark and introduces the vital components in a generic spatial data management system. The third and fourth sections then discuss the ongoing efforts and experience in spatial-temporal data and spatial data analytics, respectively. The fifth part finally concludes this talk to help the audience better grasp the overall content and points out future research directions.