Big Data

2019-09-09

17:30 Tim Allison: Evaluating Content/Text Extraction at Scale with Apache Tika

2019-09-10

11:15 Wangda Tan: YuniKorn: A Universal Resource Scheduler for both YARN and Kubernetes

11:15 Ryan Blue: Apache Iceberg: a table format for distributed databases

12:15 Lohit VijayaRenu: Managing Hundreds of Petabytes of data in the Cloud

12:15 John Youngseok Yang: Optimizing Big Data Pipelines with Apache Nemo (Incubating)

14:30 Zhankun Tang, Sunil Govind: First Step to Hybrid Cloud Computation: Elastic YARN and Kubernetes

14:30 vinoth chandar, Balaji Varadarajan: Apache Hudi (Incubating) : the past, present and future of efficient data lake architectures

15:30 Jackie Jiang, Seunghyun Lee: Apache Pinot (incubating): Building Realtime Analytics Applications at LinkedIn Scale

15:30 Chen Liang: Wire Encryption In HDFS: Protect Your Data From Others, Not Yourself

17:00 Dinesh Chitlangia: Ozone: Evolving HDFS Scalability to new heights & built-in GDPR Compliance

17:00 De Li: Apache Doris (incubating) -- A simple and single tightly coupled olap system

18:00 Surekha Saharan: Inside Apache Druid: Built for High-Performance Real-Time Analytics

18:00 Márton Elek: Hadoop Storage in the Cloud Native Era

2019-09-11

11:00 Jitendra Pandey, Suma Shivaprasad: Apache Hadoop 3.x State of The Union and Upgrade Guidance

11:00 Haisheng Yuan: Building BigData Query Optimization with Apache Calcite – Best Practices from Alibaba MaxCompute

12:00 Mohammad Islam, Xinli Shang: Schema-Controlled HDFS Column Encryption and Use Cases

12:00 Masahiro Ito: Lessons Learned from Leveraging Real-Time Power Consumption Data with Apache Kudu

14:15 Daoyuan Wang: Using Relational Cache to Boost Apache Spark SQL

14:15 Owen O'Malley: Protect your Private Data in your Hadoop Clusters with ORC Column Encryption

15:15 Lee Rhodes: DataSketches - The Required Toolkit for the Analysis of Big Data

15:15 Bharat Viswanadham, Anu Engineer: Building S3 over Ozone : Making a Cloud Native File System

16:45 Jagadish Venkatraman: Samza 1.0: How we scaled stream processing at LinkedIn

16:45 Vitor Wakim: From Postgres to an In-Memory Grid with Apache Ignite

17:45 Wenli Zhang: Web-based Interactive Big Data Visualization

17:45 Chengzhi Zhao: Building Data Platform for your Next Meetup Event with Apache Foundation on Cloud

2019-09-12

09:00 Kai Waehner: Spoilt for Choice – Kafka Streams vs. KSQL for Stream Processing on top of Apache Kafka

10:00 Josh Fischer: Streamlining Streaming System Management with Apache Heron

13:00 Kenneth Knowles, Julian Hyde: One SQL to Rule Them All – a Syntactically Idiomatic Approach to Management of Streams and Tables

14:00 Boyang Jerry Peng: Interactive querying of streams using Apache Pulsar

15:30 Xiaolong.Ran: Serverless Event Streaming with Pulsar Function: Use Cases and Best Practices

16:30 Valentin Zickner: Event Sourcing with Spring Boot and Apache Kafka

17:30 Jia Zhai, Penghui Li: Building Zhaopin's enterprise event bus based on Apache Pulsar