DRS: Dynamic Resource Scheduling for Real-Time Analytics over Fast Streams

Tom Z.J. Fu, Jianbing Ding, Richard T.B. Ma, Marianne Winslett, Yin Yang, Zhenjie Zhang

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

81 Citations (Scopus)

Abstract

In a data stream management system (DSMS), users register continuous queries, and receive result updates as data arrive and expire. We focus on applications with real-time constraints, in which the user must receive each result update within a given period after the update occurs. To handle fast data, the DSMS is commonly placed on top of a cloud infrastructure. Because stream properties such as arrival rates can fluctuate unpredictably, cloud resources must be dynamically provisioned and scheduled accordingly to ensure real-time response. It is essential, for the existing systems or future developments, to possess the ability of scheduling resources dynamically according to the current workload, in order to avoid wasting resources, or failing in delivering correct results on time. Motivated by this, we propose DRS, a novel dynamic resource scheduler for cloud-based DSMSs. DRS overcomes three fundamental challenges: (a) how to model the relationship between the provisioned resources and query response time (b) where to best place resources, and (c) how to measure system load with minimal overhead. In particular, DRS includes an accurate performance model based on the theory of Jackson open queueing networks and is capable of handling arbitrary operator topologies, possibly with loops, splits and joins. Extensive experiments with real data confirm that DRS achieves real-time response with close to optimal resource consumption.

Original languageEnglish
Title of host publicationProceedings - 2015 IEEE 35th International Conference on Distributed Computing Systems, ICDCS 2015
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages411-420
Number of pages10
ISBN (Electronic)9781467372145
DOIs
Publication statusPublished - 22 Jul 2015
Event35th IEEE International Conference on Distributed Computing Systems, ICDCS 2015 - Columbus, United States
Duration: 29 Jun 20152 Jul 2015

Publication series

NameProceedings - International Conference on Distributed Computing Systems
Volume2015-July

Conference

Conference35th IEEE International Conference on Distributed Computing Systems, ICDCS 2015
Country/TerritoryUnited States
CityColumbus
Period29/06/152/07/15

Keywords

  • data stream analytics
  • resource scheduling

Fingerprint

Dive into the research topics of 'DRS: Dynamic Resource Scheduling for Real-Time Analytics over Fast Streams'. Together they form a unique fingerprint.

Cite this