For the last experiment, we will show the impact of low-degree ordering applied to preloading mechanism with low-degree selection scheme. Only BFS algorithm is used for this experiment be- cause as shown in the previous experiments, all otherBFS-like algorithmshows similar tendency to BFS algorithm. Therefore, we use only BFS algorithm as representative algorithm. To see the impact of low-degree ordering, we compare only two schemes: using only preloading mechanism with low-degree selection (LS) and applying low-degree ordering to optimize (LSLO). We run the algorithm 20 times and average the execution time. The execution time is normalized to the result of using preloading mechanism only (LS).
Figure 11: Evaluation of ordering impact. LS is low-degree selection only and LSLO is low-degree selection combined with low-degree ordering.
As demonstrated in Figure 11, applying ordering is an obvious winner compared to the preloading-only mechanism. Compared to the LS, LSLO shows maximally 57% faster execution time and 29% faster on average. Moreover, original preloading mechanism uses index map that hasN entries if there areN vertice’s edges are preloaded. To store the map information, it needs 8N size of memory. Generally, low-degree vertices are cached, the size ofN is large which means memory cost is also large. However, if we use low-degree ordering mechanism, we need tuple information consist of(degree, start offset, number of vertices). Using the tuple information and if we let the number of tuple as Nt which is the same as the number of different degrees, the memory consumption will be12Ntwhich is much smaller than8N. (Nt<< N) Therefore, using the low-degree ordering can optimize performance not only execution time, but also memory consumption.
20
VI Related Works
Disk-based Graph Processing System
GraphChi is the first disk-based graph processing system that made large-scale graph analysis possible on a single machine [6]. GraphChi is mainly designed for HDDs and eliminates random disk accesses with its Parallel Sliding Windows technique. TurboGraph is a disk-based graph engine for SSDs [7]. In TurboGraph, adjacency lists of requested vertices are randomly accessed in each iteration. Its pin-and-slide execution model overlaps random I/O with CPU computa- tion and fully utilizes the I/O parallelism inherent within SSDs. Its successor, TurboGraph++, further extends the system for a distributed cluster [8]. As TurboGraph and GraphChi make use of the page cache for its random I/O, our optimizations can be easily applied to their systems.
While the vertex-centric computation model is widely adopted in large-scale graph process- ing, recent studies on disk-based graph systems propose an alternativeedge-centriccomputation that streams edges into memory to exploit the locality of edge access [9, 10]. The edge-centric model shows good performance for algorithms accessing the entire graph repeatedly. However, the model is not as efficient for algorithms that iteratively access different subsets of a graph as is done with BFS-like algorithms. Recently, Maass et al. proposed a hybrid model in Mosaic that supports both the vertex-centric and edge-centric models at the same time to efficiently utilize CPUs and co-processors [11].
In semi-external graph engines, vertex attributes are stored in main memory for fast up- dates [1,12–14]. Pearce et al. proposed asynchronous optimization techniques for graph traversal algorithms for semi-external graph processing [13]. FlashGraph implements several I/O opti- mizations for SSDs and SSD arrays such as merging I/O requests for higher throughput and overlapping I/O and computation [1]. We build on top of these optimizations and propose meth- ods that exploit the structural properties of the input graph.
Several other I/O optimization methods for disk-based graph processing have recently been proposed. Vora et al. employs a dynamic partitioning scheme that prevents loading unnecessary edges on disk [15]. GridGraph supports 2D edge partitioning to reduce I/O access [10]. In Graphene, a bitmap based asynchronous I/O optimization is applied to efficiently merge small I/O requests [16]. Our proposed optimization techniques are applicable on top of these I/O optimizations as we have shown with Graphene.
Main Memory Graph Processing
For large-scale graph processing, vertex-centric computation model was initially proposed in the context of distributed in-memory systems. Google’s Pregel is the first system adopting the
21
vertex-centric computation model [3]. GraphLab is a distributed machine learning framework that employs a variant of the vertex-centric model that supports asynchronous computation [17].
Its successor, PowerGraph, adopts an efficient graph partitioning scheme that considers the power-law degree distributions of real-world graphs [18].
To ease the programming difficulty of large-scale graph analysis, SociaLite supports declar- ative query language based on Datalog [19]. Wang et al. also demonstrate the performance and scalability of Datalog-based graph processing [20].
Galois is a parallel graph processing system based on an implicitly parallel vertex itera- tor [21]. Green-Marl is a domain-specific language for writing parallel graph algorithms for shared-memory [22]. Ligra is a light-weight framework with low-level primitives to implement graph processing systems [23].
Wei et al. studied optimizing the performance of graph algorithms for main memory graph processing [2]. The authors develop a graph ordering calledGorderthat optimizes the locality of updating vertex attributes. While Gorder is designed for main memory graph systems, Norder is designed to optimize the performance of disk-based graph engines.
22
VII Conclusion
In this paper, we analyzed graph algorithms on disk-based graph engines. The results tell us that we can categorize graph algorithms into three groups based on I/O performance and request patterns: BFS-like algorithm,PageRank-like algorithm and Subgraph mining algorithm.
We figured out that for BFS-like algorithm, the indifference of the number of I/O requests and the original page cache hit ratio is relatively low. Based on the results, we concluded that the original page cache is not suitable solution to optimize I/O performance for graph processing.
In this paper, we proposeSelective Preloading Mechanism, to optimize the I/O performance for graph processing. We reduced unnecessary memory consumption and allocation and evic- tion overhead using vertex-granular preloading mechanism. We selected two types of vertices:
Greedy and Low-degree. Both can optimize I/O performance compared to page cache and can be used in proper situations. We also proposed low-degree graph ordering that can be utilized with preloaded mechanism using low-degree selection scheme which reduces index look-up time by making possible to calculate the indexes.
The optimization has evaluated four real-world graphs. Greedy selection shows performance improvement up to 76% faster maximally and 28% on average compared to page cache. Low- degree selection shows performance improvement up to 75% faster maximally and 19% on av- erage compared to page cache. By applying the low-degree ordering for low-degree selection, the performance is improved 57% faster maximally and memory consumption also decreased efficiently. Moreover, we figured out that the page utilization has increased, which indicates the I/O optimization is the key factor of disk-based graph engine.
23
References
[1] D. Zheng, D. Mhembere, R. Burns, J. Vogelstein, C. E. Priebe, and A. S. Szalay, “Flash- Graph: Processing billion-node graphs on an array of commodity SSDs,” in Proceedings of the USENIX Conference on File and Storage Technologies (FAST 15), 2015, pp. 45–58.
[2] H. Wei, J. X. Yu, C. Lu, and X. Lin, “Speedup graph processing by graph ordering,” in Proceedings of the International Conference on Management of Data (SIGMOD 16), 2016, pp. 1813–1828.
[3] G. Malewicz, M. H. Austern, A. J. Bik, J. C. Dehnert, I. Horn, N. Leiser, and G. Czajkowski,
“Pregel: a system for large-scale graph processing,” in Proceedings of the ACM SIGMOD International Conference on Management of Data (SIGMOD 10), 2010, pp. 135–146.
[4] U. Brandes, “A faster algorithm for Betweenness Centrality,” The Journal of Mathematical Sociology, vol. 25, no. 2, pp. 163–177, 2001.
[5] “The koblenz network collection.” http://konect.uni-koblenz.de/.
[6] A. Kyrola, G. Blelloch, and C. Guestrin, “GraphChi: Large-scale graph computation on just a PC,” in Proceedings of the USENIX Symposium on Operating Systems Design and Implementation (OSDI 12), 2012, pp. 31–46.
[7] W.-S. Han, S. Lee, K. Park, J.-H. Lee, M.-S. Kim, J. Kim, and H. Yu, “TurboGraph: a fast parallel graph engine handling billion-scale graphs in a single PC,” in Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD 13), 2013, pp. 77–85.
[8] S. Ko and W.-S. Han, “TurboGraph++: A scalable and fast graph analytics system,” in Proceedings of the 2018 International Conference on Management of Data (SIGMOD 18), 2018, pp. 395–410.
[9] A. Roy, I. Mihailovic, and W. Zwaenepoel, “X-stream: Edge-centric graph processing using streaming partitions,” in Proceedings of the Symposium on Operating Systems Principles (SOSP 13), 2013, pp. 472–488.
[10] X. Zhu, W. Han, and W. Chen, “GridGraph: Large-scale graph processing on a single ma- chine using 2-level hierarchical partitioning.” in Proceedings of the USENIX Annual Tech- nical Conference (ATC 15), 2015, pp. 375–386.
24
[11] S. Maass, C. Min, S. Kashyap, W. Kang, M. Kumar, and T. Kim, “Mosaic: Processing a trillion-edge graph on a single machine,” inProceedings of the Twelfth European Conference on Computer Systems (EuroSys 17), 2017, pp. 527–543.
[12] P. Kumar and H. H. Huang, “G-store: High-performance graph store for trillion-edge pro- cessing,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC 16), 2016, pp. 830–841.
[13] R. Pearce, M. Gokhale, and N. M. Amato, “Multithreaded asynchronous graph traversal for in-memory and semi-external memory,” inProceedings of the ACM/IEEE International Conference for High Performance Computing, Networking, Storage and Analysis (SC 10), 2010, pp. 1–11.
[14] Z. Shao, J. He, H. Lv, and H. Jin, “Fog: A fast out-of-core graph processing framework,”
International Journal of Parallel Programming, pp. 1–14, 2016.
[15] K. Vora, G. H. Xu, and R. Gupta, “Load the edges you need: A generic I/O optimization for disk-based graph processing.” in Proceedings of the USENIX Annual Technical Conference (ATC 16), 2016, pp. 507–522.
[16] H. Liu and H. H. Huang, “Graphene: Fine-grained IO management for graph computing,”
in Proceedings of the USENIX Conference on File and Storage Technologies (FAST 17), 2017, pp. 285–300.
[17] Y. Low, D. Bickson, J. Gonzalez, C. Guestrin, A. Kyrola, and J. M. Hellerstein, “Distributed GraphLab: a framework for machine learning and data mining in the cloud,” Proceedings of the VLDB Endowment, vol. 5, no. 8, pp. 716–727, 2012.
[18] J. E. Gonzalez, Y. Low, H. Gu, D. Bickson, and C. Guestrin, “PowerGraph: Distributed graph-parallel computation on natural graphs.” in Proceedings of the Symposium on Oper- ating Systems Design and Implementation (OSDI 12), 2012, pp. 17–30.
[19] J. Seo, J. Park, J. Shin, and M. S. Lam, “Distributed SociaLite: A datalog-based language for large-scale graph analysis,” Proceedings of the VLDB Endowment, vol. 6, no. 14, pp.
1906–1917, 2013.
[20] J. Wang, M. Balazinska, and D. Halperin, “Asynchronous and fault-tolerant recursive dat- alog evaluation in shared-nothing engines,” Proceedings of the VLDB Endowment, vol. 8, no. 12, pp. 1542–1553, 2015.
[21] D. Nguyen, A. Lenharth, and K. Pingali, “A lightweight infrastructure for graph analytics,”
in Proceedings of the ACM Symposium on Operating Systems Principles (SOSP 13), 2013, pp. 456–471.
25
[22] S. Hong, H. Chafi, E. Sedlar, and K. Olukotun, “Green-Marl: a DSL for easy and efficient graph analysis,” ASPLOS XVII Proceedings of the seventeenth international conference on Architectural Support for Programming Languages and Operating Systems, pp. 349–362, 2012.
[23] J. Shun and G. E. Blelloch, “Ligra: a lightweight graph processing framework for shared memory,” inProceedings of the 18th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming (PPoPP 13), 2013, pp. 135–146.
26