• Tidak ada hasil yang ditemukan

Selective Preloading Optimization for Disk-based Graph Engines

N/A
N/A
Protected

Academic year: 2023

Membagikan "Selective Preloading Optimization for Disk-based Graph Engines"

Copied!
37
0
0

Teks penuh

(1)

저작자표시-비영리-변경금지 2.0 대한민국 이용자는 아래의 조건을 따르는 경우에 한하여 자유롭게

l 이 저작물을 복제, 배포, 전송, 전시, 공연 및 방송할 수 있습니다. 다음과 같은 조건을 따라야 합니다:

l 귀하는, 이 저작물의 재이용이나 배포의 경우, 이 저작물에 적용된 이용허락조건 을 명확하게 나타내어야 합니다.

l 저작권자로부터 별도의 허가를 받으면 이러한 조건들은 적용되지 않습니다.

저작권법에 따른 이용자의 권리는 위의 내용에 의하여 영향을 받지 않습니다. 이것은 이용허락규약(Legal Code)을 이해하기 쉽게 요약한 것입니다.

Disclaimer

저작자표시. 귀하는 원저작자를 표시하여야 합니다.

비영리. 귀하는 이 저작물을 영리 목적으로 이용할 수 없습니다.

변경금지. 귀하는 이 저작물을 개작, 변형 또는 가공할 수 없습니다.

(2)

Master’s Thesis

Selective Preloading Optimization for Disk-based Graph Engines

Jung Hyun Kim

Department of Electrical and Computer Engineering (Computer Science and Engineering)

Graduate School of UNIST

2019

[UCI]I804:31001-200000179674 [UCI]I804:31001-200000179674 [UCI]I804:31001-200000179674 [UCI]I804:31001-200000179674 [UCI]I804:31001-200000179674 [UCI]I804:31001-200000179674

(3)

Selective Proloading Optimization for Disk-based Graph Engines

Jung Hyun Kim

Department of Electrical and Computer Engineering (Computer Science and Engineering)

Graduate School of UNIST

(4)
(5)
(6)

Abstract

As the web and social network market is getting bigger, large-scale graphs are flooding out continuously. Therefore, there are many recent studies to effectively deal with such large-scale graphs. Among them, a disk-based graph engine is used to process a large-scale graph on a single machine. It can process a large-scale graph with a relatively small memory size by storing the graph data on disks. However, it needs to read the graph data from disk, the I/O performance critically affects to the overall processing performance. To reduce the performance degradation due to a huge number of I/Os, the state of art graph engine utilizes page cache to reduce the number of direct I/O request to disk. However, because of the characteristics of graph process- ing, page cache hit ratio is low. Moreover, there exist allocation and eviction overhead even page cache’s performance is poor. It other words, for graph processing, page cache cannot efficiently optimize the overall performance.

This paper presents Selective Preloading: a novel I/O optimization technique for graph pro- cessing system. In the paper, different to the page cache, graph processing system preloads vertices’ edge lists whose page utilization is low on the memory space to utilize memory space better and reduce the overhead due to page cache allocation and eviction. We propose two selection schemes: Greedyand Low-Degree. We can choose selection scheme based on the situa- tion.Greedyselection gives higher performance but also higher pre-processing time. In contrast, Low-degreeselection have low pr-processing time but lower performance compared toGreedyse- lection scheme. To optimizeLow-degreeselection’s performance, we propose low-degree ordering to reduce index look-up time in preloaded vertices. We demonstrate that BFS-like algorithms show the worst I/O locality among graph algorithms, so we set the target algorithms as BFS-like algorithms.

We have extensively evaluated our optimization with four real-world graphs. Our preloading optimization improve performance up to 76%, average 29% compared to page cache when using Greedy selection and up to 75%, average 19% when usingLow-degree selection. When applying low-degree ordering to Low-degree selection, the performance improved 57% compared to using only preloading mechanism. The memory consumption is also efficiently decreased when we use low-degree ordering. Moreover, we demonstrate that out optimizations utilizes pages better than page cache: utilization increased two to three times compared to the page cache.

(7)
(8)

Contents

I Introduction . . . 1

II Graph Processing Systems . . . 3

2.1 In-memory & Distributed Graph Processing System . . . 3

2.2 Disk-based Graph Processing System . . . 3

III Analysis of Graph Processing . . . 5

3.1 Graph Processing on Disk-based Graph Processing System . . . 5

3.2 Uniform Edge List Accesses . . . 7

3.3 Low Utilization of Page Cache . . . 8

3.4 Impact of Graph Ordering . . . 9

IV Selective Preloading . . . 11

4.1 Preloading Mechanism . . . 11

4.2 Vertex Selection Schemes . . . 12

4.3 Low-Degree Graph Ordering . . . 14

V Evaluation . . . 16

5.1 Execution Time . . . 17

5.2 Analysis of Page Utilization . . . 19

5.3 Ordering Effectiveness . . . 20

2

(9)

VI Related Works . . . 21 VII Conclusion . . . 23 References . . . 24

(10)

List of Figures

1 Architecture of disk-based graph processing system (Semi-external-memory system) 4 2 Distribution of the number of edge list accesses: aggregated vertices with the same

in-degree. Darker point represents larger number of vertices than lighter point . . 7 3 Page cache hit ratio and normalized execution time of various graph algorithms

while scaling up the page cache size from 5% to 30% of the input graph size . . . 8 4 Example of two different graph orderings . . . 10 5 Disk access pattern in the BFS algorithm with starting vertex as 1 and 3 for each. 10 6 Normalized execution time of various graph algorithms according to the four

different graph orderings . . . 10 7 Architecture of pre-loading mechanism. The above architecture shows the original

page cache and below architecture with pre-loading mechanism. . . 12 8 Scoring result for LiveJournal graph with BFS algorithm. Score for vertices of the

same out-degree are averaged. . . 14 9 Normalized execution time of various selection schemes. Six BFS-like algorithms

and four real-world graphs are used for experiments. The execution time is nor- malized to the page cache only . . . 18 10 Evaluation of page utilization with various selections for preloading . . . 19 11 Evaluation of ordering impact. LS is low-degree selection only and LSLO is low-

degree selection combined with low-degree ordering. . . 20

(11)

I Introduction

As the social network service is getting more popular, bunch of data is flooding out. More- over, to deal with big-data, the data is represented using graph format to indicate the relations between data. It can be used to analyze current trend and predict the future behaviors. There- fore, recent studies about graph analytic systems to analyze large-scale network. Among them, because of the disk size is getting larger and price is getting cheaper, disk-based graph engine which aims to deal with large-scale graphs in a single machine is generally used.

Disk-based graph engines store input graphs on external storage such as HDDs and SSDs.

When the graph analytic algorithm processes, the graph engine keeps requesting the graph data to disk. Vertex attributes are also stored in disk and updated as algorithms proceed. However, because the vertex attributes size is much smaller than the graph data, recent trend is keeping the vertex attributes in memory to improve the overall graph processing performance. This kind of system is called semi-external graph engines.

Disk-based graph engine, particularly for semi-external graph engine, the performance is dependent on the disk I/O performance to read the graph data which is in disk. Therefore, several optimization techniques are suggested to improve the I/O performance. Among them, FlashGraph, the state of the art disk-based graph engine uses page cache to reduce total request amount to disk. It reads graph data from disk by page granularity. In other words, if the re- questing page is already in page cache, it doesn’t need to request it to disk but read from page cache which has much lower latency. However, for the BFS-like graph algorithms, they never visit the node that already has visited. It means that each vertex’s edge list requests only once in a whole processing. This leads low page cache hit ratio. Even the page cache doesn’t work well, there still exists the allocation and eviction overhead which is not negligible.

In this paper, we presentSelective Preloading Mechanism: a novel solution for low utilization of page cache. We aim our optimization toBFS-like algorithmbecausePageRank-like algorithm are sequentially access their edge lists which doesn’t suit to our model and Subgraph mining algorithmhave locality which means using page cache well. Here are our contributions.

Contributions At first, we propose preloading mechanism. It preloads vertices’ edge lists which make page utilization of page cache poor. This can optimize graph processing system by increasing page utilization and reducing the unnecessary page allocation and eviction overhead in page cache. Then, we propose two selection schemes: Greedy selection and Low-degree se- lection.Greedy selection finds sub-optimal vertices for page cache andLow-degree selectionjust preloads low-degree vertices’ edge lists based on analysis. Both can be used in a proper situation.

Moreover, we propose low-degree ordering for graph storing scheme which can improve graph 1

(12)

processing performance when used with low-degree selecting preloading mechanism. This can be done by eliminating index look up time.

We implemented and extensively evaluated the optimization with four real-world graphs.

We split the original page cache into two parts, one part is for page cache and the other is for selective preloading. We left the page cache space because not all the vertices are malfunctioning with page cache. We only preloads the vertices which shows low page utilization. Moreover, as we aimed for BFS-like algorithm, which are the most affected, our evaluations is focused on those algorithms. We show that our optimization improves performance up to 76% only using preloading mechanism when using Greedy selection. Moreover, when combined with low-degree ordering, the performance increases 57% compared to preloading only mechanism.

Paper Organization The paper is composed as following order. Section II briefly intro- duces the graph processing systems, specially, focus on disk-based graph processing systems which we are aimed for. Section III, we analyze the graph processing on disk-based graph pro- cessing systems and shows the ineffectiveness of current systems. In Section IV, we propose our optimization:Selective Preloadingin the following order: preloading mechanism, vertex selection schemes and low-degree ordering. We evaluate our optimization in Section V. At last, we refers the works that are related to our optimization in Section VI and concludes our paper in Section VII.

2

(13)

II Graph Processing Systems

As graph data analysis is being an important to figure out important patterns or meanings of large-scale data, systems for processing such a large scale graph are needed. Among them, we’ll discuss about three graph processing systems, differentiate according to the graph data storing scheme: in-memory graph processing system, distributed graph processing system and disk-based graph processing system.

2.1 In-memory & Distributed Graph Processing System

In-memory graph processing system is the simplest one among them. In this system, all the vertex and edge data are located on memory space. This system offers user the easiest program- ming APIs because they don’t have to consider about requesting to other machines or disk.

Moreover, networking overhead in memory is low, the processing speed is the fastest among three systems. However, it uses only a single machine which means the memory space scalability is limited, it cannot process large-scale graph data.

To resolve the problem of in-memory graph processing system, distributed graph processing system has proposed. Distributed graph processing system uses multiple machines to store the large-scale graph data by using partitioning scheme. Each machine processes algorithm using the partial graph data. Then, the main machine aggregates the results generated by multiple machines. It can efficiently deal with a large-scale graph data and the processing performance can be improved linearly according to the number of machines. However, it is hard to construct the environments due to the cost and as the number of machine increases, networking overhead is getting larger which reduces the performance improvement.

2.2 Disk-based Graph Processing System

Even distributed graph processing system can process large-scale graphs, there exists prob- lem of building such environment. Therefore, disk-based graph processing system has proposed to process large-scale graphs in a single machine. Disk-based graph processing system stores graph data in external memory which is much cheaper and can store all graph data in a single machine. The processing mechanism will be explained in section 3.1.

We can separate disk-based graph processing system based on storing graph data scheme:

external-memory system and semi-external-memory system. There are two types of graph data:

adjacency list and vertex attributes. Adjacency list is graph algorithm’s traversal data. It repre- sents graph layout composed of the set of edges, which makes the system can propagate messages following the edges. Vertex attribute data is graph algorithm’s computational data. It represents

3

(14)

computational intermittent or final result and usually defined by user algorithm. The external- memory system is storing all adjacency list data and vertex attributes data in disk because memory space is not enough to hold whole data. It has to request not only the adjacency list data but also vertex attributes’ data to compute. However, as memory space is getting larger, it can hold vertex attributes data. By using the notion, the semi-external-memory system has pro- posed which stores adjacency data in external memory while holding vertex attributes data on main memory. This can reduce the number of I/O requests which occurs due to vertex attributes.

We used semi-external-memory system which is much more the state of the art technique. Figure 1 shows the general architecture of disk-based graph processing system.

Figure 1: Architecture of disk-based graph processing system (Semi-external-memory system)

4

(15)

III Analysis of Graph Processing

In this paper, we will distinguish graph analytic algorithms into three types: BFS-like algo- rithm, PageRank-like algorithm and Subgraph matching algorithms. The criteria for division is processing mechanism and I/O access pattern. BFS-like algorithm are those that start from a given set of vertices and propagates activation state by traversing their neighbor vertices, recur- sively. The algorithms inBFS-like algorithmare breadth first search, diameter of the input graph, betweenness centrality, shortest paths, computation of information diffusion, etc.PageRank-like algorithm is algorithms that enumerate vertices in a graph and then compute their attributes based on their neighbors’ attributes. This group includes PageRank, personalized PageRank, label propagation for detecting communities, etc. Lastly, Subgraph mining algorithm is those that find query graphs which may have significant meaning in a given graph. It contains algo- rithms such as counting triangles, mining motifs, etc. The I/O access pattern is also different among these algorithms. For example, I/O access is more random inBFS-like algorithm than in other algorithms, whilePageRank-like algorithmis less affected by the layout of graphs on disks than the other algorithms. In this paper, we observed that BFS-like algorithmshows the worst utilization of page cache, we focus on these algorithms for our optimization.

In this section, we will show the overall analysis of graph processing mechanism. At first, in section 3.1, we will discuss how the graph processing works on disk-based graph engines, specifi- cally, semi-external graph engine. Then, in section 3.2, we will demonstrate the uniform access of the number of edge references according to the degree of vertices. At section 3.3, we will discuss the ineffectiveness of page cache which the state-of-the-art disk-based graph processing systems are using to optimize I/O performance. At section 3.4, we explain what the graph ordering is and how it impacts on the processing performance.

3.1 Graph Processing on Disk-based Graph Processing System

Most graph processing systems use vertex-centric programming model for graph processing (Shown as in Algorithm 1). Graph algorithms written in the vertex-centric model run iteratively, with a varying subset of vertices activated per iteration depending on the algorithm type. The activation of vertices can be induced generally in two ways: receiving messages in the previous iteration or algorithm itself make all vertices to be activated. When a vertex is activated in an iteration, the system processes vertex program with the vertex attributes to update. When the process is done, it requests edge list when it needs to propagate the updated values to neighbor vertices. The request makes graph processing system send a request to disk to retrieve the edge- list. The request and retrieve mechanism is done in page granularity. The page size is normally 1KB to a few MBs. The edge list is in page cache to reduce the direct I/O to disk. It is controlled by either graph processing system or the file system itself.

5

(16)

Algorithm 1 Vertex-Centric Programming Model V : Vertex set

Va: Activated vertices set→Va⊆V E : Edge set →(V ×V)

G: Graph Instance →(V, E)

1: Va← Start vertices

2: whileVa is not empty do

3: Vtemp ←φ

4: foru∈Va do

5: process(u)

6: Nu ← Neighbors ofu

7: forv∈Nu do

8: if vneeds update then

9: Send message tov withu’s updated value

10: Vtemp ←v

11: end if

12: end for

13: end for

14: Va ←Vtemp

15: end while

The graph data layout stored in disk is generally represented in sequential order. The order is determined by the vertex ID. There are some graph processing systems which offer data par- titioning schemes to reduce disk I/O. The most common partitioning is vertex-cut or edge-cut partitioning. However, some systems optimize partitioning such as hybrid-cut partitioning to re- duce disk I/O which directly leads to performance improvement in disk-based graph processing system. Moreover, when adjacent pages are requested due to edge list requests with close vertex IDs, the requests may be merged for higher throughput.

When the edge lists of activated vertices are accessed, pages containing these edges are loaded into the page cache. As page units are large, edges of other activated vertices may also be in the page which is already loaded and hence, we don’t need to send a request to disk. To achieve high cache utilization, which leads to the high performance, we want the page cache to contain as many edge lists of activated vertices as possible. This requires the vertices in the input graph to be ordered such that the edges of activated vertices in the same iteration are stored closely.

6

(17)

3.2 Uniform Edge List Accesses

In this section, we will discuss about the number of edge list accesses according to the de- gree of vertices. In this paper, we use preloading mechanism to reduce the request to disk. For that, selecting proper vertices is important for performance improvement. Simply, selecting and loading vertices which requests to disk frequently seems efficient. Moreover, degree of vertices is critically affect to utilization of page because of limited page size. Therefore, analyzing access pattern for vertices according to vertices’ degree is important.

(a) (b)

(c) (d)

Figure 2: Distribution of the number of edge list accesses: aggregated vertices with the same in-degree. Darker point represents larger number of vertices than lighter point

Figure 2 shows the average and standard deviation of the aggregated number of edge list request of vertices which have the same in-degree. Initially, we assumed that the number of requests of vertices would be roughly proportional to its in-degree, because requests are sent over edges in a vertex-centric model. However, Figure 2 demonstrates that the gap between low in-degree vertices and high in-degree vertices (blue dots) is not big, except for triangle count- ing(TC). For the BFS experiment, it shows variance in the number of requests because we run BFS algorithm with randomly selected source vertices and aggregate the total request counts.

Triangle counting shows considerably large variance, which distinguishes it from the others.

7

(18)

The reason for this uniformity in edge list access is that when runningBFS-like algorithmin the vertex-centric computation model, the edge list of a vertex is accessed only once per itera- tion. Moreover, only a subset of the activated vertices needs to send messages to their neighbor vertices. Thus, the maximum number of edge request is generally smaller than the number of total iterations.

The observation that there are no substantial differences in the number of edge requests tells us that strategies such as simply storing frequently accessed edge lists in memory are not effective approaches for improving performance. Based on this observation, we will suggest vertex selection scheme suitable for graph processing.

3.3 Low Utilization of Page Cache

As BFS-like algorithmrunning on disk-based system has only a subset of active vertices on each iteration, I/O access locality is poor, which makes the page cache ineffective for these kinds of algorithms. To quantify this notion, we performed analysis using the algorithms provided in FlashGraph [1] and observe the hit ratio and normalized execution time by changing the size of page cache. We used five algorithms: breadth-first search (BFS), diameter of the graph (DIAM) and betweenness centrality (BC) for BFS-like algorithm, PageRank (PR) for PageRank-like al- gorithmand triangle counting (TC) forSubgraph-mining algorithm.

Figure 3: Page cache hit ratio and normalized execution time of various graph algorithms while scaling up the page cache size from 5% to 30% of the input graph size

Figure 3 depicts the results with page cache sizes ranging from 5% to 30% of the input graph run on the Twitter dataset. All algorithms show hit ratios of less than 30% except for triangle counting. Triangle counting algorithm isSubgraph-mining algorithmwhich utilizes spatial local- ity of graph well. More importantly, for BFS-like algorithm, the performance is improved only by 5% to 10% even with 6 factors increased in page cache size.

8

(19)

Our conclusion is that for BFS-like algorithm, page cache doesn’t efficiently work to reduce I/O performance due to the characteristic of algorithms. Moreover, because it is “cache” which executes in run-time, there exist allocation and eviction cost, even the utilization of the page is poor. Therefore, we propose Selective Preloading Mechanism to mitigate this problems. The details of optimization proposed in section 4.1

3.4 Impact of Graph Ordering

Graph ordering is a technique that reassign the vertices’ ID. Figure 4 shows the example of two different graph orderings. This technique may affect to the I/O performance because for disk-based graph processing systems, graph data stored in disk sorted by vertex ID. Therefore, the access pattern on disk can be relatively sequential according to how the vertices are or- dered. Figure 5 shows the difference of disk access pattern according to the graph layout. For ordering (a), the disk access pattern is relatively sequential compared to (b). Both SSDs and HDDs show faster performance with sequential reads than random reads. Thus, how a graph is stored and accessed by the running algorithms has substantial impact on performance. In this section, we perform experiments to help us understand the performance impact of graph layouts.

For the experiments, we restructure a graph in four different orderings and measure the performance of the graph algorithms. Here are the four orderings: randomly assign the vertex IDs (Random), sorted by PageRank values (PageRank), sorted in low-degree vertex with lower vertex ID (LowD) and Gorder [2] which was proposed to improve the locality of access to ver- tices and their edge lists for main memory graph systems. Then, the graph is stored in CSR (Compressed Sparse Row) format, where the edge lists of vertices are ordered by their vertex IDs and stored sequentially.

Figure 6 compares the performance of three BFS-like algorithms with four orderings on LiveJournal dataset. We can see that the algorithms perform consistently better with Gorder than with random ordering. Moreover, we can see that low-degree ordering shows reasonable performance compared to Gorder. Clearly, ordering strongly affects the performance of BFS-like algorithms. In Section 4.3, we presents graph ordering that optimizes graph processing perfor- mance when combined with our preloading mechanism.

9

(20)

Figure 4: Example of two different graph orderings

Figure 5: Disk access pattern in the BFS algorithm with starting vertex as 1 and 3 for each.

Figure 6: Normalized execution time of various graph algorithms according to the four different graph orderings

10

(21)

IV Selective Preloading

In this section, we will briefly describe about our optimization: Selective Preloading. The idea and implementation is simple. We separate page cache space into two parts and assign one to be used for preloading vertices which makes the page cache utilization poor. After than, we select vertices and put the edge lists of them into the preloading memory space. We load two types of vertices’ edge lists: greedy selected vertices and low-degree vertices. We will explain how this preloading mechanism can solve the original page cache’s problem and why the vertices are selected. Moreover, we will re-order the graph that can optimize our low-degree vertices preloading mechanism.

4.1 Preloading Mechanism

The idea is based on the observation that for page granularity caching scheme allocates un- necessary memory. The page cache may evict the page even the page used only a small part of memory. This seems can be solved by simply allocating not page granularity but vertex gran- ularity. However, there are two problems of vertex granular caching. 1) The size of edge lists are different each other. So, if we cache in vertex granularity, when a cache entry needs to be changed, memory fragmentation problem can occur which may leads additional overhead. 2) BFS-like algorithmdoesn’t have temporal locality which means that when a vertex is activated once, it takes long time to be activated or never activates again. Hence, the possibility that the vertex is re-activated before the page evicts from the page cache is very low. Even such a low efficiency, page cache still needs dynamic allocation and eviction cost.

To alleviate the inefficiency, we propose preloading technique which loads the edge lists that makes the page cache utilization poor before the process starts. The preloading mech- anism improves the I/O performance by reducing the unnecessary memory consumption and allocate/eviction cost of the page which may use only a small part of it. We separate page cache space into two parts and use one as the original page cache and the other for preloading space.

Figure 7 shows the architecture of original system and the system with preloading mechanism.

As shown in the figure, 3rd and 4th pages show low utilization of page cache: using tiny part of page. Therefore, we select such harmful vertices’ edge lists and loads on the separated space. We left the page cache area because there are some edge lists or pages that are already utilizing page cache well such as 1st page in Figure 7. In other words, selecting the vertices whose edge lists are needed to be preloaded is an important scheme. We will give two selection schemes in section 4.2.

11

(22)

Figure 7: Architecture of pre-loading mechanism. The above architecture shows the original page cache and below architecture with pre-loading mechanism.

4.2 Vertex Selection Schemes

As we mentioned in the previous section, we will discuss which vertices should be selected for preloading. First, we formulate the penalty of using page cache. The goal of our preloading mechanism is to minimize the equation: i.e. the page cache can be utilized as well as possible.

Equation (1) shows our target model.

minimize

P R F(P R) =X

v∈V

X

(v,u1)∈E u1∈P R/

1 r(u1)di(u1)

1 P

(v,u2)∈E u2∈P(u1)

u2∈P R/

do(u2)

subject to X

v∈P R

deg(v)≤M, X

v∈P R

deg(v)≥M−

(1)

• is a small positive number

• P Rrepresents the set of preloaded vertices

• E is the set of all edges in the given graph

• P(u) is the set of vertices whose edge lists are stored in the same page as vertex u

• r(u) is the expected number of requests to the page whereuis stored, which we assume to be proportional to the number of vertices whose edge lists are stored in the page

• di(u),do(u) are the in-degree and out-degree of vertex u, respectively

• M represents the size of available memory for preloading

12

(23)

Function F represents the penalty occurred by misuse of pages in the cache. That is, if the cache is fully utilized, there is no penalty. Hence, our goal is to minimized F.

However, solving the equation is NP-hard problem because it’s a similar problem to Knap- sack problem. Therefore, we used general way to solve NP-hard problem: Greedy algorithm.

By using the greedy algorithm, we can find suboptimal vertices to be preloaded. However, for LiveJournal(LJ), which is relatively small dataset, the pre-processing time is about188seconds, which is much larger than the actual BFS runtime:3∼4 seconds. Therefore, we propose much simpler selection algorithm to reduce the pre-processing time.

There are two things to be considered when we select vertices: 1) High probability of re- questing it’s edge list. This is because, if a selected vertex never requests it’s edge list, we don’t actually use the edge list which is already resides on the memory, i.e. waste of memory space. 2) Low memory cost when we load the vertex’s edge list. The preloading memory space is limited, we should consider the size of edge list of each vertex. Moreover, if a vertex has high-degree, it may already utilize page cache well. Therefore, we can generate a score equation by considering two conditions in Equation (2):

s(v) = f(vin)

vout (2)

• s(v) is the score of v

• vin and vout are in-degree and out-degree of the vertex v

• f(vin) gives a probability of a vertex with in-degree asvin request it’s edge list

f function means the values of the graph in Figure 2. By using the scoring function, we can assume that the vertices that return higher score can optimize I/O performance when preloaded on the memory space. We applied the scoring function to real graph, LiveJournal(LJ). The result is illustrated in Figure 8.

As Figure 8 illustrates, compared to high-degree vertices, low-degree vertices shows higher scores. However, to get the exact result of scoring function, we need to calculate the probability function f. It is redundant job to run algorithm to get the probability function. Moreover, we need to calculate all the vertices’ score to select "best" vertices. However, most graph data shows similar tendency with Figure 8. It shows that low-degree vertices generally higher scores:

i.e. reasonable vertices to be selected. Moreover, it doesn’t have to run the entire algorithm, but only sorting, the pre-processing time is lower. Therefore, selecting the low-degree vertices as preloading vertices is an efficient way to optimize I/O performance. In the next section, we will show graph ordering scheme that can optimize low-degree preloading scheme.

13

(24)

Figure 8: Scoring result for LiveJournal graph with BFS algorithm. Score for vertices of the same out-degree are averaged.

4.3 Low-Degree Graph Ordering

In this section, we propose the simple but efficient graph ordering that is suitable for our low-degree preloading mechanism. For preloading mechanism, there is a problem of index look- up overhead of preloaded vertices. For original vertex IDs’, low degree vertices may scatter on the graph. To store them in the preloading area, we need to make an indexing map. When processing large-scale graphs, the number of preloaded vertices is also large, which makes the index look-up time longer.

However, we can reduce the index look-up time when we use low-degree selection scheme for preloading mechanism with simple ordering:low-degree graph ordering. This can eliminates the index look-up overhead because we can compute the index with the number of previous vertices.

Moreover, because we don’t have to store the indexing map, the memory cost will be reduced.

Equation 3 shows how we can compute a vertex’s index.

vidx= (vid−start(deg(vid)))∗deg(vid) + X

d<deg(vid)

(end(d)−start(d))∗d (3)

• vidx is the index of vertex of loaded list

• vid is the vertex ID

• deg(v) represents the degree of vertexv

• start(d)/end(d) represents the start/end vertex ID among the vertices with degreed

14

(25)

However, to use the low-degree ordering for our optimization, we need to consider the possi- bility of performance degradance due to the low-degree ordering. However, as discussed in section 3.4 and shown in Figure 6 low-degree ordering shows performance improvement compared to other graph orderings and reasonable performance compared to Gorder which is the state of the art ordering. Low-degree ordering itself shows poor performance compared to Gorder but it can be used with preloading mechanism, it can show higher performance improvement.

15

(26)

V Evaluation

We evaluate the performance of our optimization with six algorithms: breadth first search (BFS), diameter of graph (DIAM), betweenness centrality (BC), shortest paths (SP), all-pair shortest paths (APSP) and finding weakly connected components (WCC). All the algorithms are BFS-like algorithms which are fit to our purposed model. BFS and SP are implemented as de- scribed in Pregel [3]. For APSP, we sampled 128 source vertices and compute the distances from those sources using SP. APSP runs in multiple steps and in each step it computes the distances from eight source vertices. BC is implemented using the algorithm proposed by Bran-des [4]. It runs SP from each source vertex and counts the number of paths passed for each vertex. This is repeated for all the (source) vertices. As computation is intense, an approximate approach is taken by computing the centrality scores with 128 randomly sampled source vertices. WCC is implemented in the typical manner of propagating component IDs for each vertex and then computing the minimum IDs.

Table 1: Hardware and Software speculations

Hardware

Intel Xeon E5-2683 v4 Samsung 128GB DRAM Intel 400GB SSD with SATA 6.0Gb/s Software Ubuntu 16.04 LTS

FlashGraph v0.3.2

We performed our experiments on a machine with Intel Xeon E5-2683 v4. CPU has 16-cores that serves hyper-threading. In this paper, we used single threads for each core. Machine has 128GB DRAM and Intel 400GB SSD with SATA 6.0GB/s interface. We set the page size as 8KB which is generally used. For all experiments, we used only a single disk. The machine runs Ubuntu 16.04 as an operating system. We implement and evaluate our model on FlashGraph [1], a semi-external graph engine optimized for SSDs. We choose FlashGraph because 1) it is a rep- resentative semi-external graph engine, 2) it is recently developed and thus, most known I/O optimizations are provided, and 3) it is actively maintained and core graph algorithms are al- ready implemented in the system. The experimental environments are summarized in Table 1.

The evaluation is performed with four real-world networks: Flickr, LiveJournal, Twitter and Friendster. Flickr and LiveJournal are relatively small data which are used to show the scala- bility of our model. All the dataset is downloaded from KONECT [5]. The informations of each data are summarized in Table 2.

16

(27)

Table 2: Datasets [5]

Graph |V| |E| Size (GB)

Flickr 2,302,925 33,140,017 0.42

LiveJournal 4,847,571 68,475,391 1.02

Twitter 52,579,682 1,963,263,821 16.1

Friendster 68,349,466 2,586,147,869 21.2

5.1 Execution Time pre-processing

First, we evaluated pre-processing time. The pre-processing means the time cost of selecting vertices. We compare two approaches: using greedy algorithm for selection and low-degree ver- tex selection. Table 3 shows the result of pre-processing execution time among three datasets:

LiveJournal, Twitter and Friendster.

Table 3: Pre-processing execution time

Graph Greedy Low-degree

LiveJournal 3.8 min 3.3 sec

Twitter 2.9 hour 30.1 sec

Friendster 3.5 hour 43.1 sec

For greedy selection scheme, the cost of selecting vertices is extremely high: maximally 3.5 hour for Friendster dataset which is the biggest dataset. However, single graph analytic algo- rithms don’t run that long, i.e. the execution time for betweenness centrality algorithm with Friendster dataset which shows the longest execution time is under 800 seconds and for BFS algorithm, it takes under 60 seconds. The algorithm can run multiple times to get the similar execution time with the pre-processing time, betweenness centrality algorithm should run 22∼ 23 times and for BFS algorithm, it should run over 300 times. In other words, greedy selection scheme can be used in such situation: graph dataset doesn’t change for a long time or performs large times of graph analytic algorithms.

In contrast, low-degree vertex selecting scheme shows quite reasonable pre-processing time.

Compared to the greedy selection, the pre-processing time is up to 417×faster with Friendster dataset and minimally69×faster with LiveJournal dataset. This is because low-degree selection only needs to sort the vertices according to the degree and choose front-most vertices. The sorting can be done in O(nlogn) by using quicksort algorithm. The number of vertices can be simply calculated because the degree itself shows the size of each vertex’s edge list. Therefore,

17

(28)

we can conclude that selecting low-degree vertices can be used in such situation: graph dataset changes frequently or performs small times of graph analytic algorithms.

Overall execution time

Now, we will discuss overall performance improvement using preloading mechanism. We scale up the memory space used by page cache and preloading from 10% to 40% of input graph size.

We compare 5 schemes: greedy selection (Greedy), low-degree vertices selection (LowD), high- degree vertices selection (HighD), random vertices selection (Rand) and using only page cache.

The result is shown in Figure 9.

(a) BFS (b) DIAM (c) BC

(d) SP (e) APSP (f) WCC

Figure 9: Normalized execution time of various selection schemes. Six BFS-like algorithms and four real-world graphs are used for experiments. The execution time is normalized to the page cache only

The result is normalized to the execution time that uses page cache. As the result shows that the using greedy selection scheme for preloading shows the best performance among all: up to 76% and 29% on average compared to the page cache only. Low-degree vertex selection shows lower performance improvement compared to greedy selection but improved up to up to 75%

and 19% on average compared to the page cache only.

18

(29)

5.2 Analysis of Page Utilization

In this section, we will show how much the preloading mechanism improved page utilization compared to the page cache. The utilization is measured in the following way. For the entire execution of the algorithms, for each page that is retrieved to the page cache, individual 64 byte granularity units of the page are monitored throughout while the page resides in the page cache. Utilization of the page, calculated when the page is evicted, is the fraction of the total accessed units within the page size. Page utilization that is reported is the average of all the pages evicted as well as those still residing in the page cache at the end of the algorithm execution.

Figure 10: Evaluation of page utilization with various selections for preloading

Figure 10 shows the result of page utilization among various selecting schemes. As shown in the figure, greedy selection and low-degree selection utilize page in higher rate compared to the random selection. Also, compared to the page cache, which was shown in Figure 3, preloading mechanism can utilize page much better. Utilization has increased up to around 60% which is 20% to 30% higher than page cache itself.

19

(30)

5.3 Ordering Effectiveness

For the last experiment, we will show the impact of low-degree ordering applied to preloading mechanism with low-degree selection scheme. Only BFS algorithm is used for this experiment be- cause as shown in the previous experiments, all otherBFS-like algorithmshows similar tendency to BFS algorithm. Therefore, we use only BFS algorithm as representative algorithm. To see the impact of low-degree ordering, we compare only two schemes: using only preloading mechanism with low-degree selection (LS) and applying low-degree ordering to optimize (LSLO). We run the algorithm 20 times and average the execution time. The execution time is normalized to the result of using preloading mechanism only (LS).

Figure 11: Evaluation of ordering impact. LS is low-degree selection only and LSLO is low-degree selection combined with low-degree ordering.

As demonstrated in Figure 11, applying ordering is an obvious winner compared to the preloading-only mechanism. Compared to the LS, LSLO shows maximally 57% faster execution time and 29% faster on average. Moreover, original preloading mechanism uses index map that hasN entries if there areN vertice’s edges are preloaded. To store the map information, it needs 8N size of memory. Generally, low-degree vertices are cached, the size ofN is large which means memory cost is also large. However, if we use low-degree ordering mechanism, we need tuple information consist of(degree, start offset, number of vertices). Using the tuple information and if we let the number of tuple as Nt which is the same as the number of different degrees, the memory consumption will be12Ntwhich is much smaller than8N. (Nt<< N) Therefore, using the low-degree ordering can optimize performance not only execution time, but also memory consumption.

20

(31)

VI Related Works

Disk-based Graph Processing System

GraphChi is the first disk-based graph processing system that made large-scale graph analysis possible on a single machine [6]. GraphChi is mainly designed for HDDs and eliminates random disk accesses with its Parallel Sliding Windows technique. TurboGraph is a disk-based graph engine for SSDs [7]. In TurboGraph, adjacency lists of requested vertices are randomly accessed in each iteration. Its pin-and-slide execution model overlaps random I/O with CPU computa- tion and fully utilizes the I/O parallelism inherent within SSDs. Its successor, TurboGraph++, further extends the system for a distributed cluster [8]. As TurboGraph and GraphChi make use of the page cache for its random I/O, our optimizations can be easily applied to their systems.

While the vertex-centric computation model is widely adopted in large-scale graph process- ing, recent studies on disk-based graph systems propose an alternativeedge-centriccomputation that streams edges into memory to exploit the locality of edge access [9, 10]. The edge-centric model shows good performance for algorithms accessing the entire graph repeatedly. However, the model is not as efficient for algorithms that iteratively access different subsets of a graph as is done with BFS-like algorithms. Recently, Maass et al. proposed a hybrid model in Mosaic that supports both the vertex-centric and edge-centric models at the same time to efficiently utilize CPUs and co-processors [11].

In semi-external graph engines, vertex attributes are stored in main memory for fast up- dates [1,12–14]. Pearce et al. proposed asynchronous optimization techniques for graph traversal algorithms for semi-external graph processing [13]. FlashGraph implements several I/O opti- mizations for SSDs and SSD arrays such as merging I/O requests for higher throughput and overlapping I/O and computation [1]. We build on top of these optimizations and propose meth- ods that exploit the structural properties of the input graph.

Several other I/O optimization methods for disk-based graph processing have recently been proposed. Vora et al. employs a dynamic partitioning scheme that prevents loading unnecessary edges on disk [15]. GridGraph supports 2D edge partitioning to reduce I/O access [10]. In Graphene, a bitmap based asynchronous I/O optimization is applied to efficiently merge small I/O requests [16]. Our proposed optimization techniques are applicable on top of these I/O optimizations as we have shown with Graphene.

Main Memory Graph Processing

For large-scale graph processing, vertex-centric computation model was initially proposed in the context of distributed in-memory systems. Google’s Pregel is the first system adopting the

21

(32)

vertex-centric computation model [3]. GraphLab is a distributed machine learning framework that employs a variant of the vertex-centric model that supports asynchronous computation [17].

Its successor, PowerGraph, adopts an efficient graph partitioning scheme that considers the power-law degree distributions of real-world graphs [18].

To ease the programming difficulty of large-scale graph analysis, SociaLite supports declar- ative query language based on Datalog [19]. Wang et al. also demonstrate the performance and scalability of Datalog-based graph processing [20].

Galois is a parallel graph processing system based on an implicitly parallel vertex itera- tor [21]. Green-Marl is a domain-specific language for writing parallel graph algorithms for shared-memory [22]. Ligra is a light-weight framework with low-level primitives to implement graph processing systems [23].

Wei et al. studied optimizing the performance of graph algorithms for main memory graph processing [2]. The authors develop a graph ordering calledGorderthat optimizes the locality of updating vertex attributes. While Gorder is designed for main memory graph systems, Norder is designed to optimize the performance of disk-based graph engines.

22

(33)

VII Conclusion

In this paper, we analyzed graph algorithms on disk-based graph engines. The results tell us that we can categorize graph algorithms into three groups based on I/O performance and request patterns: BFS-like algorithm,PageRank-like algorithm and Subgraph mining algorithm.

We figured out that for BFS-like algorithm, the indifference of the number of I/O requests and the original page cache hit ratio is relatively low. Based on the results, we concluded that the original page cache is not suitable solution to optimize I/O performance for graph processing.

In this paper, we proposeSelective Preloading Mechanism, to optimize the I/O performance for graph processing. We reduced unnecessary memory consumption and allocation and evic- tion overhead using vertex-granular preloading mechanism. We selected two types of vertices:

Greedy and Low-degree. Both can optimize I/O performance compared to page cache and can be used in proper situations. We also proposed low-degree graph ordering that can be utilized with preloaded mechanism using low-degree selection scheme which reduces index look-up time by making possible to calculate the indexes.

The optimization has evaluated four real-world graphs. Greedy selection shows performance improvement up to 76% faster maximally and 28% on average compared to page cache. Low- degree selection shows performance improvement up to 75% faster maximally and 19% on av- erage compared to page cache. By applying the low-degree ordering for low-degree selection, the performance is improved 57% faster maximally and memory consumption also decreased efficiently. Moreover, we figured out that the page utilization has increased, which indicates the I/O optimization is the key factor of disk-based graph engine.

23

(34)

References

[1] D. Zheng, D. Mhembere, R. Burns, J. Vogelstein, C. E. Priebe, and A. S. Szalay, “Flash- Graph: Processing billion-node graphs on an array of commodity SSDs,” in Proceedings of the USENIX Conference on File and Storage Technologies (FAST 15), 2015, pp. 45–58.

[2] H. Wei, J. X. Yu, C. Lu, and X. Lin, “Speedup graph processing by graph ordering,” in Proceedings of the International Conference on Management of Data (SIGMOD 16), 2016, pp. 1813–1828.

[3] G. Malewicz, M. H. Austern, A. J. Bik, J. C. Dehnert, I. Horn, N. Leiser, and G. Czajkowski,

“Pregel: a system for large-scale graph processing,” in Proceedings of the ACM SIGMOD International Conference on Management of Data (SIGMOD 10), 2010, pp. 135–146.

[4] U. Brandes, “A faster algorithm for Betweenness Centrality,” The Journal of Mathematical Sociology, vol. 25, no. 2, pp. 163–177, 2001.

[5] “The koblenz network collection.” http://konect.uni-koblenz.de/.

[6] A. Kyrola, G. Blelloch, and C. Guestrin, “GraphChi: Large-scale graph computation on just a PC,” in Proceedings of the USENIX Symposium on Operating Systems Design and Implementation (OSDI 12), 2012, pp. 31–46.

[7] W.-S. Han, S. Lee, K. Park, J.-H. Lee, M.-S. Kim, J. Kim, and H. Yu, “TurboGraph: a fast parallel graph engine handling billion-scale graphs in a single PC,” in Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD 13), 2013, pp. 77–85.

[8] S. Ko and W.-S. Han, “TurboGraph++: A scalable and fast graph analytics system,” in Proceedings of the 2018 International Conference on Management of Data (SIGMOD 18), 2018, pp. 395–410.

[9] A. Roy, I. Mihailovic, and W. Zwaenepoel, “X-stream: Edge-centric graph processing using streaming partitions,” in Proceedings of the Symposium on Operating Systems Principles (SOSP 13), 2013, pp. 472–488.

[10] X. Zhu, W. Han, and W. Chen, “GridGraph: Large-scale graph processing on a single ma- chine using 2-level hierarchical partitioning.” in Proceedings of the USENIX Annual Tech- nical Conference (ATC 15), 2015, pp. 375–386.

24

(35)

[11] S. Maass, C. Min, S. Kashyap, W. Kang, M. Kumar, and T. Kim, “Mosaic: Processing a trillion-edge graph on a single machine,” inProceedings of the Twelfth European Conference on Computer Systems (EuroSys 17), 2017, pp. 527–543.

[12] P. Kumar and H. H. Huang, “G-store: High-performance graph store for trillion-edge pro- cessing,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC 16), 2016, pp. 830–841.

[13] R. Pearce, M. Gokhale, and N. M. Amato, “Multithreaded asynchronous graph traversal for in-memory and semi-external memory,” inProceedings of the ACM/IEEE International Conference for High Performance Computing, Networking, Storage and Analysis (SC 10), 2010, pp. 1–11.

[14] Z. Shao, J. He, H. Lv, and H. Jin, “Fog: A fast out-of-core graph processing framework,”

International Journal of Parallel Programming, pp. 1–14, 2016.

[15] K. Vora, G. H. Xu, and R. Gupta, “Load the edges you need: A generic I/O optimization for disk-based graph processing.” in Proceedings of the USENIX Annual Technical Conference (ATC 16), 2016, pp. 507–522.

[16] H. Liu and H. H. Huang, “Graphene: Fine-grained IO management for graph computing,”

in Proceedings of the USENIX Conference on File and Storage Technologies (FAST 17), 2017, pp. 285–300.

[17] Y. Low, D. Bickson, J. Gonzalez, C. Guestrin, A. Kyrola, and J. M. Hellerstein, “Distributed GraphLab: a framework for machine learning and data mining in the cloud,” Proceedings of the VLDB Endowment, vol. 5, no. 8, pp. 716–727, 2012.

[18] J. E. Gonzalez, Y. Low, H. Gu, D. Bickson, and C. Guestrin, “PowerGraph: Distributed graph-parallel computation on natural graphs.” in Proceedings of the Symposium on Oper- ating Systems Design and Implementation (OSDI 12), 2012, pp. 17–30.

[19] J. Seo, J. Park, J. Shin, and M. S. Lam, “Distributed SociaLite: A datalog-based language for large-scale graph analysis,” Proceedings of the VLDB Endowment, vol. 6, no. 14, pp.

1906–1917, 2013.

[20] J. Wang, M. Balazinska, and D. Halperin, “Asynchronous and fault-tolerant recursive dat- alog evaluation in shared-nothing engines,” Proceedings of the VLDB Endowment, vol. 8, no. 12, pp. 1542–1553, 2015.

[21] D. Nguyen, A. Lenharth, and K. Pingali, “A lightweight infrastructure for graph analytics,”

in Proceedings of the ACM Symposium on Operating Systems Principles (SOSP 13), 2013, pp. 456–471.

25

(36)

[22] S. Hong, H. Chafi, E. Sedlar, and K. Olukotun, “Green-Marl: a DSL for easy and efficient graph analysis,” ASPLOS XVII Proceedings of the seventeenth international conference on Architectural Support for Programming Languages and Operating Systems, pp. 349–362, 2012.

[23] J. Shun and G. E. Blelloch, “Ligra: a lightweight graph processing framework for shared memory,” inProceedings of the 18th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming (PPoPP 13), 2013, pp. 135–146.

26

(37)

Gambar

Figure 1: Architecture of disk-based graph processing system (Semi-external-memory system)
Figure 2: Distribution of the number of edge list accesses: aggregated vertices with the same in-degree
Figure 3: Page cache hit ratio and normalized execution time of various graph algorithms while scaling up the page cache size from 5% to 30% of the input graph size
Figure 6: Normalized execution time of various graph algorithms according to the four different graph orderings
+7

Referensi

Dokumen terkait

iii P R O G R E SS Jurnal Pendidikan Agama Islam Daftar Isi Salam redaksi : ...i Daftar Isi : ...ii PROBLEMATIKAPEMBELAJARAN PENDIDIKAN AGAMA ISLAM DI LEMBAGA PENDIDIKAN