Graph algorithms are an effective way to detect malicious activity within your company’s network. Pretty much in the same way, they serve better recommendations to your customers, solve supply chain problems in manufacturing and distribution, and detect fraudulent transactions or insurance claims.
Olga Razvenskaia and I will explain why and demonstrate how Neo4j Graph Analytics for Snowflake can use two graph-based algorithms, K-Nearest Neighbours (KNN) and GraphSAGE to detect network intrusion. All this can be done while your data remains securely in the Snowflake enclave .
Before we get going, I want to give a huge shout-out to the authors of the academic paper who proved the effectiveness of this approach in “Optimizing IoT Intrusion Detection — A Graph Neural Network Approach with Attribute-Based Graph Construction” by Tien Ngo, Jiao Yin, Yong-Feng Ge, Hua Wang, published in the Special Issue Data Privacy Protection in the Internet of Things .
It was their thorough research at the Institute for Sustainable Industries and Liveable Cities, Victoria University, Australia, that inspired this blog.
Gratitude also goes to Mohanad Sarhan, Siamak Layeghy, and Marius Portmann, the curators of the underlying data set , who kindly granted us permission to use it in our blog. Good public data sets are hard to find, so we really appreciate the great work at the University of Queensland, Australia.
Tien Ngo, Jiao Yin, Yong-Feng Ge, and Hua Wang's own words explain the challenges with traditional models best and how their approach solves the business problem.
Highlights from the paper’s abstract
“complexity and heterogeneity of the Internet of Things (IoT) ecosystem present significant challenges for developing effective intrusion detection systems. …
existing approaches primarily construct graphs based on physical network connections which may not effectively capture node representations.... (Ed. this is also why relational databases arent suitable.)
Instead of relying on physical links, the TKSGF constructs graphs based on Top-K attribute similarity ( Ed. find the most similar) , ensuring a more meaningful representation of node relationships…
maintaining scalability. Furthermore, we conducted extensive experiments to analyze the impact of graph directionality (directed vs. undirected), different K values, and various GNN architectures and configurations on detection performance….
benchmark demonstrated that our proposed framework consistently outperformed traditional machine learning methods “
K-Nearest Neighbour (KNN) is a simple, fast, supervised machine learning method that can be used to classify data . It works by identifying labelled data points (in our case, intrusions) that are closest to a new, unlabeled data point, which could be a benign or malicious event, and creating relationships to provide classification.
You could use KNN to classify inventory, insurance claims or customers who interact with your services. For intrusion detection, we will create a series of relationships among the most similar intrusion events. Because KNN is a “lazy learner”, it computes the predictions as it goes and does not store the resulting model.
GraphSAGE (Graph Sample and Aggregated) is different to most graph algorithms. It’s an inductive learning* model that generates embeddings by sampling and aggregating features from the surrounding structure of the graph. This approach captures the chain of events that lead to the system’s compromise. The data set is labelled, which means we can run supervised machine learning because we know whether the event is one of the attacks (see below). Unlike KNN, GraphSAGE does store the resulting model, which can then be used for prediction time and time again.
*Inductive learning derives general rules, patterns, or principles from specific observations, data points, or examples.
Note: Neo4j’s GraphSAGE implementation is also capable of unsupervised learning, where you don’t have labelled data and semi-supervised machine learning, where you have a mix of labelled and unlabelled data.
For the model, we need to encode the data in the IOT table as an 8-dimensional vector, with a float per feature. This enables us to use cosine to measure the distance (or difference) between the data points.
Download the IoT-BoT-data from here , upload it to a stage in Snowflake and copy it into a database table, e.g. NETFLOWDATASET.DATA.NF_BOT_IOT . The row_id is a unique identifier of a “connection event” as a node, and we use rand to split the data for train and test data.
, create the unlabelled data set
Neo4j Graph Analytics for Snowflake is available in the Snowflake Marketplace as a native application that runs on Snowpark Container Services (SPCS). It installs a series of compute pools, enabling you to use the size that best suits the data set. An estimation procedure can be used to pick the right compute pool size. The application provides a set of procedures to run an algorithm over data from your Snowflake tables and write the results to new tables.
Once you have installed, granted and activated the application, click launch and run the grants below so it can read and write to the database tables.
Note: The following assumes you kept the application's default name during installation.
We treat every row in the original dataset as a connection event. We use KNN to determine similarity between pairs of events stored as vectors.
Instead of relying on physical links, the TKSGF constructs graphs based on Top-K attribute similarity, ensuring a more meaningful representation of node relationships
Two nodes are considered connected if they have a cosine similarity of more than 0.8 (‘similarityCutoff’: 0.8) ; every node is limited to no more than 9 connections (‘topK’: 9) . The default sampleRate value of 0.5 is used, which means that only 50% of all possible pairs are checked for every node. This improves the runtime of the algorithm without impacting the quality of the results.
Neo4j Graph Analytics for Snowflake enables us to run these complex algorithms directly where the data lives — inside Snowflake. When KNN is run, the data will be projected into memory within the Large X64 SPCS instance, and the results will be written to the REL_IS_SIMILAR table .
KNN run took 0m 45s to compute the similarities with 600100 nodes, compute the results and write the data back to the output table REL_IS_SIMILAR for use with GraphSAGE.
, we need to create the connections (i.e., actual relationships) between the IoT devices, add a label column to the data, and convert the VEC column from an array type to a VECTOR(FLOAT,8) type.
Because GraphSAGE is computationally intensive (essentially deep learning on graphs), we will use Snowflake’s GPU_NV_S compute pool. You can use one of the CPU instances provided if you don’t have access to GPUs, but it will take longer.
GraphSAGE uses the node table IOT_CONNECTIONS with labels set to NULL for the test events (the ones we want to predict), and the REL_IS_SIMILAR table created by KNN as the relationship table. The REL_IS_SIMILAR connects up to 9 events num using Samples = [9,9] that meet the 0.8 cutoff. GraphSAGE builds a persistent model in SPCS storage called gs_tksgf_iot, using supervised machine learning and node similarity. Persisting the model to storage lets you use it in your pipelines to make predictions whenever you want.
Our graph is made up of attacks, and the relationships are formed from the similarity to other attacks.
Note on the configuration:
splitRatios’: {‘TRAIN’: 0.70, ‘TEST’: 0.15, ‘VALID’: 0.15} .
This configuration tells our native app how to handle the data and split it into 70% training, 15% test, and 15% validation sets for measuring the training progress — this is not connected to the train/test split we introduced above.
It took 5 Minutes to train the model, which is stored within SPCS with the name gs_tksgf_iot . Graph Analytics for Snowflake won't overwrite the model if it exists. Please refer to the ‘ Model Catalog Operations ’ for procedures for listing and dropping models.
We can now use the trained model to predict whether an event in the test set is malicious or benign. When you start pulling operational event data from your network, you can predict whether it is under attack.
It took 1m43s to make the predictions.
To calculate the prediction results and compare them with the findings in the paper.
Macro-F1 score reported in the results table below vs the results in the paper are 0.811 vs. 0.821, and the weighted F1-score reported is 0.985 vs. 0.985 from the paper. This means we have reproduced the results in the Academic paper, validating their research and our implementation on Snowflake.
The application creates GraphSAGE’s models in SPCS storage, which can be managed with the following commands. Documentation available here .
In this blog, Olga Razvenskaia and Stu Moore demonstrate that Neo4j Graph Analytics for Snowflake is highly effective at detecting network intrusions.
Neo4j Graph Analytics for Snowflake is available in the Snowflake Marketplace and i ncludes a 30-day free trial .
Plus, there are lots of other great examples available on GitHub and Snowflake’s Developer Guides and check out our new Product Page on Neo4j.com for more resources.
The full story
This article is one source in a clustered incident — the cluster page carries the summary, timeline and every other outlet covering it.
