How to deploy GraphDB in GCP¶
What’s in this document?
GraphDB can be deployed on Google Cloud Platform (GCP) by following the general installation instructions. You can find information regarding the costs of running a GraphDB instance on the Google Cloud Platform website.
This documentation will walk you through the process of setting up the necessary environment for deploying GraphDB on Google Cloud Platform.
Note
Ontotext maintains a Terraform module that automates the entire procedure of deploying GraphDB on Google Cloud Platform. Learn more about how to use it at our GitHub repository.
Architecture¶
The GraphDB architecture diagram showcases the deployment architecture for GraphDB on Google Compute Engine instance in Google Cloud Platform. The diagram illustrates the key components, and their interactions to provide a high-level understanding of the system’s architecture and how it should be deployed. There are no third-party integration points on the default GraphDB deployment.
Prerequisites¶
There are several prerequisites for running a GraphDB instance on GCP:
Active GraphDB license required to use the Enterprise functionalities of the database
Access to an GCP account
Note
We recommend the use of an Identity and Access Management user for the deployment instead of a root user account.
Technical requirements¶
The following Google Cloud Platform services are required to complete the GraphDB deployment on GCP:
Service |
Description |
|---|---|
Google Compute Engine (GCE) |
Server instance that are used to host the database application |
Cloud Balancing |
Used for load balancing the GraphDB cluster nodes |
Google Cloud Persistent Disk |
Persistent Disk volumes are used for storing the data |
Google Cloud Identity and Access Management (IAM) |
Provides user and access management for your GraphDB deployment |
Google Cloud Secrets Manager |
Used to store various GraphDB configurations |
Google Cloud Storage |
Used to store GraphDB Backups |
Cloud Monitoring |
Used to monitor the status of the GraphDB cluster |
Required skills¶
The following skills and knowledge are typically required in order to successfully deploy GraphDB on GCE Instance:
Google Cloud Platform Fundamentals |
Familiarity with Google Cloud Platform (GCP) and understanding of its core concepts, such as Google Cloud Engine instances, security, VPCs and IAM roles. Knowledge of how to navigate the Google Cloud Platform Web Console and interact with GCP services is essential. |
|---|---|
Google Cloud Engine Instance Management |
Proficiency in creating and managing Google Cloud Engine instances. This includes selecting the appropriate instance type, configuring security settings, managing storage (Cloud Persistent Disk Volumes), and understanding Google Cloud Engine instance lifecycle management. |
Networking and Security |
Understanding of networking concepts in GCP, including VPC (Virtual Private Cloud) configuration, subnets, routing tables, and firewall rules. Knowledge of how to set up inbound and outbound traffic rules to allow communication with GraphDB. |
Linux Administration |
Proficiency in Linux command-line interface (CLI) and basic administration tasks. This includes SSH access to GCE instances, navigating the file system, managing permissions, installing packages, and configuring system settings. |
Database Management |
Knowledge of GraphDB and its deployment requirements. Understanding of how to configure GraphDB settings, including database storage, memory allocation, and repository creation. |
Database Backup and Recovery |
Familiarity with backup and recovery strategies for GraphDB on GCP. Knowledge of GCP services like Google Cloud Storage for data backups and restoration processes. |
Monitoring and Troubleshooting |
Proficiency in monitoring the health and performance of GraphDB instances on GCP. Understanding of logging, monitoring and troubleshooting techniques using GCP Cloud Monitoring, GCE instance logs, and GraphDB diagnostic tools. |
High Availability and Scalability |
Knowledge of implementing high availability and scalability for GraphDB on GCP. This may involve using features like GCP Auto Scaling, load balancers, and multi-zone deployments. |
Infrastructure as Code (IaC) |
Familiarity with Infrastructure as Code principles and tools like Terraform. This enables automating the provisioning and configuration of GraphDB infrastructure on GCP. |
Security Best Practices |
Understanding of security best practices for GCP deployments, including data encryption, access controls, identity and access management, and compliance considerations. |
Setting up your Virtual Private Cloud (VPC)¶
Go to the and click on .
Under , select the option. This will allow you to configure the subnets in the regions.
Provide a descriptive by which to recognize your VPC.
Use the table below as a reference to define the to allow inbound connections to the GraphDB instance.
Name 1
Type
Targets
Filters
Protocols / ports
Action
Priority 2
rule-allow-custom
Ingress
Aply to all
IP ranges: 10.128.0.0/9
all
Allow
65,000
rule-allow-icmp
Ingress
Aply to all
IP ranges: 0.0.0.0/0
icmp
Allow
65,000
rule-allow-rdp
Ingress
Aply to all
IP ranges: 0.0.0.0/0
Allow
65,000
rule-allow-ssh
Ingress
Aply to all
IP ranges: 0.0.0.0/0
Allow
65,000
rule-deny-all-ingress
Ingress
Aply to all
IP ranges: 0.0.0.0/0
all
Deny
66,000
rule-allow-all-egress
Egress
Aply to all
IP ranges: 0.0.0.0/0
all
Allow
66,000
- 1
The names in the table are illustrative and describe what the rule does. Use names that make sense for your deployment.
- 2
The priorities given in the table are illustrative. What is important is that the priority of rules
rule-deny-all-ingresandrule-allow-all-egressis higher than of the rest of the rules.
Click on
Tip
It may take a while for the VPC to deploy. You will know that it is done deploying when you the message “Successfully created network” appears on the bottom of your screen.
Setting up your Cloud DNS private hosted zone¶
The GraphDB Raft implementation requires static addresses. This is achieved by creating a private hosted zone in Google Cloud DNS and registering the instances there.
From the Cloud DNS dashboard, click on .
Under Zone type, select .
Provide a Zone name.
Provide a DNS name such as
graphdb.cluster.Under the section select the networks from which you want to have access to the hosted zone.
Click
Note
Later, you will also need to create “A” records for the instances.
Creating a Cloud Storage Bucket¶
Tip
This step is optional, but recommended.
GraphDB can store backups to Cloud Storage and, if needed, restore from them.
To create an Cloud Storage bucket:
From the Cloud Storage console, click on .
Provide a globally unique name.
Select your region, scroll down and click .
In order to block all public access when you create a bucket you should tick Enforce public access prevention on this bucket on the .
Importing a certificate into GCP Certificate Manager¶
Tip
This step is optional.
While serving GraphDB requests over a secured and encrypted connection is not strictly required, it is nevertheless highly recommended. This section goes over the process of importing a certificate into GCP Certificate Manager so that you can use that certificate later when creating the load balancer.
Go to the console and click on .
Paste or upload your Certificate.
Paste or upload your Private Key.
Warning
You need to remove the password for the key before pasting it.
You can optionally add labels, then click on
Setting up the Load Balancer¶
Go to Load Balancing service and click on .
Click on and provide the following configurations:
Type of load balancer: Network Load Balancer
Proxy or passthrough: Proxy load balancer
Public facing or internal: Public facing (external)
Tip
Otherwise GraphDB will not be accessible externally.
Global or single region deployment: Best for global workloads
Load balancer generation option: Global external proxy Network Load Balancer
Then click on the button.
On the next screen, provide a unique Load Balancer name.
Select your from the Region dropdown field.
Choose the in which you want to deploy the load balancer.
Under Backend configuration:
Backend type: Select .
Protocol: Select .
Under Instance Groups, configure the following parameters:
Instance group: Select the instance group that you used to deploy the GraphDB nodes.
Port numbers: 7200
Under Healthcheck configuration, configure the following parameters:
Protocol: HTTP
Port: 7201
Request path: /rest/cluster/node/status
Under Frontend configuration, configure the following parameters:
Name: Provide a unique name.
Protocol: SSL
Port: 443
Certificate: Choose the certificate that you have created in the previous section.
Click .
Setting up service account¶
The GCE instances require certain permissions in order to perform several tasks in GCP. This section describes the permissions they need, what they are used for, and how to create them.
To achieve all of this, create an instance profile and then attach it to the instances:
Go to the Identity and Access Management (IAM) dashboard and select from the navigation menu on the left.
On the service account details provide the following configurations:
Define name for the Service account name.
Define Service account description.
On the Grant this service account access to project (optional) step you can specify the role for the service that you want to access from the GCE instances.
On the Grant users access to this service account (optional) step you can specify Service account user roles or Service account admin roles that can access the service account.
Click on .
Note
If you are planning to store GraphDB backups to Cloud Storage, the GCE instance needs to be able to read, write, and list objects in Cloud Storage service. Perform the steps from Setting up service account, but under Grant this service account access to project, select .
Launching GraphDB instances¶
Before you can launch your GraphDB instance, you will need to create an instance template with an existing GraphDB image. You can do it by using the following command using the gcloud command line interface:
gcloud beta compute instance-templates create graphdb-instance-template \
--project=example-project-name --machine-type=n2-highmem-8 \
--network-interface=network=default,network-tier=PREMIUM,stack-type=IPV4_ONLY \
--instance-template-region=us-central1 --maintenance-policy=MIGRATE \
--provisioning-model=STANDARD \
--scopes=https://www.googleapis.com/auth/cloud-platform \
--create-disk=auto-delete=yes,boot=yes,device-name=graphdb-instance-template,image=projects/graphdb-public/global/images/ontotext-graphdb-10-7-6-202410151959,mode=rw,size=20,type=pd-balanced \
--create-disk=device-name=persistent-disk-1,mode=rw,size=500,type=pd-balanced \
--no-shielded-secure-boot --shielded-vtpm --shielded-integrity-monitoring \
--reservation-affinity=any
Setting up Instance Group¶
In order to create an instance group based on the instance template that you have created in the previous section, execute the following steps in the GCP Web Console.
Go to Compute Engine Service.
On the left tab select option.
Click on button and fill the fields below.
Name: Define a name that you want to use for the instance template.
Description (Optional): Provide a description to help easily identify the deployed resources.
Instance template: Select the instance template that you have created via the gcloud CLI, as described in Launching GraphDB instances.
Location: Choose and check if the three zones in the region you selected are selected (they should be selected by default).
Autoscaling: Set both Minimum number of instances and Maximum number of instances to 3.
Port Mapping: 7200
Click on the button
Creating a cluster¶
In order to create the cluster, you will need to get the address of the GCE instances.
Go to the dashboard and open your private hosted zone.
Write down the names of all records with type “A”.
Once you’ve written down all records names, you can create a cluster by following the Creating and managing a cluster documentation from one of the instances.
Tip
The recommended way to gain access to the instances is by going to Compute Engine and selecting the instance you want to access, then clicking on .
Opening GraphDB instances¶
Once all instances are running, you can launch your GraphDB instance:
Access the .
Copy the DNS name.
Paste it in the address bar of your browser and press Enter.
Updating GraphDB configurations and versions¶
Since GraphDB and its configurations are baked into the Instance Image, you will need to recreate the GCE instances when updating your GraphDB configuration to a newer minor version. You can do this by either manually stopping each individual instance, or scaling the cluster out, and then scaling it back in. This section describes both methods in detail.
Stopping individual GCE instances¶
The faster way to update to a newer minor version and its configuration is to stop each individual GCE instance. The downside to this method is that you are decreasing the cluster HA. In other words, if a node fails while recreating an instance, the cluster will be unable to process writes. The process is simple:
Update the Instance Image in the instance template.
Terminate the instances one by one, starting with the follower nodes, and leaving the leader node to be the last instance to be terminated.
See also
To avoid compatibility issues, also refer to the Migrating GraphDB configurations documentation.
Warning
When you terminate an instance, wait for the new one to be started. Then verify that it has successfully rejoined the cluster and that it is in sync before proceeding with the next one.
Scaling the cluster out and then back in¶
You can also recreate the GCE instances by scaling the cluster out and in. The advantage of this approach is that the HA will not be impacted. However, the cluster will need to replicate its state to the new nodes. This can take a significant amount of time, especially with bigger-sized repositories.
Update the Instance Image in the launch instance template.
Double the size of the cluster.
Note
Change the minimum, maximum and desired size of the auto scaling group.
Once the new instances are started, join them to the cluster and wait until they are healthy and in sync with the cluster.
Note
Make sure to join the nodes with a single API call to avoid replicating the cluster state multiple times.
Add scale in protection on the new nodes.
Remove the old nodes from the cluster.
Change the minimum, maximum and desired size of the auto scaling group to their original values.
Remove the scale in protection.