How to deploy GraphDB in GCP

GraphDB can be deployed on Google Cloud Platform (GCP) by following the general installation instructions. You can find information regarding the costs of running a GraphDB instance on the Google Cloud Platform website.

This documentation will walk you through the process of setting up the necessary environment for deploying GraphDB on Google Cloud Platform.

Note

Ontotext maintains a Terraform module that automates the entire procedure of deploying GraphDB on Google Cloud Platform. Learn more about how to use it at our GitHub repository.

Architecture

The GraphDB architecture diagram showcases the deployment architecture for GraphDB on Google Compute Engine instance in Google Cloud Platform. The diagram illustrates the key components, and their interactions to provide a high-level understanding of the system’s architecture and how it should be deployed. There are no third-party integration points on the default GraphDB deployment.

GraphDB architecture diagram showcasing deployment on GCE instance, detailing key components and their interactions for a comprehensive system overview.

Prerequisites

There are several prerequisites for running a GraphDB instance on GCP:

Note

We recommend the use of an Identity and Access Management user for the deployment instead of a root user account.

Technical requirements

The following Google Cloud Platform services are required to complete the GraphDB deployment on GCP:

Service

Description

Google Compute Engine (GCE)

Server instance that are used to host the database application

Cloud Balancing

Used for load balancing the GraphDB cluster nodes

Google Cloud Persistent Disk

Persistent Disk volumes are used for storing the data

Google Cloud Identity and Access Management (IAM)

Provides user and access management for your GraphDB deployment

Google Cloud Secrets Manager

Used to store various GraphDB configurations

Google Cloud Storage

Used to store GraphDB Backups

Cloud Monitoring

Used to monitor the status of the GraphDB cluster

Required skills

The following skills and knowledge are typically required in order to successfully deploy GraphDB on GCE Instance:

Google Cloud Platform Fundamentals

Familiarity with Google Cloud Platform (GCP) and understanding of its core concepts, such as Google Cloud Engine instances, security, VPCs and IAM roles. Knowledge of how to navigate the Google Cloud Platform Web Console and interact with GCP services is essential.

Google Cloud Engine Instance Management

Proficiency in creating and managing Google Cloud Engine instances. This includes selecting the appropriate instance type, configuring security settings, managing storage (Cloud Persistent Disk Volumes), and understanding Google Cloud Engine instance lifecycle management.

Networking and Security

Understanding of networking concepts in GCP, including VPC (Virtual Private Cloud) configuration, subnets, routing tables, and firewall rules. Knowledge of how to set up inbound and outbound traffic rules to allow communication with GraphDB.

Linux Administration

Proficiency in Linux command-line interface (CLI) and basic administration tasks. This includes SSH access to GCE instances, navigating the file system, managing permissions, installing packages, and configuring system settings.

Database Management

Knowledge of GraphDB and its deployment requirements. Understanding of how to configure GraphDB settings, including database storage, memory allocation, and repository creation.

Database Backup and Recovery

Familiarity with backup and recovery strategies for GraphDB on GCP. Knowledge of GCP services like Google Cloud Storage for data backups and restoration processes.

Monitoring and Troubleshooting

Proficiency in monitoring the health and performance of GraphDB instances on GCP. Understanding of logging, monitoring and troubleshooting techniques using GCP Cloud Monitoring, GCE instance logs, and GraphDB diagnostic tools.

High Availability and Scalability

Knowledge of implementing high availability and scalability for GraphDB on GCP. This may involve using features like GCP Auto Scaling, load balancers, and multi-zone deployments.

Infrastructure as Code (IaC)

Familiarity with Infrastructure as Code principles and tools like Terraform. This enables automating the provisioning and configuration of GraphDB infrastructure on GCP.

Security Best Practices

Understanding of security best practices for GCP deployments, including data encryption, access controls, identity and access management, and compliance considerations.

Setting up your Virtual Private Cloud (VPC)

  1. Go to the VPC Networks and click on Create VPC Network.

  2. Under Subnet Creation mode, select the Automatic option. This will allow you to configure the subnets in the regions.

  3. Provide a descriptive Name by which to recognize your VPC.

  4. Use the table below as a reference to define the Firewall Rules to allow inbound connections to the GraphDB instance.

    Name 1

    Type

    Targets

    Filters

    Protocols / ports

    Action

    Priority 2

    rule-allow-custom

    Ingress

    Aply to all

    IP ranges: 10.128.0.0/9

    all

    Allow

    65,000

    rule-allow-icmp

    Ingress

    Aply to all

    IP ranges: 0.0.0.0/0

    icmp

    Allow

    65,000

    rule-allow-rdp

    Ingress

    Aply to all

    IP ranges: 0.0.0.0/0

    tcp:3389

    Allow

    65,000

    rule-allow-ssh

    Ingress

    Aply to all

    IP ranges: 0.0.0.0/0

    tcp:22

    Allow

    65,000

    rule-deny-all-ingress

    Ingress

    Aply to all

    IP ranges: 0.0.0.0/0

    all

    Deny

    66,000

    rule-allow-all-egress

    Egress

    Aply to all

    IP ranges: 0.0.0.0/0

    all

    Allow

    66,000

    1

    The names in the table are illustrative and describe what the rule does. Use names that make sense for your deployment.

    2

    The priorities given in the table are illustrative. What is important is that the priority of rules rule-deny-all-ingres and rule-allow-all-egress is higher than of the rest of the rules.

  5. Click on Create VPC

    Tip

    It may take a while for the VPC to deploy. You will know that it is done deploying when you the message “Successfully created network” appears on the bottom of your screen.

Setting up your Cloud DNS private hosted zone

The GraphDB Raft implementation requires static addresses. This is achieved by creating a private hosted zone in Google Cloud DNS and registering the instances there.

  1. From the Cloud DNS dashboard, click on Create DNS Zone.

  2. Under Zone type, select Private.

  3. Provide a Zone name.

  4. Provide a DNS name such as graphdb.cluster.

  5. Under the Networks section select the networks from which you want to have access to the hosted zone.

  6. Click Create

Note

Later, you will also need to create “A” records for the instances.

Creating a Cloud Storage Bucket

Tip

This step is optional, but recommended.

GraphDB can store backups to Cloud Storage and, if needed, restore from them.

To create an Cloud Storage bucket:

  1. From the Cloud Storage console, click on Create.

  2. Provide a globally unique name.

  3. Select your region, scroll down and click Create bucket.

In order to block all public access when you create a bucket you should tick Enforce public access prevention on this bucket on the Choose how to control access to objects section.

Importing a certificate into GCP Certificate Manager

Tip

This step is optional.

While serving GraphDB requests over a secured and encrypted connection is not strictly required, it is nevertheless highly recommended. This section goes over the process of importing a certificate into GCP Certificate Manager so that you can use that certificate later when creating the load balancer.

  1. Go to the Certificate Manager console and click on Create.

  2. Paste or upload your Certificate.

  3. Paste or upload your Private Key.

    Warning

    You need to remove the password for the key before pasting it.

  4. You can optionally add labels, then click on Create

Setting up the Load Balancer

  1. Go to Load Balancing service and click on Create load balancer.

  2. Click on Network Load Balancer and provide the following configurations:

    • Type of load balancer: Network Load Balancer

    • Proxy or passthrough: Proxy load balancer

    • Public facing or internal: Public facing (external)

      Tip

      Otherwise GraphDB will not be accessible externally.

    • Global or single region deployment: Best for global workloads

    • Load balancer generation option: Global external proxy Network Load Balancer

  3. Then click on the Configure button.

On the next screen, provide a unique Load Balancer name.

  1. Select your Region from the Region dropdown field.

  2. Choose the Network in which you want to deploy the load balancer.

Under Backend configuration:

  1. Backend type: Select Instance group.

  2. Protocol: Select TCP.

  3. Under Instance Groups, configure the following parameters:

  • Instance group: Select the instance group that you used to deploy the GraphDB nodes.

  • Port numbers: 7200

  1. Under Healthcheck configuration, configure the following parameters:

    • Protocol: HTTP

    • Port: 7201

    • Request path: /rest/cluster/node/status

  2. Under Frontend configuration, configure the following parameters:

    • Name: Provide a unique name.

    • Protocol: SSL

    • Port: 443

    • Certificate: Choose the certificate that you have created in the previous section.

  3. Click Create.

Setting up service account

The GCE instances require certain permissions in order to perform several tasks in GCP. This section describes the permissions they need, what they are used for, and how to create them.

To achieve all of this, create an instance profile and then attach it to the instances:

  1. Go to the Identity and Access Management (IAM) dashboard and select Service accounts from the navigation menu on the left.

  2. On the service account details provide the following configurations:

    • Define name for the Service account name.

    • Define Service account description.

  3. On the Grant this service account access to project (optional) step you can specify the role for the service that you want to access from the GCE instances.

  4. On the Grant users access to this service account (optional) step you can specify Service account user roles or Service account admin roles that can access the service account.

  5. Click on Done.

Note

If you are planning to store GraphDB backups to Cloud Storage, the GCE instance needs to be able to read, write, and list objects in Cloud Storage service. Perform the steps from Setting up service account, but under Grant this service account access to project, select Cloud Storage Admin Role.

Launching GraphDB instances

Before you can launch your GraphDB instance, you will need to create an instance template with an existing GraphDB image. You can do it by using the following command using the gcloud command line interface:

gcloud beta compute instance-templates create graphdb-instance-template \
--project=example-project-name --machine-type=n2-highmem-8 \
--network-interface=network=default,network-tier=PREMIUM,stack-type=IPV4_ONLY \
--instance-template-region=us-central1 --maintenance-policy=MIGRATE \
--provisioning-model=STANDARD  \
--scopes=https://www.googleapis.com/auth/cloud-platform \
--create-disk=auto-delete=yes,boot=yes,device-name=graphdb-instance-template,image=projects/graphdb-public/global/images/ontotext-graphdb-10-7-6-202410151959,mode=rw,size=20,type=pd-balanced \
--create-disk=device-name=persistent-disk-1,mode=rw,size=500,type=pd-balanced \
--no-shielded-secure-boot --shielded-vtpm --shielded-integrity-monitoring \
--reservation-affinity=any

Setting up Instance Group

In order to create an instance group based on the instance template that you have created in the previous section, execute the following steps in the GCP Web Console.

  1. Go to Compute Engine Service.

  2. On the left tab select Instance Groups option.

  3. Click on Create button and fill the fields below.

    • Name: Define a name that you want to use for the instance template.

    • Description (Optional): Provide a description to help easily identify the deployed resources.

    • Instance template: Select the instance template that you have created via the gcloud CLI, as described in Launching GraphDB instances.

    • Location: Choose Multiple zones and check if the three zones in the region you selected are selected (they should be selected by default).

    • Autoscaling: Set both Minimum number of instances and Maximum number of instances to 3.

    • Port Mapping: 7200

  4. Click on the Create button

Creating a cluster

In order to create the cluster, you will need to get the address of the GCE instances.

  1. Go to the Cloud DNS dashboard and open your private hosted zone.

  2. Write down the names of all records with type “A”.

Once you’ve written down all records names, you can create a cluster by following the Creating and managing a cluster documentation from one of the instances.

Tip

The recommended way to gain access to the instances is by going to Compute Engine and selecting the instance you want to access, then clicking on SSH.

Opening GraphDB instances

Once all instances are running, you can launch your GraphDB instance:

  1. Access the Load balancer.

  2. Copy the DNS name.

  3. Paste it in the address bar of your browser and press Enter.

Updating GraphDB configurations and versions

Since GraphDB and its configurations are baked into the Instance Image, you will need to recreate the GCE instances when updating your GraphDB configuration to a newer minor version. You can do this by either manually stopping each individual instance, or scaling the cluster out, and then scaling it back in. This section describes both methods in detail.

Stopping individual GCE instances

The faster way to update to a newer minor version and its configuration is to stop each individual GCE instance. The downside to this method is that you are decreasing the cluster HA. In other words, if a node fails while recreating an instance, the cluster will be unable to process writes. The process is simple:

  1. Update the Instance Image in the instance template.

  2. Terminate the instances one by one, starting with the follower nodes, and leaving the leader node to be the last instance to be terminated.

See also

To avoid compatibility issues, also refer to the Migrating GraphDB configurations documentation.

Warning

When you terminate an instance, wait for the new one to be started. Then verify that it has successfully rejoined the cluster and that it is in sync before proceeding with the next one.

Scaling the cluster out and then back in

You can also recreate the GCE instances by scaling the cluster out and in. The advantage of this approach is that the HA will not be impacted. However, the cluster will need to replicate its state to the new nodes. This can take a significant amount of time, especially with bigger-sized repositories.

  1. Update the Instance Image in the launch instance template.

  2. Double the size of the cluster.

    Note

    Change the minimum, maximum and desired size of the auto scaling group.

  3. Once the new instances are started, join them to the cluster and wait until they are healthy and in sync with the cluster.

    Note

    Make sure to join the nodes with a single API call to avoid replicating the cluster state multiple times.

  4. Add scale in protection on the new nodes.

  5. Remove the old nodes from the cluster.

  6. Change the minimum, maximum and desired size of the auto scaling group to their original values.

  7. Remove the scale in protection.