Backup and restore

GraphDB supports the backup and restore of both a single GraphDB instance and a cluster through its recovery REST API. Both partial (per-repository) and full recovery procedures are available with optional inclusion of user account data. Backups can be created both locally and on select cloud offerings.

Note

As with all operations that involve a REST API, in order to perform a backup or a restore procedure:

  • The respective GraphDB instance must be online.

  • The cluster must be writable — in other words, the majority of its nodes must be active.

Starting with version 10.4.0, GraphDB uses LZ4 compression for backup and restore. Parallel backup compression and parallel S3 streaming are only available with licenses with four or more licensed cores.

Warning

Compressed backups created with GraphDB 10.4.0 and newer are not backward compatible and cannot be restored with older versions of GraphDB. However, restore procedures are backward compatible — in other words, in GraphDB 10.4.0 and newer you are able to restore backups created with older versions of GraphDB.

Planning a backup

Whether you want to be able to quickly recover your data in case of failure or perform routine admin operations such as upgrading a GraphDB instance, it is important to prepare an optimal backup & restore procedure.

There are various factors to take into consideration when designing a backup strategy, such as:

  • Optimal timing for downtime tolerance for applying backup

  • Read-only tolerance on a single node setup for creating a backup

  • Load-balanced backup creation (backup is created by one of the followers, so if a quorum exists, updates will be processed)

  • Scope of the backed-up data (for example, full or per-repository backup, or whether user accounts and settings are included)

  • Available system resources and specifically ensuring enough disk space for backup

  • Frequency of backup creation

Note

Only administrators are able to create and restore backups.

It’s worth keeping in mind that creating and restoring a backup will change the read and write behavior of GraphDB. These changes are affected by the type of operation (read or write), the type of backup (full or partial), and the deployment (a single GraphDB instance or a cluster).

Based on those criteria, an operation can be affected in one of three ways:

  • ✅ Allowed (not affected by the backup or restore)

  • ❌ Rejected

  • ⏳ Waiting (it will be processed once the backup or restore procedure has been completed)

Backup

Backup type

Deployment

Write

Read

Write in affected repository

Write in unaffected repository

Read in affected repository

Read in unaffected repository

Partial

Cluster

N/A

✅

✅

✅

N/A

N/A

Single node

N/A

✅

❌

✅

N/A

N/A

Full

Cluster

✅

✅

N/A

N/A

N/A

N/A

Single node

❌

✅

N/A

N/A

N/A

N/A

Restore

Restore type

Deployment

Write

Read

Write in affected repository

Write in unaffected repository

Read in affected repository

Read in unaffected repository

Partial

Cluster

N/A

N/A

⏳

⏳

❌

✅

Single node

N/A

N/A

⏳

✅

❌

✅

Full

Cluster

❌

❌

N/A

N/A

N/A

N/A

Single node

❌

❌

N/A

N/A

N/A

N/A

Note

For the purposes of this table, restoreSystemData should be considered as full restore since it also affects user data, which is not repository-specific.

Planning cloud backups

You can also create a backup saved in the cloud, and restore from backup stored on cloud storage. The supported cloud storage options are Amazon S3 and S3-compatible services, Azure Blob, and Google Cloud Storage.

Cloud backup and restore have the same options as regular GraphDB backup and regular GraphDB restore, with an additional bucketUri parameter that contains all the information about the cloud bucket.

Creating an Amazon S3 or compatible services bucket

Amazon S3

For Amazon’s S3, the bucketUri parameter uses the following format:

s3:///<bucket-name>/<backup-name>?region=<AWSRegion>&AWS_ACCESS_KEY_ID=<key-id>&AWS_SECRET_ACCESS_KEY=<access-key>

S3 compatible services

For S3 compatible services, the bucketUri parameter uses the following format:

s3://[<endpoint-hostname>:<endpoint-port>]/<bucket-name>/<backup-name>?region=<AWSRegion>&AWS_ACCESS_KEY_ID=<key-id>&AWS_SECRET_ACCESS_KEY=<access-key>

The endpoint-hostname and endpoint-port values are only used for S3 compatible services.

If either AWS_ACCESS_KEY_ID or AWS_SECRET_ACCESS_KEY isn’t provided, it will fall back to the AWS standardized credential providers.

If region is not provided, it will attempt to use the default AWS region provider chain to resolve the configured region. If that fails, it will default to us-east-1 — US East (N. Virginia). Furthermore, if the provided region does not match the one where the bucket is located, once a connection to the bucket has been established, all subsequent requests will automatically be directed to the region the bucket resides in.

In addition to the default backup options, you can also set up these global parameters in your graphdb.properties file before starting your GraphDB instance:

Property name

Description

Default value

graphdb.s3.tls.enabled

Enables TLS for secure connection to S3 compatible services

false

graphdb.s3.backup.httpclient.write.timeout

Timeout in seconds for a cloud backup’s single part upload

3600

graphdb.s3.backup.disable.checksum

Disables the S3 checksum validation when uploading backups.

Warning

Set to true only if the S3 compatible service you are using doesn’t support trailing checksums.

false

Creating an Azure Blob bucket

For Azure Blob, the bucketUri parameter uses the following format:

az://<container-name>/<backup-name>?blob_storage_account=<storage_account_name>

The container-name, backup-name and storage_account_name values are required.

For shared account key, the bucketUri parameter uses the following format:

az://<container-name>/<backup-name>?blob_storage_account=<storage_account_name>&blob_access_key=<SAKey>

The SAS token is already in Uri format, which is why it is appended after the storage account.

Warning

Some SAS token generators include a ? in front of the SAS token. The ? is a delimiter character — it is not a part of the SAS token and should not be included.

Example token usage

az://<container-name>/<backup-name>?blob_storage_account=<storage_account_name>&sv=2022-11-02&sr=b&sig=<signature>&sp=rcw

If both the shared account key and the SAS token aren’t provided, Azure will attempt to authenticate from a default identity chain. For more information, check the Azure documentation on the DefaultAzureCredential Class.

Creating a Google Cloud Storage bucket

For Google Cloud Storage, the bucketUri parameter uses the following format:

gs://<bucket-name>/<backup-name>

The credentials must be provided as files. If no authentication file is provided, the authentication will fall back to the application’s default credential chain. To learn how to create credentials, check the Google Cloud Platform documentation on creating credentials for service accounts.

Creating a backup

As mentioned, backups can be either covering all repositories (full data backup) or only selected existing repositories (partial data backup), and they may also include the user accounts and settings.

  • Full data backups will wait until all running updates are completed.

  • Partial data backups will wait until running updates on the affected repositories are completed.

Note

Cluster backup creation is lock-free, meaning that by leveraging the multiple instances and quorum mechanism, the cluster can create a backup while simultaneously processing updates if the deployment has more than 2 nodes.

A GraphDB instance can be backed up using the /rest/recovery/backup endpoint. To create a backup, simply POST an HTTP request as shown further down below.

Tip

The backup is created in a .tar archive. If the backup creation has been successful, the archive should contain a .success file. If there is no such success file, the backup is corrupted and GraphDB will not be able to restore data from it.

Backup options

The following parameters can be configured when creating a backup:

Option

Description

repositories

List of repositories to be backed up. Specified as JSON in the request body.

  • If the parameter is missing, all repositories will be included in the backup.

  • If it is an empty list ([]), no repositories will be included in the backup.

  • Otherwise, the repositories from the list will be included in the backup.

backupSystemData

Determines whether user account data such as user accounts, saved queries, or visual graphs, among other. should be included in the backup. Specified as JSON in the request body. Boolean, the default value is false.

Full data backup

Here is an example curl request for full data backup creation without system data (i.e., user accounts and settings):

curl -X POST -OJ -H 'Content-Type: application/json' '<base_url>/rest/recovery/backup'

This does the following:

  • Backs up all data in all repositories.

  • Does not include user accounts and settings because backupSystemData = false by default (see the above Backup options).

  • Creates the backup as a new file of the type backup-yyyy-mm-dd-hh-mm-ss.tar.

Note

This is an archive file that you do not need to extract — it is to be used as is.

To set the name of the backup yourself, replace -OJ with --output <backup-name>, i.e.:

curl -X POST --output <backup-name> -H 'Content-Type: application/json' '<base_url>/rest/recovery/backup'

Partial data backup

Here is an example curl request for partial data backup creation without system data:

curl -X POST -OJ -H 'Content-Type: application/json' -d '{
   "repositories":["<repo_name>"]
}' '<base_url>/rest/recovery/backup'

Which does the following:

  • Backs up one or more repositories that are explicitly named.

  • Does not include user accounts and settings as backupSystemData = false by default.

You can also use --output <backup-name> instead of -OJ if you want to customize the name of the backup as shown above.

Note

If a POST request does not include a list of repositories for backup, it will automatically create a full data backup.

Full data and system backup

Here is an example curl request for full data and system backup creation with system data:

curl -X POST -OJ -H 'Content-Type: application/json' -d '{
   "backupSystemData": true
}' '<base_url>/rest/recovery/backup'

Which does the following:

  • Backs up all data in all repositories.

  • Backup includes user accounts and settings as backupSystemData = true is explicitly provided.

Backup of only system data

Here is an example curl request for creating a backup of system data only:

curl -X POST -OJ -H 'Content-Type: application/json' -d '{
   "repositories" : [], "backupSystemData": true
}' '<base_url>/rest/recovery/backup'

Which does the following:

  • Backup includes user accounts and settings as backupSystemData = true is explicitly provided.

  • No repositories are included in the backup as repositories is an empty list (repositories: []).

    Note

    If this parameter is not provided, all repositories will be included in the backup.

Creating a cloud backup

The GraphDB instance uses a different endpoint when creating a backup saved in the cloud — /rest/recovery/cloud-backup. This endpoint takes several parameters.

  • Repositories (optional): List of repositories to be backed up. If nothing is passed for it, the default values of the options will be used. See the section on backup options above for further details.

  • backupSystemData (optional): Determines whether user account data. If nothing is passed for it, the default values of the options will be used. See the section on backup options above for further details.

  • bucketUri (required)

  • authenticationFile (optional)

Below are example cURL request for full data backup creation with system data.

Creating a backup in Amazon S3 and compatible services

curl -X POST --header 'Content-Type: multipart/form-data' --header 'Accept: application/json' \
  -F 'params={"repositories" :["<repo-name>"], "backupSystemData" :<boolean>,
  "bucketUri": "s3:///<bucket_name>/<backup_name>?region=<region>&AWS_ACCESS_KEY_ID=<key_id>&AWS_SECRET_ACCESS_KEY=<key>"}' \
  '<base_url>/rest/recovery/cloud-backup'

Creating a backup in Azure Blob

curl -X POST --header 'Content-Type: multipart/form-data' --header 'Accept: application/json' \
  -F 'params={"repositories": ["<repo-name>"],"backupSystemData": <boolean>,
  "bucketUri": "az://<container-name>/<backup-name>?blob_storage_account=<storage_account_name>&blob_access_key=<key>"}' \
  '<base_url>/rest/recovery/cloud-backup'

Creating a backup in Google Cloud Platform

curl -X 'POST' \
  -H 'accept: application/json' \
  -H 'Content-Type: multipart/form-data' \
  -F 'params={"repositories":["<repo-name>"],"backupSystemData":<boolean>,"bucketUri":"gs://<bucket>/<backup-name>"}' \
  -F 'authenticationFile=@/path/to/file/google-credentials.json' \
  '<base_url>/rest/recovery/cloud-backup'

Tip

The examples for Full data backup, Partial data backup, Full data and system backup, and Backup of only system data shown above are also valid for the cloud backup. As long as the cloud backup is provided with the same parameters and the bucketUri is valid, the resulting backup .tar file should be the same.

Restoring from a backup

A GraphDB instance or cluster can be restored to a backed-up state through the /rest/recovery/restore endpoint.

The recovery procedure in the cluster is treated as a simple update as it leverages the Raft protocol that allows a set of distributed nodes to act as one.

Warning

If the backup is corrupted, GraphDB will not be able to restore data from it. A healthy backup should contain a .success file in its package. If there is no such file, the backup is considered corrupted.

Note

It is recommended to perform cluster transaction log truncate operations after a successful data restore, as the transaction log will use more storage space upon a backup/restore procedure.

To restore a backup, simply POST an HTTP request as shown below.

Restore options

The following parameters can be configured when restoring from a backup:

Option

Description

repositories

List of repositories to recover from the backup. Specified as JSON in the request body.

  • If the parameter is missing, all repositories that are in the backup will be restored.

  • If it is an empty list ([]), no repositories from the backup will be restored.

  • Otherwise, the repositories from the list will be restored.

restoreSystemData

Determines whether GraphDB should restore user account data such as user accounts, saved queries, visual graphs etc. from a backup or continue with the their current state. Specified as JSON in the request body. If no system data is found in the backup, an error will be returned. Boolean, the default is false.

removeStaleRepositories

Cleans other existing repositories on the GraphDB instance where the restore is done. The default is false, meaning that no repositories will be cleaned.

Full data restore preserving other repositories

If we have successfully created a backup and want to completely revert to the backed-up state while preserving the existing repositories on the instance where we are restoring, we can use the below curl request example. No additional parameters are provided, meaning that defaults are applied.

curl -X POST '<base_url>/rest/recovery/restore' \
-H 'Content-Type: multipart/form-data' \
-F 'params={
        };type=application/json' \
-F file=@./<full-data-backup-name.tar>

Note

The full-data-backup-name.tar file must be a Full data backup.

Full data restore with replace

We can also apply a backup and remove repositories that are not restored from it.

curl -X POST '<base_url>/rest/recovery/restore' \
-H 'Content-Type: multipart/form-data' \
-F 'params={
        "removeStaleRepositories": true \
    };type=application/json' \
-F file=@./<full-data-backup-name.tar>

What this does:

  • Removes other repositories on the instance where the backup is applied as removeStaleRepositories = true.

  • Does a full data restore as the repositories parameter is not provided.

Partial data restore

Here, we need to provide the names of the repositories that we want to restore as values for the repositories parameter.

curl -X POST '<base_url>/rest/recovery/restore' \
-H 'Content-Type: multipart/form-data' \
-F 'params={
        "repositories" : ["<repo-name>"] \
    };type=application/json};type=application/json' \
-F file=@./<full-data-backup-name.tar>

Restoring only system data

To restore only the system data from a backup, we can use the following curl request:

curl -X POST '<base_url>/rest/recovery/restore' \
-H 'Content-Type: multipart/form-data' \
-F 'params={
        "repositories" : [], \
        "restoreSystemData": true \
    };type=application/json}' \
-F file=@./<full-data-system-backup-name.tar>

What this does:

  • User account data is restored as restoreSystemData = true.

  • No repositories are restored as the repositories parameter is an empty list ([]).

    Note

    The full-data-system-backup-name.tar file must contain system data, i.e., the backup must be created with backupSystemData = true as shown in the section Full data and system backup.

Restoring from a cloud backup

The GraphDB instance uses a different endpoint when restoring from a backup saved on cloud storage — /rest/recovery/cloud-restore.

Below are example cURL request for applying a backup and removing all repositories that are not restored from it. In case you want to do a full backup (as opposed to a partial one), omit the repositories parameter.

Restoring a backup in Amazon S3 and compatible services

curl -X POST --header 'Content-Type: multipart/form-data' --header 'Accept: application/json' \
  -F 'params={"repositories":["<repo-name>"],"removeStaleRepositories":<boolean>,"restoreSystemData":<boolean>,
  "bucketUri":"s3:///<bucket_name>/<backup_name>?region=<region>&AWS_ACCESS_KEY_ID=<key_id>&AWS_SECRET_ACCESS_KEY=<key>"}' \
  '<base_url>/rest/recovery/cloud-restore'

Restoring a backup in Azure Blob

curl -X POST --header 'Content-Type: multipart/form-data' --header 'Accept: application/json' \
  -F 'params={"repositories":["<repo-name>"],"removeStaleRepositories":<boolean>,"restoreSystemData":<boolean>,
  "bucketUri":"az://<container-name>/<backup-name>?blob_storage_account=<storage_account_name>&blob_access_key=<key>"}' \
  '<base_url>/rest/recovery/cloud-restore'

Restoring a backup in Google Cloud Storage

curl -X 'POST' \
  -H 'accept: application/json' \
  -H 'Content-Type: multipart/form-data' \
  -F 'params={"repositories":["<repo-name>"],"removeStaleRepositories":<boolean>,"restoreSystemData":<boolean>,"bucketUri":"gs://<bucket>/<backup-name>"}' \
  -F 'authenticationFile=@/path/to/file/google-credentials.json' \
  '<base_url>/rest/recovery/cloud-restore'

Tip

The examples for Full data restore preserving other repositories, Full data restore with replace, Partial data restore, and Restoring only system data shown above are also valid for the cloud restore. As long as the cloud restore is provided with the parameters and the bucketUri is a valid GraphDB backup file, the resulting restore should be the same.

Backup compatibility between different versions of GraphDB

When restoring a backup, compatibility between different versions of GraphDB mostly depends on whether you are restoring between different patch versions, minor versions, or major versions of GraphDB. GraphDB strives to present backward backup compatibility in most cases and generally you should be able to restore backups made with older versions of GraphDB than the one you are currently using, but this may not always be the case.

Backup-restore compatibility between different patch versions

We always allow unconditional backward and forward backup-restore between patch versions of the same minor release. For example, backup from any GraphDB 11.3 patch release, such as 11.3.3, should be restored on any 11.3.x version, even if it is a previous patch version.

Backup-restore compatibility between different minor versions

Forward backup-restore is supported between minor versions. A backup created on 11.2.x should be restored on any patch release for 11.3, 11.4, and so on.

Backward backup-restore between minor versions is supported by default, but is not unconditional. For some minor versions, support for backward backup-restore may be dropped to accomodate for fixing major and critical bugs, or for improvements and new feautres. In such case this will be explicitly documented in the Migrating GraphDB configurations page.

Backup-restore compatibility between different major versions

Forward backup-restore is supported by default between a new major version and the latest previous minor release by default. Exceptions may occur and will be explicitly documented in the Migrating GraphDB configurations page. For example you can restore from 10.8.x to any patch release for GraphDB 11 (10.8 is the last 10 minor version).

Backup-restore skipping a major version is not supported. In other words, restoring a backup from GraphDB 9 to GraphDB 11 is not possible.

Backward backup-restore is not supported between major versions. This will not be documented explicitly for each new major version.

Monitoring your recovery operations

You can monitor your backups through Monitor ‣ Backup and Restore. The backup monitoring interface displays the backups and restores that are currently underway, and displays additional information, such as the recovery operation type (backup or restore), the user who initiated the procedure, the affected repositories, how much time has elapsed since the procedure was initiated, and the snapshot options.

This interface also allows you to temporarily Pause updates of the table so that you can copy text from it.

GraphDB Workbench monitoring interface displaying the progress of backup and recovery operations.

Warning

The Pause button doesn’t pause the recovery operation itself — it merely freezes the information presented in Workbench and prevents the table from updating.

You can also access the Backup and Restore interface through the The Workbench notification area notification area on the top of every Workbench view.

Notification area in GraphDB Workbench indicating access to global monitoring features, including backup and restore monitoring.

Monitoring your recovery operations through cURL

Track backup operations

GET /rest/monitor/backup

Example:

curl -X GET --header 'Accept: application/json' '<base_url>/rest/monitor/backup'

This command will return the current recovery operation. The response will specify whether the operation is a backup or a restore.