Directories and configuration properties

GraphDB relies on several main directories for configuration, logging, and data.

Directories

GraphDB home

The GraphDB home defines the root directory where GraphDB stores all of its data. The home can be set through the system or config file property graphdb.home.

The default value for the GraphDB home directory depends on how you run GraphDB:

  • Running as a standalone server: the default is the same as the distribution directory.

  • All other types of installations: OS-dependent directory.

    • On Mac: ~/Library/Application Support/GraphDB.

    • On Windows: \Users\<username>\AppData\Roaming\GraphDB.

    • On Linux and other Unixes: ~/.graphdb.

Note

In the unlikely case of running GraphDB on an ancient Windows XP, the default directory is \Documents and Settings\<username>\Application Data\GraphDB.

GraphDB does not store any files directly in the home directory, but uses the following subdirectories for data or configuration:

Data directory

The GraphDB data directory defines where GraphDB stores repository data. The data directory can be set through the system or config property graphdb.home.data. The default value is the data subdirectory relative to the GraphDB home directory.

Configuration directory

The GraphDB configuration directory defines where GraphDB looks for user-definable configuration. It can be set through the system property graphdb.home.conf.

Note

It is not possible to set the config directory through a config property as the value needs to be set before the config properties are loaded.

The default value is the conf subdirectory relative to the GraphDB home directory.

Work directory

The GraphDB work directory defines where GraphDB stores non-user-definable configuration. The work directory can be set through the system or config property graphdb.home.work. The default value is the work subdirectory relative to the GraphDB home directory.

Logs directory

The GraphDB logs directory defines where GraphDB stores log files. The logs directory can be set through the system or config property graphdb.home.logs. The default value is the logs subdirectory relative to the GraphDB home directory.

Note

When running GraphDB as deployed .war files, the logs directory will be a subdirectory graphdb within the Tomcat’s logs directory.

Tip

Even though GraphDB provides the means to specify separate custom directories for data, configuration and so on, it is recommended to specify the home directory only. This ensures that every piece of data, configuration, or logging, is within the specified location.

Step-by-step guide:

  1. Choose a directory for GraphDB home, e.g., /opt/graphdb-instance.

  2. Create the directory /opt/graphdb-instance.

  3. (Optional) Copy the subdirectory conf from the distribution into /opt/graphdb-instance.

  4. Start GraphDB with graphdb -Dgraphdb.home=/opt/graphdb-instance.

GraphDB creates the missing subdirectories data, conf (if you skipped that step), logs, and work.

Distribution directory

The distribution directory stores the GraphDB binaries and additional files such as sample data, template files, and command line tools.

When running GraphDB Desktop, the simplest way to find the distribution directory is by clicking the Open dist directory button on the GraphDB Desktop application window. See Configuring your running GraphDB instance for more information on the application window.

When running the GraphDB standalone server (as opposed to the GraphDB Desktop) the distribution directory is the directory where the downloaded distribution file was unzipped.

To learn more about the structure of the distribution directory, see Structure of the distribution package.

Checking the configured directories

When GraphDB starts, it logs the actual value for each of the above directories, e.g.:

GraphDB Home directory: /path-to-graphdb/app
GraphDB Config directory: /path-to-graphdb/app/conf
GraphDB Data directory: /path-to-graphdb/app/data
GraphDB Work directory: /path-to-graphdb/app/work
GraphDB Logs directory: /path-to-graphdb/app/logs

Configuration

There is a single graphdb.properties config file for GraphDB. It is provided in the distribution under conf/graphdb.properties, where GraphDB loads it from.

This file contains a list of config properties defined in the following format:

propertyName = propertyValue, i.e., using the standard Java properties file syntax.

Each config property can be overridden through a Java system property with the same name, provided in the environment variable GDB_JAVA_OPTS, or in the command line.

Storing property values in a separate file

You can also store the value for any configuration property in a separate text file, and point to that file from graphdb.properties. The path can be either relative of absolute. Relative paths are relative to the GraphDB /conf directory found in the distribution folder.

The text file containing the value must contain a single non-blank line that contains only the property value, although blank lines before or after are allowed. The file has a size limit of 1024 bytes.

Property names pointing to a file end with @file. For example, graphdb.auth.ldap.bind.userDn.password=PASSWORD_VALUE is how the LDAP bind user password property would appear in graphdb.properties (in this example, with a value of PASSWORD_VALUE). The same can be achieved by setting graphdb.auth.ldap.bind.userDn.password@file=/opt/mysecret.txt, where the file mysecret.txt contains that same value PASSWORD_VALUE as a single line of text.

This mechanism provides more flexibility to configure your GraphDB instances, as you won’t need to configure the individual graphdb.properties file for each separate instance. For example, a password or a token can be generated and stored in a separate file, referenced in graphdb.properties by using the suffix @file, and distributed among all relevant GraphDB instances without editing each individual graphdb.properties file.

Configuring GraphDB through environment variables

You can also configure GraphDB through the environment variables.

GraphDB reads everything that starts with graphdb., graphdb_, graphdb__ or graphdb___. Environment variables configured this way take precedence over what has been configured in graphdb.properties.

If you configure GraphDB this way, you can circumvent the limited character set available for environment variables in many systems by using the following special character processing:

Environment variable example

How it is processed by GraphDB

Period .

graphdb.property

No changes (graphdb.property)

Underscore _

graphdb_property

No changes (graphdb_property)

Double underscores __

graphdb__property

Processed to . (graphdb.property)

Triple underscore __

graphdb___property

Processed to - (graphdb-property)

Example configurations of GraphDB through externalized environment variables

Configuring GraphDB in Docker

This example shows how to configure GraphDB’s GPT settings with environment variables from Docker Compose.

version: '3.7'

services:
  graphdb:
    image: ontotext/graphdb:10.7.0
    restart: unless-stopped
    ports:
      - "7201:7201"
    environment:
      GDB_JAVA_OPTS: -Xms4G -Xmx4G
      graphdb.connector.port: "7201"
      graphdb.openai.api-key: "<secret-token>"
    volumes:
      - graphdb_data:/opt/graphdb/home

volumes:
  graphdb_data:

Configuring OIDC in Kubernetes

Create a configuration map with the following data and provide it as environment to the pod or container:

apiVersion: v1
kind: ConfigMap
metadata:
  name: graphdb-oidc
data:
  graphdb.auth.methods: gdb,openid
  graphdb.auth.openid.issuer: https://login.microsoftonline.com/[ID]/v2.0
  graphdb.auth.openid.client_id: [CLIENT-ID]
  graphdb.auth.openid.username_claim: email
  graphdb.auth.openid.auth_flow: code
  graphdb.auth.openid.token_type: id
  graphdb.auth.database: local
  graphdb.auth.oauth.roles_claim: roles
  graphdb.auth.oauth.roles_prefix: GDB_
  graphdb.auth.oauth.default_roles: ROLE_USER

And then use it in GraphDB’s pods with:

envFrom:
  - configMapRef:
      name: graphdb-oidc

Changing the cluster secret in Kubernetes

This example highlights why using a Secret object is better than using GDB_JAVA_OPTS when running GraphDB in Kubernetes. Usually, such sensitive configurations are created beforehand and have to be provided as environment, or even mounted and read as files.

Create a secret with the following content and provide it as environment.

apiVersion: v1
kind: Secret
metadata:
  name: graphdb-cluster-secret
data:
  graphdb.auth.token.secret: [YOUR-SECRET]

And then use it in a pod:

envFrom:
- secretRef:
    name: graphdb-cluster-secret

Configuration properties

The properties are of five types and are detailed below.

General properties

The general properties define some basic configuration values that are shared with all GraphDB components and types of installation:

Property name

Description

graphdb.home

Defines the GraphDB home directory

graphdb.home.data

Defines the GraphDB data directory

graphdb.home.conf

(only as a system property) Defines the GraphDB conf directory

graphdb.home.work

Defines the GraphDB work directory

graphdb.home.logs

Defines the GraphDB logs directory

graphdb.dist

If graphdb.dist is set and graphdb.home is not, GraphDB will look for the data, conf, logs, etc. directories there (unless they are explicitly set).

graphdb.workbench.home

The place where the source for GraphDB Workbench is located

graphdb.jsonld.whitelist

Sets whitelist for JSON-LD and NDJSON-LD resources necessary for importing and exporting data to certain JSON-LD or NDJSON-LD document forms.

graphdb.license.file

Sets a custom path to the license file to use

graphdb.page.cache.size

The amount of memory to be taken by the page cache

graphdb.pidfile

The full path to the file where the GraphDB process ID is stored

graphdb.foreground

Tells GraphDB not to close stdout/stderr, but the user can choose whether to daemonize or not

graphdb.heapdump.enable

GraphDB can dump the heap on out-of-memory errors in order to provide insight to the cause for excessive memory usage. This property enables or disables the heap dump. Default is true.

graphdb.heapdump.path

File to write the heap dump to. The default is the heapdump.hprof file in the configured logs directory.
See also the properties graphdb.home and graphdb.home.logs.

graphdb.max.active.repositories

Used to define the maximum number of active repositories, as described in the Limiting the number of active repositories section of the Repository Caching Mechanism documentation. Setting this property to less than 1 (preferably 0) will deactivate the mechanism. The default is 100.

graphdb.repository.expiry.after.minutes

Sets a timeout period in minutes after which a repository that has not been accessed will be shut down by the Repository Caching Mechanism. A value of 0 disables the mechanism. The default is 0.

graphdb.inference.buffer

Buffer size (the number of statements) for each load stage in parallel import. Defaults to 200,000 statements.
See also graphdb.inference.concurrency.

graphdb.inference.concurrency

Number of inference threads in parallel import. The default value is the number of cores of the machine processor.
See also graphdb.inference.buffer.

graphdb.ontop.jdbc.path

GraphDB directory for the JDBC driver used in the creation of Ontop repositories. Use it when you want to set it to a directory different from the lib/jdbc one where the driver is normally placed.

graphdb.lucene.maxclause

Sets the number of maximum number of clauses permitted per Lucene query. The default value is 4096.

Workbench properties

In addition to the standard GraphDB command line parameters, the GraphDB Workbench can be controlled with the following parameters (they should be of the form -Dparam=value):

Parameter

Description

Default value

graphdb.workbench.cors.enable

Enables cross-origin resource sharing.

false

graphdb.workbench.cors.origin

Sets the allowed Origin value for cross-origin resource sharing. This can be a comma-delimited list or a single value. The value “*” means “allow all origins” and it works with authentication too.

*

graphdb.workbench.cors.expose-headers

As per GraphDB’s compliance with the Access-Control-Expose-Headers, when the two parameters above are enabled, this parameter exposes headers other than the CORS-safelisted request headers. They are exposed in a comma-delimited list.

Example:
graphdb.workbench.cors.enable=true

graphdb.workbench.cors.origin=*

graphdb.workbench.cors.expose-headers="content,location"

If no value is set, only the CORS-safelisted request headers will be exposed.

graphdb.workbench.maxConnections

Sets the maximum number of concurrent connections to a remote GraphDB instance.

200

graphdb.workbench.datadir

Sets the directory where the workbench persistence data will be stored.

${user.home}/.graphdb-workbench/

graphdb.workbench.importDirectory

Changes the location of the file import folder.

${user.home}/graphdb-import/

graphdb.workbench.maxUploadSize

Sets the maximum upload size for importing local files. The value must be in bytes.

1 gb

URL properties

Tip

Jump ahead to Typical use cases for a list of examples that cover URL properties usage.

In certain cases, GraphDB needs to construct a URL that refers to itself:

  • The repository list in Setup ‣ Repository manager where each repository provides a link that can be used to access the repository via the REST API.

When GraphDB is accessed directly (without a reverse proxy), it will figure out the correct URLs based on the URL of incoming requests. For example, if GraphDB is accessed using the URL http://graphdb.example.com:7200/, it will construct URLs like http://graphdb.example.com:7200/repositories/repoId.

When GraphDB is accessed via a reverse proxy, the server will not see the actual URL used to access the server and thus it cannot determine a valid external URL on its own. There are two specific setups:

  • The external URL as seen via the proxy uses / as its root, for example, http://rdf.example.com/.

    • GraphDB will map the external / to its own / automatically, no need to add or change any configuration.

    • GraphDB will still not know how to construct external URLs, so setting graphdb.external-url is recommended even though it might appear to work without setting it.

  • The external URL as seen via the proxy uses /something as its root (i.e., something in addition to the /), for example, http://example.com/rdf.

    • GraphDB cannot map this automatically and needs to be configured using the property graphdb.vhosts or graphdb.external-url (see below).

    • This will instruct GraphDB that URLs beginning with http://example.com/rdf/ map to the root path / of the GraphDB server.

The URL properties determine how GraphDB constructs URLs that refer to itself, as well as what URLs are recognized as URLs to access the GraphDB installation. GraphDB will try to auto-detect those values based on URLs used to access it, and the network configuration of the machine running GraphDB. In certain setups involving virtualization or a reverse proxy, it may be necessary to set one or more of the following properties:

Property

Description

graphdb.vhosts

A comma-delimited list of virtual host URLs that can be used to access GraphDB. Setting this property is necessary when GraphDB needs to be accessed behind a reverse proxy and the path of the external URL is different from /, for example http://example.com/rdf.

graphdb.external-url

Sets the canonical external URL. This property implies graphdb.vhosts. If you have provided an explicit value for both graphdb.vhosts and graphdb.external-url, then the URL specified for graphdb.external-url must be one of the URLs in the value for graphdb.vhosts.

When a reverse proxy is in use and most users will access GraphDB through the proxy, it is recommended to set this property instead of, or in addition to graphdb.vhosts, as it will let GraphDB know that the canonical external URL is the one as seen through the proxy.

Tip

Prior to GraphDB 9.8, only the graphdb.external-url property existed. You can keep using it as is.

graphdb.external-url.enforce.transactions

Determines whether it is necessary to rewrite the Location header when no proxy is configured. Setting this property to true will use the graphdb.external-url when building the transaction URLs.

Set it to true when the returned URLs are incorrect due to missing or invalid proxy configurations. Set it to false when the server can be called on multiple addresses, as it will override the returned address to the one defined by the graphdb.external-url.

Boolean, default is false.

graphdb.hostname

Overrides the hostname reported by the machine.

Enabling the configuration will use the graphdb.external-url when building the transaction URLs. It should be used when the returned URLs are not correct due to missing or invalid proxy configurations. The configuration should not be used when the server can be called on multiple addresses as it will override the returned address to a single one defined by the graphdb.external-url.

Note

For remote locations, the URLs are always constructed using the base URL of the remote location as specified when the location was attached.

Typical use cases
  1. GraphDB is behind a reverse proxy whose URL path is / and most clients will use the proxy URL.

    This setup will appear to work out-of-the box without setting any of the URL properties but it is recommended to set graphdb.external-url. Example URLs:

    • Internal URL: http://graphdb.example.com:7200/

    • External URL used by most clients: http://rdf.example.com/

    The corresponding configuration is:

    # Recommended even though it may appear to work without setting this property
    graphdb.external-url = http://rdf.example.com/
    
  2. GraphDB is behind a reverse proxy whose URL path is /something and most clients will use the proxy URL.

    This configuration requires setting graphdb.external-url (recommended) or graphdb.vhosts to the correct URLs as seen externally through the proxy. Example URLs:

    • Internal URL: http://graphdb.example.com:7200/

    • External URL used by most clients: http://example.com/rdf/

    The corresponding configuration is:

    # Required and recommended
    graphdb.external-url = http://example.com/rdf/
    
    # Non-recommended alternative to the above
    #graphdb.vhosts = http://example.com/rdf/
    

Cluster properties

Parameter

Description

Default value

graphdb.cluster.sync.timeoutS

Specifies the maximum duration in seconds that the system waits for the completion of cluster creation before timing out.

600 (10 minutes)

graphdb.cluster.catchup.snapshot.replication.timeoutS

Specifies the maximum duration in seconds that the system waits to replicate the cluster data on a new node before timing out.

10800 (3 hours)

graphdb.cluster.catchup.follower.sync.timeoutS

Specifies the maximum duration in seconds that the system waits for a follower node to catch up with the latest cluster data and get in sync with the other nodes before timing out.

1800 (30 minutes)

graphdb.cluster.proxy.socketTimeout

Specifies the maximum duration in seconds a node waits for various operations (such as connecting to a server, reading data from the socket, or writing data to the socket after establishing a connection with another node in the cluster group) before timing out. Part of custom GraphDB HTTP client configurations representing defaults for request redirection.

-1 (indefinite)

graphdb.cluster.proxy.maxConnectionsTotal

Defines the default maximum number of concurrent connections that a node can establish to other nodes within the cluster group. Part of custom GraphDB HTTP client configurations.

50000

graphdb.cluster.proxy.maxConnectionsPerRoute

Defines the default maximum number of concurrent connections that a node can establish to another node within the cluster group. Part of custom GraphDB HTTP client configurations.

30000

graphdb.cluster.proxy.connectionTimeoutS

Specifies the maximum duration in seconds a node is willing to wait to establish a connection to another node in the cluster group before timing out. Part of custom GraphDB HTTP client configurations.

15

graphdb.cluster.proxy.socket.soTimeout

Specifies the maximum time in milliseconds a node is willing to wait to read data after establishing a connection with another node in the cluster group before timing out. Part of custom GraphDB HTTP client configurations.

18000000 (5 hours) if TLS is enabled, otherwise 0 indicating no timeout (which defaults to 30 seconds).

graphdb.cluster.tls.trust.own.certificate

Determines whether the cluster should accept the provided SSL certificate as a trusted one.

TRUE

graphdb.cluster.node.report.staleAfterSeconds

Specifies the time in seconds after which a new system report will be triggered upon receiving a customer request, provided the specified time has elapsed since the previous report was made.

60

graphdb.cluster.observers.autoUnregisterAfterMinutes

Specifies the time in minutes after which an observer is automatically unregistered if it has not received an update within that timeframe.

15

The raft properties below control the security configuration for the gRPC communications.

Parameter

Description

Example value

graphdb.raft.security.mode

Determines the security mode configuration.

DEFAULT, TLS, NONE

Warning

If set to TLS while one or more of the other TLS-related properties are not configured properly, the server may not be able to start.

graphdb.raft.security.certificateRevocationListFile

Determines a list of digital certificates that have been revoked by the issuing certificate authority (CA) before their actual or assigned expiration date.

graphdb.raft.security.certificateVerification

Determines the settings for certificate verification during security communication.

required, optional, optionalNoCA

graphdb.raft.security.certificateChainFile

Specifies the file path or location where the SSL/TLS certificate chain file is stored. The format is PEM-encoded.

graphdb.raft.security.certificateFile

Specifies the file path or location where the SSL/TLS certificate file is stored. The format is PEM-encoded.

graphdb.raft.security.certificateKeyFile

Specifies the file path or location where the private key associated with the SSL/TLS certificate is stored. The format is PEM-encoded.

graphdb.raft.security.certificateKeyPassword

Specifies the password or passphrase used to protect the private key associated with the SSL/TLS certificate.

graphdb.raft.security.algorithm

Specifies a particular algorithm or configuration parameter within the JSSE framework used for securing communication within the Raft-based distributed system.

graphdb.raft.security.keyAlias

Specifies the alias of the key within the JSSE framework keystore that is used for SSL/TLS operations.

graphdb.raft.security.keystoreFile

Specifies the location of the keystore file within the JSSE framework.

graphdb.raft.security.keystorePass

Specifies the password or the passphrase of the keystore file within the JSSE framework.

graphdb.raft.security.keystoreProvider

Specifies the fully qualified class name of the provider class used to access the keystore file within the JSSE framework.

graphdb.raft.security.keystoreType

Specifies the type or format of the keystore file within the JSSE framework.

graphdb.raft.security.truststoreFile

Specifies the file path or location of the truststore file.

"javax.net.ssl.trustStore"

graphdb.raft.security.truststorePass

Determines the password or passphrase required to access the truststore file.

"javax.net.ssl.trustStorePassword"

graphdb.raft.security.truststoreProvider

Determines the provider class used to access the truststore file.

"javax.net.ssl.trustStoreProvider", "javax.net.ssl.keyStoreProvider"

graphdb.raft.security.truststoreType

Determines the type or format of the truststore file.

"javax.net.ssl.trustStoreType", "javax.net.ssl.keyStoreType", "JKS"

graphdb.raft.security.rootCerts

Specifies the root certificates that the system should trust when establishing secure connections.

GraphQL properties

The GraphQL properties control how the GraphQL mechanism operates in GraphDB. Most of these properties correspond to similar properties in Semantic Objects, which have been listed in the table below for the purposes of migrating from Semantic Objects to GraphDB.

Parameter

Description

Default value

graphdb.graphql.query.optimizations.mutationMode

Specifies the write mode to the underlying GraphDB repository.

The possible values are:

  • default: Placeholder for the application default. The default value.

  • read_write: Modifications will affect the existing data in the repository. By default, all data will be written to the default graph, but this parameter also allows writing in a custom graph passed in the mutation request. Default behavior.

  • changes: Modifications will affect the existing data in the repository. All data inserts will be done in either per-entity graphs or in a custom graph passed in the mutation request.

  • read_only: Modifications will not be possible and will always fail.

Corresponding Semantic Objects property:

sparql.optimizations.mutationMode

default

graphdb.graphql.query.optimizations.disableUnionToLateralOptimization

Controls if union optimization should be disabled. The union optimization includes how SPARQL is generated for multi-valued properties in the presence of filters in the GraphQL queries. When enabled, lateral nested selects will be used to fetch multi-valued properties instead of SPARQL UNION.

Corresponding Semantic Objects property:

sparql.optimizations.disableUnionToLateralOptimization

false

graphdb.graphql.response.compactErrorMessages

Controls how the GraphQL Responder processes errors that are returned to the client. When true, it will force the GraphQL Responder to override most of the returned messages and to return a single compacted one for each of the following cases: lack of permissions, missing required data, multiple values for single-valued properties and unexpected or bad data format. This could be overridden by adding compactErrorMessages: true in the config section of the SOML or by passing the errorsFormat: FULL argument to a GraphQL query or mutation.

Corresponding Semantic Objects property:

graphql.response.compactErrorMessages

false

graphdb.graphql.response.json.nullArrays

Controls how multi-valued properties without values are represented in the JSON response. When set to true, a null will be returned instead of an empty array []. The effect of this is that properties defined as nonNullable: true (represented as [Type]! or [Type!]!) would destroy the parent if no values are present or if the non-nullable property is null.

Corresponding Semantic Objects property:

graphql.response.json.nullArrays

false

graphdb.graphql.mutation.validation.asyncEnabled

Enables or disables asynchronous mutation validation.

Corresponding Semantic Objects property:

graphql.validation.asyncValidationEnabled

true

graphdb.graphql.mutation.validation.failureMode

Determines the behavior of mutation validators when evaluation exceptions occur.

The possible values are:

  • default

  • ignore

  • warn

  • failure

  • fail

Corresponding Semantic Objects property:

graphql.validation.validationFailureMode

Note

This setting only applies to exceptions that are generated during the evaluation process and not to errors produced during regular validator operations.

default

graphdb.graphql.mutation.validation.asyncTimeoutSeconds

Configures the maximum duration that an asynchronous validation should wait to complete in seconds. When this time is reached the behavior will be based on the configured graphdb.graphql.mutation.validation.failureMode.

Corresponding Semantic Objects property:

graphql.validation.asyncValidationTimeoutSeconds

60

graphdb.graphql.mutation.validation.maxConcurrentPerRequest

Configures how many concurrent requests could be send to the database during concurrent mutation validation. A request with multiple mutations will share this limit.

Corresponding Semantic Objects property:

graphql.validation.maxConcurrentValidationsPerRequest

4

graphdb.graphql.mutation.generation.enabled

Enables or disables the generation functionality.

Corresponding Semantic Objects property:

graphql.mutation.generation.enabled

true

graphdb.graphql.introspectionQueryCache.enabled

Enables or disables introspection query caching. When set to true, introspection queries will be cached until the schema is changed. The cache key building ignores the query whitespace characters, as well as any comments.

Corresponding Semantic Objects property:

graphql.introspectionQueryCache.enabled

true

graphdb.graphql.introspectionQueryCache.config

Configures the cache behavior such as maximum size, eviction policy, and concurrency. For all possible configurations, see the GuavaCache documentation and the CacheBuilderSpec.

Corresponding Semantic Objects property:

graphql.introspectionQueryCache.config

maximumSize=100,initialCapacity=2,softValues,expireAfterAccess=30m

graphdb.graphql.enableReducedSchema

Enables or disables reducing of the generated GraphQL schema as much as possible. Although this typically results in a smaller schema, it may also reduce the dynamic extensibility of the schema, such as when merging two GraphQL schemas. When set to true, the resulting GraphQL schema will exclude scalar types and their associated input types that are not used in the generated GraphQL schema.

Corresponding Semantic Objects property:

graphql.enableReducedSchema

true

graphdb.graphql.enableOutputValidations

Enables or disables output data validation. When set to false, it will be less strict and will only fail on incompatible types.

Corresponding Semantic Objects property:

graphql.enableOutputValidations

true

graphdb.graphql.query.depthLimit

Limits the maximum depth of a GraphQL query. Queries that have a depth greater than the value defined by this property will be rejected.

Corresponding Semantic Objects property:

graphql.query.depthLimit

15

graphdb.graphql.endpointsCacheConfig

Configures the cache eviction for active GraphQL endpoints in a repository. This allows up to 100 GraphQL endpoints per repository before starting to force the unloading of endpoints of memory.

Tip

This will not block the access to more endpoints. Instead, extra endpoints will be loaded and others will be evicted.

Note

  • Loading an endpoint is associated with slow query response time for the query that is causing the endpoint to load for the first time.

  • Since the configuration has softValues, some or all of the endpoints could be unloaded in case of low server memory, which could result in slow queries overall as the endpoints will be unloaded upon the processing of the request.

maximumSize=100,expireAfterAccess=4h,softValues

graphdb.graphql.sparql.federated.services.<service_id>

Declares a federated SPARQL service.

Corresponding Semantic Objects property:

sparql.federated.services.<service_id>

none

Network properties

The network properties control how the standalone application listens on a network. These properties correspond to the attributes of the embedded Tomcat Connector. Each property is composed of the prefix graphdb.connector. + the relevant Tomcat Connector attribute. SSL properties are composed slightly differently, taking the prefix graphdb.connector.ssl. + the relevant Tomcat Connector SSL attribute.

In addition to those properties, the property graphdb.connector.ssl.enabled does not correspond to a single Tomcat Connector attribute, which allows you to enable SSL easily without requiring several different properties.

The most important property is graphdb.connector.port, which defines the port to be used. The default is 7200.

In addition, the sample config file provides an example for setting up SSL.

Note

The graphdb.connector.<xxx> properties are only relevant when running GraphDB as a standalone application.

Engine properties

You can configure the GraphDB Engine through a set of properties composed of the prefix graphdb.engine. + the relevant engine property. These properties correspond to the properties that can be set when creating a repository through the Workbench or through a .ttl file.

Note

The properties defined in the config override the properties for each repository, regardless of whether you created the repository before or after setting the global value of an engine property. As such, the global override should be used only in specific cases. For normal everyday needs, set the corresponding properties when you create a repository.

Property name

Description

Default value

graphdb.engine.entity-pool-implementation

Defines the Entity Pool implementation for the whole installation. Possible values are transactional or classic.

The default value is transactional. The transactional-simple implementation is not supported anymore.

graphdb.persistent.parallel.inferencers

Since GraphDB 8.6.1, inferencers for our Parallel loader are shut down at the end of each transaction to minimize GraphDB’s memory footprint. For cases where a lot of small insertions are done in a quick succession that can be a problem, as inferencer initialization times can be fairly slow. This setting reverts to the old behavior where inferencers are only shut down when the repository is released.

false

graphdb.engine.entity.validate

A global setting that ensures IRI validation in the entity pool. It is performed only when an IRI is seen for the first time (i.e., when being created in the entity pool). For consistency reasons, not only IRIs coming from RDF serializations, but also all new IRIs (via API or SPARQL), will be validated in the same way. This property can be turned off by setting its value to false.

true

graphdb.health.minimal.free.storage.warn

Defines the percentage of free disk space that triggers a warning that GraphDB is running low on disk space.

20

graphdb.health.minimal.free.storage.error

Defines the percentage of free disk space that triggers an error that GraphDB is running low on disk space.

10

graphdb.health.minimal.free.storage.fatal

Defines the percentage of free disk space that causes GraphDB to prevents new transactions from being started and log a fatal state error.

5

graphdb.health.minimal.free.storage.enabled

Enables or disables disk space health checks when the repository is initiated and before a transaction begins.

true

graphdb.health.minimal.free.storage.asyncCheck

Enables or disables asynchronous disk space health checks during writing every 5 seconds by default — controlled by graphdb.health.minimal.free.storage.asyncCheck.intervalMS — and additionally makes checks on every 10,000 statements by default, controlled by graphdb.health.continues.disk.check.operations.

true

graphdb.health.minimal.free.storage.asyncCheck.intervalMS

Controls the number of seconds a disk space health check is triggered if graphdb.health.minimal.free.storage.asyncCheck is set to true.

5

graphdb.health.continues.disk.check.operations

Controls the number of added statements that trigger disk space health checks if graphdb.health.minimal.free.storage.asyncCheck is set to true.

10000

Warning

To disable the low disk space health check mechanism, set both graphdb.health.minimal.free.storage.enabled and graphdb.health.minimal.free.storage.asyncCheck to false. The Low disk space health checks page provides more information on the validation checks and the different log messages.

Note

Note that IRI validation makes the import of broken data more problematic - in such a case, you would have to change a config property and restart your GraphDB instance instead of changing the setting per import.

OpenAI properties

The OpenAI properties are a set of properties used to configure the use of OpenAI GPT models. Except for the graphdb.openai.api-key setting, all of these settings are either optional or have a default setting.

Property name

Description

Default value

graphdb.openai.api-key

The authentication token for the OpenAI API. To obtain an authentication token, first create an account by clicking Sign up on the OpenAI home page. Then, you can create a key on their API Keys page.

graphdb.openai.url

The base URL for the OpenAI functionality. This setting can be used to connect to another compatible provider such as Azure OpenAI. An Azure OpenAI endpoint URL will follow this model: https://<some-id>.openai.azure.com/

graphdb.openai.auth

The authentication method to use for the chat completions endpoint. Optional. This property does not need to be configured for Azure OpenAI as it is detected automatically. The possible values are:

  • auto: Automatically determines the authentication method.

  • bearer: Sends the token via the HTTP header “Authorization: Bearer <token>”.

  • api-key: Used only with Azure OpenAI. Sends the token via the HTTP header api-key: <token>.

  • custom: Sends a token that must consist of a header and value separated by a colon (for example, my-header:my-auth-value). GraphDB will send it as the HTTP header “my-header: my-auth-value”.

  • none: No authentication headers will be sent.

auto

graphdb.openai.project

Configures the OpenAI project via the OpenAI-Project header to specify which project is used.

graphdb.openai.organization

Configures the OpenAI organization via the OpenAI-Project header to specify which organization is used.

Talk to Your Graph properties

The Talk to Your Graph properties are a set of properties used to configure the Talk to Your Graph mechanism and its agents. These properties begin with the prefix graphdb.ttyg.

Property name

Description

graphdb.ttyg.installation.id

Installation ID for the GraphDB instance. This property is relevant only when using the same OpenAI project or OpenAI Azure deployment across multiple GraphDB installations, and can be set to a unique value to prevent the different GraphDB instances from seeing each other’s agents. The default value is __default__ and allows agents and chats created with that value to be seen by the same GraphDB even if you change it to a custom value at a later point.

Default value: __default__

graphdb.ttyg.timeout

The timeout in seconds when calling the OpenAI API for Talk to Your Graph only.

Default value: 90

graphdb.ttyg.base.instructions

The default base instructions for the agents. Changing the default value is not recommended unless with a very specific intent.

Default value:

You are Quadro, an assistant developed by Ontotext and you can answer questions about data stored in GraphDB.

graphdb.ttyg.sparql.instructions

The default SPARQL queries instructions for the agents. Used only if SPARQL search is enabled.

Default value:

If you need to write a SPARQL query, use only the classes and properties provided in the schema and don’t invent or guess any. Always try to return human-readable names or labels and not only the IRIs. If SPARQL fails to provide the necessary information you can try another tool too.

graphdb.ttyg.ontology.introduction

Default ontology introduction instructions for the agents. Used only if SPARQL search is enabled.

Default value:

The ontology schema to use in SPARQL queries is:

graphdb.ttyg.http.timeout.connect

HTTP connect timeout in milliseconds to use in HTTP requests to external services. Currently, only ChatGPT Retrieval uses HTTP connections.

Default value: 30000

graphdb.ttyg.http.timeout.socket

HTTP socket timeout in milliseconds to use in HTTP requests to external services. Currently, only ChatGPT Retrieval uses HTTP connections.

Default value: 30000

graphdb.ttyg.chat.messages.limit

The max number of messages shown for each chat.

Default value: 100

graphdb.ttyg.fts.query.template

The SPARQL template for the full-text search query method.

Default value:

PREFIX onto: <http://www.ontotext.com/>
DESCRIBE ?iri {
   ?x onto:fts (%s '*') {
       ?x ?p ?iri .
   } UNION {
       ?iri ?p ?x .
   }
}

Where %s is replaced by the value generated from the model.

graphdb.ttyg.similarity.query.template

The SPARQL template for the semantic similarity search query method. The value of the first %s is replaced by the name of the similarity index itself, as defined by the user. The value of the second %s is replaced by the value generated from the model. The value of %f is replaced by the similarity score threshold, as defined by the user.

The value of this property applies to the whole project not only for the particular agent. When introducing changes to the template, make sure to preserve the order of the name of the similarity index itself (the first %s), and the value generated from the model (second %s). Failing to do so will break the query method. We also recommend against changing the three placeholders manually.

Default value:

PREFIX sim: <http://www.ontotext.com/graphdb/similarity/>
PREFIX sim-index: <http://www.ontotext.com/graphdb/similarity/instance/>
DESCRIBE ?documentID WHERE {
    ?search a sim-index:%s ;
            sim:searchTerm %s ;
            sim:documentResult ?result .
    ?result sim:value ?documentID ;
            sim:score ?score.
    FILTER(?score >= %f)
}

graphdb.ttyg.fts_iri.query.template

The SPARQL template for the full-text search for IRI discovery query method.

Default value:

PREFIX rdfs: <http://www.w3.org/2000/01/rdf-schema#>
PREFIX skos: <http://www.w3.org/2004/02/skos/core#>
PREFIX onto: <http://www.ontotext.com/>
SELECT ?label ?iri {
    ?label onto:fts (%s '*') .
    ?iri rdfs:label|skos:prefLabel ?label .
}

Where %s is replaced by the value generated from the model.

SPARQL functions and SPARQL explain

A set of properties that begin with the prefix graphdb.gpt-sparql and are used to configure the use of the OpenAI SPARQL functions and SPARQL explain extension.

Property name

Description

Default value

graphdb.gpt-sparql.model

Configures the model or Azure OpenAI model deployment ID to use for SPARQL functions and SPARQL explain. When used for Azure OpenAI, it is used to determine the model deployment ID for the full URL to the Chat Completions API endpoint.

gpt-4o

graphdb.gpt-sparql.api-version

Configures the API version for Azure OpenAI. The default value is 2024-06-01 and it may change in the future. Unless you need a specific API version, it’s recommended to leave this property unset.

2024-06-01

graphdb.gpt-sparql.timeout

Configures the timeout in seconds when calling the OpenAI API for SPARQL functions and SPARQL explain only.

90

Configuring logging

GraphDB uses logback to configure logging. The default configuration is provided as logback.xml in the GraphDB conf directory.