Loading data

What’s in this document?

GraphDB offers multiple interfaces for loading RDF data in your repository. Each interface supports the loading of any RDF format, but the different interfaces have different limitations and advantages.

You can access these interfaces through Workbench or through the Import RDF tool. You can also load RDF data via HTTP with curl.

GraphDB’s data loading interfaces

Interface

Use cases

Mode

Speed

Workbench import of a local RDF file

Small files, limited to 1GB by default

Online parallel

Moderate speed

Workbench import of a remote RDF file

Small files, limited to 1GB by default

Online parallel

Moderate speed

Workbench import of a text snippet

Small text snippets

Online parallel

Moderate speed

Workbench import of a server file

No limits on the file size

Online parallel

Fast, ignoring all HTTP protocol overheads

SPARQL endpoint

No limits on the file size

Online parallel

Moderate speed

ImportRDF Load

Batch import of very big files

Initial offline import with no plugins

Fast, with small speed degradation

ImportRDF Preload

Import huge datasets with no inference

Initial offline import with no inference and plugins

Ultra-fast without speed degradation

Tip

It’s often useful to use GraphDB’s rdfvalidator command line utility to check that an RDF file parses properly before attempting to load it.

Updating data in GraphDB is done via smart updates using server-side SPARQL templates.

GraphDB supports SHACL validation ensuring efficient data consistency checking.

The GraphDB sequences plugin provides transactional sequences for GraphDB. A sequence is a long counter that can be atomically incremented in a transaction to provide incremental IDs.

Loading via HTTP with curl

Using curl lets you script this call in an application. See also the Help ‣ REST API view of the GraphDB Workbench where you will find a complete reference of all REST APIs and be able to run API calls directly from the browser.

In addition to this, the RDF4J API is also available.

Most data import queries can either take the following set of attributes as an argument or return them as a response.

  • fileNames (string list): A list of the files to import.

  • importSettings (JSON object): Import settings.

    • baseURI (string): Base URI for the files to be imported.

    • context (string): Context for the files to be imported.

    • data (string): Inline data.

    • forceSerial (boolean): Force use of the serial statements pipeline.

    • name (string): Filename.

    • status (string): Status of an import — pending, importing, done, error, none, or interrupting.

    • timestamp (integer): When the import was started.

    • type (string): The type of the import.

    • replaceGraphs (string list): A list of graphs that you want to be completely replaced by the import.

    • parserSettings (JSON object): Parser settings.

      • failOnUnknownDataTypes (boolean): Fail parsing if datatypes are not recognized.

      • failOnUnknownLanguageTags (boolean): Fail parsing if languages are not recognized. Valid language tags are defined in the IANA Language Subtag Registry. Since its structure (described in BCP47 sec 3.1) is not easy to read, the script named iana-lang-tags.pl gets the registry, parses it, and writes it to a tab-delimited file. The result is saved as a Google Sheet named iana-lang-tags. You can also use custom tags and subtags by using the prefix x-.

      • normalizeDataTypeValues (boolean): Normalize recognized datatype values. For example, "01"^^xsd:integer will be converted to "1"^^xsd:integer.

      • normalizeLanguageTags (boolean): Normalize recognized language tags, as described in RFC 5646, section 2.1.1.

      • preserveBNodeIds (boolean): Use blank node IDs found in the file instead of assigning them.

      • stopOnError (boolean): Stop on error. If false, the error will be logged and parsing will continue.

      • verifyDataTypeValues (boolean): Verify values of recognized datatypes and stop on error. This applies to dates (such as "2023-02-31"^^xsd:date), numbers ("foo"^^xsd:decimal), decimal digits ("3.14"^^xsd:integer), and range errors ("-1"^^xsd:positiveInteger, "1234567890123456"^^xsd:int). It does not apply to custom datatypes or datatypes derived by restriction.

      • verifyLanguageTags (boolean): Verify language based on a given set of definitions for valid languages.

      • contextLink (string): Provide context for importing Flattened and Compacted JSON-LD files or NDJSON-LD files.

Cancel server file import operation

DELETE /rest/repositories/<repo_id>/import/server

Example:

curl -X DELETE <base_url>/rest/repositories/<repo-id>/import/server?name=<encoded_filepath>

Get server files available for import

GET /rest/repositories/<repo_id>/import/server

Example:

curl <base_url>/rest/repositories/<repo_id>/import/server

Import a server file into the repository or a named graph inside the repository

POST /rest/repositories/<repo_id>/import/server

Example:

curl -X POST --header 'Content-Type: application/json' -d '{
  "fileNames": [
    "<data_url>",
    "<data_url>"
  ]
}' <base_url>/rest/repositories/<repo_id>/import/server

Tip

Common parameters:

<base_url>: The URL host and path leading to the deployed GraphDB Workbench webapp

<repo_id>: The ID string that identifies the repository to use

<encoded_filepath>: Encoded filepath leading to a server file that is in the process of being imported

With the API, you can add a file of triples to a named graph specified on the command line. It’s important to remember that RDF4J calls the fourth part of a quad (where the graph name is stored) the context.

Example:

  curl -X POST --header 'Content-Type: text/turtle' \
--data-binary @<data_url> \  '<base_url>/repositories/<repo_id>/statements?context=<named_graph>

Warning

The graph name must include <> and the whole name must be URL encoded. In the example above, <named_graph> is the name of the graph, including the angle brackets, which must also be encoded.