Loading data¶
What’s in this document?
GraphDB offers multiple interfaces for loading RDF data in your repository. Each interface supports the loading of any RDF format, but the different interfaces have different limitations and advantages.
You can access these interfaces through Workbench or through the Import RDF tool. You can also load RDF data via HTTP with curl.
Interface |
Use cases |
Mode |
Speed |
|---|---|---|---|
Small files, limited to 1GB by default |
Online parallel |
Moderate speed |
|
Small files, limited to 1GB by default |
Online parallel |
Moderate speed |
|
Small text snippets |
Online parallel |
Moderate speed |
|
No limits on the file size |
Online parallel |
Fast, ignoring all HTTP protocol overheads |
|
No limits on the file size |
Online parallel |
Moderate speed |
|
Batch import of very big files |
Initial offline import with no plugins |
Fast, with small speed degradation |
|
Import huge datasets with no inference |
Initial offline import with no inference and plugins |
Ultra-fast without speed degradation |
Tip
It’s often useful to use GraphDB’s rdfvalidator command line utility to check that an RDF file parses properly before attempting to load it.
Updating data in GraphDB is done via smart updates using server-side SPARQL templates.
GraphDB supports SHACL validation ensuring efficient data consistency checking.
The GraphDB sequences plugin provides transactional sequences for GraphDB. A sequence is a long counter that can be atomically incremented in a transaction to provide incremental IDs.
Loading via HTTP with curl¶
Using curl lets you script this call in an application. See also the view of the GraphDB Workbench where you will find a complete reference of all REST APIs and be able to run API calls directly from the browser.
In addition to this, the RDF4J API is also available.
Most data import queries can either take the following set of attributes as an argument or return them as a response.
fileNames(string list): A list of the files to import.importSettings(JSON object): Import settings.baseURI(string): Base URI for the files to be imported.context(string): Context for the files to be imported.data(string): Inline data.forceSerial(boolean): Force use of the serial statements pipeline.name(string): Filename.status(string): Status of an import — pending, importing, done, error, none, or interrupting.timestamp(integer): When the import was started.type(string): The type of the import.replaceGraphs(string list): A list of graphs that you want to be completely replaced by the import.parserSettings(JSON object): Parser settings.failOnUnknownDataTypes(boolean): Fail parsing if datatypes are not recognized.failOnUnknownLanguageTags(boolean): Fail parsing if languages are not recognized. Valid language tags are defined in the IANA Language Subtag Registry. Since its structure (described in BCP47 sec 3.1) is not easy to read, the script named iana-lang-tags.pl gets the registry, parses it, and writes it to a tab-delimited file. The result is saved as a Google Sheet namediana-lang-tags. You can also use custom tags and subtags by using the prefixx-.normalizeDataTypeValues(boolean): Normalize recognized datatype values. For example,"01"^^xsd:integerwill be converted to"1"^^xsd:integer.normalizeLanguageTags(boolean): Normalize recognized language tags, as described in RFC 5646, section 2.1.1.preserveBNodeIds(boolean): Use blank node IDs found in the file instead of assigning them.stopOnError(boolean): Stop on error. Iffalse, the error will be logged and parsing will continue.verifyDataTypeValues(boolean): Verify values of recognized datatypes and stop on error. This applies to dates (such as"2023-02-31"^^xsd:date), numbers ("foo"^^xsd:decimal), decimal digits ("3.14"^^xsd:integer), and range errors ("-1"^^xsd:positiveInteger,"1234567890123456"^^xsd:int). It does not apply to custom datatypes or datatypes derived by restriction.verifyLanguageTags(boolean): Verify language based on a given set of definitions for valid languages.contextLink(string): Provide context for importing Flattened and Compacted JSON-LD files or NDJSON-LD files.
Cancel server file import operation
DELETE /rest/repositories/<repo_id>/import/server
Example:
curl -X DELETE <base_url>/rest/repositories/<repo-id>/import/server?name=<encoded_filepath>
Get server files available for import
GET /rest/repositories/<repo_id>/import/server
Example:
curl <base_url>/rest/repositories/<repo_id>/import/server
Import a server file into the repository or a named graph inside the repository
POST /rest/repositories/<repo_id>/import/server
Example:
curl -X POST --header 'Content-Type: application/json' -d '{
"fileNames": [
"<data_url>",
"<data_url>"
]
}' <base_url>/rest/repositories/<repo_id>/import/server
Tip
Common parameters:
<base_url>: The URL host and path leading to the deployed GraphDB Workbench webapp
<repo_id>: The ID string that identifies the repository to use
<encoded_filepath>: Encoded filepath leading to a server file that is in the process of being imported
With the API, you can add a file of triples to a named graph specified on the command line. It’s important to remember that RDF4J calls the fourth part of a quad (where the graph name is stored) the context.
Example:
curl -X POST --header 'Content-Type: text/turtle' \
--data-binary @<data_url> \ '<base_url>/repositories/<repo_id>/statements?context=<named_graph>
Warning
The graph name must include <> and the whole name must be URL encoded. In the example above, <named_graph> is the name of the graph, including the angle brackets, which must also be encoded.