Skip to content

Changelog

Version Compatibility

The JanusGraph project is growing along with the rest of the graph and big data ecosystem and utilized storage and indexing backends. Below are version compatibilities between the various versions of components. For dependent backend systems, different minor versions are typically supported as well. It is strongly encouraged to verify version compatibility prior to deploying JanusGraph.

Although JanusGraph may be compatible with older and no longer supported versions of its dependencies, users are warned that there are possible risks and security exposures with running software that is no longer supported or updated. Please check with the software providers to understand their supported versions. Users are strongly encouraged to use the latest versions of the software.

Version Compatibility Matrix

Currently supported

All currently supported versions of JanusGraph are listed below.

Info

You are currently viewing the documentation page of JanusGraph version 1.1.0. To ensure that the information below is up to date, please double check that this is not an archived version of the documentation.

JanusGraph Storage Version Cassandra HBase Bigtable ScyllaDB Elasticsearch Solr TinkerPop Spark Scala
1.2.z 2 3.11.z, 4.0.z, 5.0.z 2.6.z 1.3.0, 1.4.0, 1.5.z, 1.6.z, 1.7.z, 1.8.z, 1.9.z, 1.10.z, 1.11.z, 1.14.z 6.y 6.1-6.8.z, 7.y, 8.y, 9.y, OpenSearch 2.y, 3.y 8.11.z, 9.y 3.8.z 3.2.z 2.12.z
1.1.z 2 3.11.z, 4.0.z 2.6.z 1.3.0, 1.4.0, 1.5.z, 1.6.z, 1.7.z, 1.8.z, 1.9.z, 1.10.z, 1.11.z, 1.14.z 6.y 6.y, 7.y, 8.y 8.y 3.7.z 3.2.z 2.12.z

Info

Even so ScyllaDB is marked as N/A prior version 1.0.0 it was actually supported using cql storage option. The only difference is that from version 1.0.0 JanusGraph officially supports ScyllaDB using scylla and cql storage options and have extended test coverage for ScyllaDB.

End-of-Life

The versions of JanusGraph listed below are outdated and will no longer receive bugfixes.

JanusGraph Storage Version Cassandra HBase Bigtable ScyllaDB Elasticsearch Solr TinkerPop Spark Scala
0.1.z 1 1.2.z, 2.0.z, 2.1.z 0.98.z, 1.0.z, 1.1.z, 1.2.z 0.9.z, 1.0.0-preZ, 1.0.0 N/A 1.5.z 5.2.z 3.2.z 1.6.z 2.10.z
0.2.z 1 1.2.z, 2.0.z, 2.1.z, 2.2.z, 3.0.z, 3.11.z 0.98.z, 1.0.z, 1.1.z, 1.2.z, 1.3.z 0.9.z, 1.0.0-preZ, 1.0.0 N/A 1.5-1.7.z, 2.3-2.4.z, 5.y, 6.y 5.2-5.5.z, 6.2-6.6.z, 7.y 3.2.z 1.6.z 2.10.z
0.3.z 2 1.2.z, 2.0.z, 2.1.z, 2.2.z, 3.0.z, 3.11.z 1.0.z, 1.1.z, 1.2.z, 1.3.z, 1.4.z 1.0.0, 1.1.0, 1.1.2, 1.2.0, 1.3.0, 1.4.0 N/A 1.5-1.7.z, 2.3-2.4.z, 5.y, 6.y 5.2-5.5.z, 6.2-6.6.z, 7.y 3.3.z 2.2.z 2.11.z
0.4.z 2 2.1.z, 2.2.z, 3.0.z, 3.11.z 1.2.z, 1.3.z, 1.4.z, 2.1.z N/A N/A 5.y, 6.y 7.y 3.4.z 2.2.z 2.11.z
0.5.z 2 2.1.z, 2.2.z, 3.0.z, 3.11.z 1.2.z, 1.3.z, 1.4.z, 2.1.z 1.3.0, 1.4.0, 1.5.z, 1.6.z, 1.7.z, 1.8.z, 1.9.z, 1.10.z, 1.11.z, 1.14.z N/A 6.y, 7.y 7.y 3.4.z 2.2.z 2.11.z
0.6.z 2 3.0.z, 3.11.z 1.6.z, 2.2.z 1.3.0, 1.4.0, 1.5.z, 1.6.z, 1.7.z, 1.8.z, 1.9.z, 1.10.z, 1.11.z, 1.14.z N/A 6.y, 7.y 7.y, 8.y 3.5.z 3.0.z 2.12.z
1.0.z 2 3.11.z, 4.0.z 2.5.z 1.3.0, 1.4.0, 1.5.z, 1.6.z, 1.7.z, 1.8.z, 1.9.z, 1.10.z, 1.11.z, 1.14.z 5.y 6.y, 7.y, 8.y 8.y 3.7.z 3.2.z 2.12.z

Release Notes

Version 1.2.0 (Release Date: ???)

<dependency>
    <groupId>org.janusgraph</groupId>
    <artifactId>janusgraph-core</artifactId>
    <version>1.2.0</version>
</dependency>
compile "org.janusgraph:janusgraph-core:1.2.0"

Tested Compatibility:

  • Apache Cassandra 3.11.10, 4.0.6, 5.0.8
  • Apache HBase 2.6.0
  • Oracle BerkeleyJE 7.5.11
  • ScyllaDB 6.2.0
  • Elasticsearch 6.6.0, 7.17.8, 8.15.3, 9.5.4
  • OpenSearch 2.19.6, 3.9.0
  • Apache Lucene 9.12.3
  • Apache Solr 8.11.4, 9.10.1
  • Apache TinkerPop 3.8.2
  • Java 11, 17, 21, 25 (OLAP with Apache Spark: Java 11 and 17 only)

Installed versions in the Pre-Packaged Distribution:

  • Cassandra 5.0.9
  • Elasticsearch 7.17.29

Changes

For more information on features and bug fixes in 1.2.0, see the GitHub milestone:

Assets

Upgrade Instructions

Java 11 is now the minimum supported Java version, Java 17, 21 and 25 are supported

Starting from version 1.2.0 JanusGraph requires Java 11 or newer. Support for Java 8 has been dropped for building and running JanusGraph, JanusGraph Server, the pre-packaged distribution and the Gremlin Console. All JanusGraph artifacts are now compiled for Java 11, so applications embedding JanusGraph must run on Java 11 or newer as well. This change is required by the upgrade to Apache TinkerPop 3.8, which itself requires Java 11.

JanusGraph is now built and tested with Java 11, 17, 21 and 25. The exceptions are the OLAP modules (janusgraph-hadoop and the Hadoop/Spark based graph computer of the storage backends): Apache Spark 3.3.x, which TinkerPop's spark-gremlin is based on, only runs on Java 8 through 17, so OLAP jobs are only supported on Java 11 and 17. The pre-packaged distribution bundles Cassandra 5.0.9, which itself only runs on Java 11 and 17, so the embedded Cassandra of the janusgraph-full distribution requires Java 11 or 17 while an external Cassandra cluster can be used from JanusGraph running on any supported Java version.

Java 17 and newer enforce strong encapsulation of the JDK internals. JanusGraph itself does not need any --add-opens / --add-exports options, but some libraries which can be used with JanusGraph still rely on deep reflection (most notably the Kryo serialization used by Gryo and OLAP), in which case the options documented by TinkerPop for JDK 17 have to be passed to the JVM (JanusGraph's test suites pass them on JDK 17 and newer, see the jdk17-plus-tests profile of the root pom.xml). The Hadoop client artifacts pulled in by Spark (hadoop-client-api, hadoop-client-runtime) are aligned with the other Hadoop 3.4.3 artifacts, because their 3.3.x versions still call Subject.getSubject(), which does not work on Java 24 and newer (this affected the HBase client's user resolution).

Since the default build now targets Java 11, the separate Java 11 build variant has been removed:

  • The -Pjava-11 Maven profile no longer exists. Build JanusGraph with mvn clean install on Java 11+.
  • The janusgraph-java-11-<version>.zip and janusgraph-java-11-full-<version>.zip distribution archives are no longer produced. Use janusgraph-<version>.zip and janusgraph-full-<version>.zip instead; they are built for Java 11.
  • The -java-11 Docker image tag suffix produced by that build variant is gone. The regular janusgraph/janusgraph:<version> images are the only ones built; they already use a Java 11 runtime (eclipse-temurin:11-jre).
  • conf/jvm-8.options has been removed from the distribution. bin/janusgraph-server.sh now always reads conf/jvm-11.options unless JAVA_OPTIONS_FILE is set.
Upgrade to Apache TinkerPop 3.8.2

JanusGraph 1.2.0 upgrades Apache TinkerPop from 3.7.3 to 3.8.2. TinkerPop 3.8 is a new minor release line with a number of breaking changes to Gremlin semantics that may affect existing traversals and applications. The most notable ones are:

  • store() was removed in favor of local(aggregate()), aggregate(Scope, String) was removed and has(key, traversal) / has(T, traversal) were removed (use where() instead).
  • none() was renamed to discard(); none(P) is now a collection filtering step complementing any(P) and all(P).
  • P.getOriginalValue() was removed in favor of P.getValue().
  • java.time.OffsetDateTime replaces java.util.Date as the default date type: asDate(), dateAdd() and dateDiff() return OffsetDateTime, and dateDiff() returns milliseconds instead of seconds.
  • The repeat traversal of repeat() now consistently uses global semantics: if it contains a barrier step (for example barrier(), order() or aggregate()), all traversers of a loop enter the repeat traversal at once instead of one at a time, which changes the ordering of results and lets barrier steps see all traversers of the loop. JanusGraph's own multi-query batching steps inserted into repeat traversals do not trigger this mode, so a repeat() without user-defined barriers keeps processing its traversers in bounded batches. RepeatUnrollStrategy only unrolls simple navigation and filter steps, and cap() / inject() inside repeat() are rejected by StandardVerificationStrategy.
  • valueMap(), propertyMap(), groupCount(), sack(), dedup(), sample() and aggregate() reject more than one by() modulator.
  • property(key, value) without an explicit cardinality now always resolves the cardinality through Graph.Features.VertexFeatures.getCardinality(key). JanusGraph resolves it against the schema of the calling transaction (including property keys created but not yet committed in that transaction) and falls back to the default cardinality of the configured schema maker for unknown keys.
  • Arithmetic in sum() and sack() promotes to the next wider numeric type on overflow, split() with an empty separator splits a string into its characters, and floating-point literals in gremlin-lang scripts are parsed as Double instead of BigDecimal.
  • GraphSON 2.0 and 3.0 only deserialize TraversalStrategy implementations that are registered with TraversalStrategies.GlobalCache (all JanusGraph strategies are registered).
  • GraphSON 1.0 with embedded types (GraphSONMessageSerializerV1, TypeInfo.PARTIAL_TYPES) only deserializes explicitly allowed @class type ids. GraphSON 1.0 mappers created through JanusGraph.io() allow JanusGraph's own types automatically. A manually configured GraphSONMessageSerializerV1 needs allowedTypeIdNames: [org.janusgraph.graphdb.relations.RelationIdentifier, org.janusgraph.core.attribute.Geoshape, org.janusgraph.graphdb.tinkerpop.io.JanusGraphP] next to ioRegistries (see the commented examples in conf/gremlin-server/*.yaml), and a hand-built GraphSONMapper needs JanusGraphIoRegistryV1d0.allowGraphSONTypeIds(builder) (or addAllowedTypeIdName(...) with the names of JanusGraphIoRegistryV1d0.GRAPHSON_ALLOWED_TYPE_ID_NAMES).
  • The UnifiedChannelizer of Gremlin Server has been deprecated.

Please review the TinkerPop upgrade documentation for 3.8.0, 3.8.1 and 3.8.2 before upgrading.

Updated third-party libraries

JanusGraph 1.2.0 updates its third-party libraries to their latest versions which are compatible with Java 11 and with the libraries JanusGraph builds on (TinkerPop 3.8, Spark 3.3 and Hadoop 3.4). Applications embedding JanusGraph pick these versions up transitively. The most notable updates are:

  • Google Cloud Bigtable HBase client 2.20.2 (from 1.24.0) in janusgraph-bigtable. The connection settings described in the Bigtable documentation are unchanged.
  • janusgraph-cdc (new in 1.2.0) uses the Apache Kafka 4.3.1 clients, which require Kafka brokers 2.1 or newer.
  • Apache HBase client 2.6.6 (from 2.6.0), Apache ZooKeeper 3.9.6 (from 3.9.2), gRPC 1.84.0 (from 1.66.0), Guava 33.7.1 (from 33.3.0), Jackson 2.22 (from 2.17), Log4j 2.26.1 (from 2.23.1), HPPC 0.10.0 (from 0.9.1) and Vavr 1.0.1 (from 0.10.4).
Apache Cassandra 5.0 support

Starting from version 1.2.0 JanusGraph supports Apache Cassandra 5.0 as a storage backend. The pre-packaged janusgraph-full distribution now bundles Cassandra 5.0.9 instead of Cassandra 4.0.6. The embedded Cassandra runs on Java 11 and 17 (Cassandra 4.0 only ran on Java 8 and 11). Its cassandra/conf/cassandra.yaml is regenerated from the stock Cassandra 5.0 configuration with the JanusGraph settings (cluster name, db/cassandra data directories) applied; num_tokens stays at 256 so data directories created by earlier janusgraph-full distributions keep working. Review the Cassandra 5.0 upgrade notes before reusing an existing db/cassandra directory with the new distribution.

ElasticSearch 9 support

Starting from version 1.2.0 JanusGraph supports ElasticSearch 9.

The pre-packaged janusgraph-full distribution now bundles Elasticsearch 7.17.29 instead of Elasticsearch 7.17.8. Elasticsearch 7.17 is the last Elasticsearch line whose JVM can be Java 11: the bundled Elasticsearch runs on the JDK shipped in its Linux x86_64 tarball or, where that JDK cannot run (for example macOS or Linux on ARM), on the JDK of the host through JAVA_HOME. Elasticsearch 8 requires Java 17 and Elasticsearch 9 requires Java 21 for its JVM, and since 9.4 its launcher and native libraries are x86_64 only, so bundling them would have made the embedded Elasticsearch Linux x86_64 only. The native machine learning binaries are no longer part of the distribution (machine learning is disabled in the bundled elasticsearch.yml).

OpenSearch 2 and 3 support

Starting from version 1.2.0 JanusGraph supports OpenSearch 2 and 3 with the elasticsearch index backend. OpenSearch provides the Elasticsearch 7 API, which JanusGraph now uses when the cluster reports an OpenSearch version. Before, JanusGraph rejected OpenSearch versions as unsupported Elasticsearch versions, so OpenSearch 2 only worked with compatibility.override_main_response_version, and OpenSearch 3, which removed that setting, didn't work at all. JanusGraph rejects other OpenSearch versions, including OpenSearch 1, which reached its end of life, unless the option below is set. When JanusGraph detects OpenSearch, it ignores index.[X].elasticsearch.use-mapping-for-es7, because OpenSearch 2 removed mapping types.

The new option index.[X].elasticsearch.major-version sets the major version of the Elasticsearch API which the cluster provides (7 for OpenSearch). If it is set, JanusGraph doesn't ask the cluster for its version, and so doesn't know whether the cluster is OpenSearch. See OpenSearch for details, including the settings Amazon OpenSearch Service needs.

Elasticsearch 6.0 is no longer supported

JanusGraph 1.2.0 no longer supports Elasticsearch 6.0, and its tests no longer run against it. Elasticsearch 6.0 reached its end of life in 2019, and unlike later versions it does not run the ingest pipeline of a bulk request on the upsert of an update: a document which a mutation creates through an upsert, such as one an addition recreates because it is missing from the index, skips the pipeline set with index.[X].elasticsearch.ingest-pipeline.[mixedIndexName]. Upgrade Elasticsearch 6.0 clusters before upgrading JanusGraph; the oldest Elasticsearch 6 release tested is 6.6.0.

Solr 9 support, Solr 8 is deprecated

JanusGraph 1.2.0 uses SolrJ 9.10.1 and supports Solr 9. Solr 8.11 is still supported and tested, but deprecated: Solr 8 reached its end of life, and a future JanusGraph version will drop it. SolrJ 10 requires Java 17 (and Solr 10 Java 21), which is newer than JanusGraph's minimum Java version, so Solr 10 isn't supported yet.

JanusGraph can be upgraded before the Solr cluster. The configset in conf/solr now works with Solr 8.11 and 9.

Solr 9 removed LatLonType and the LRUCache and FastLRUCache caches, which the configset in conf/solr of earlier JanusGraph distributions used, so Solr 9 can't load the collections created with that configset. Before upgrading the Solr cluster to Solr 9, change the configset of these collections and upload it again (bin/solr zk upconfig; in the HTTP mode, change the files in the conf directory of each core):

  • Replace the location field type (JanusGraph doesn't use it) with <fieldType name="location" class="solr.LatLonPointSpatialField" docValues="true"/>. Changing only the class isn't enough, because LatLonPointSpatialField rejects the subFieldSuffix attribute of LatLonType.
  • Change the class of the caches (filterCache, queryResultCache, documentCache and perSegFilter) to solr.CaffeineCache.

The data of these collections stays readable, because Solr 9 still supports the Trie field types of that configset, although it deprecates them. Solr 9 can't open indexes which were created by Solr 7 or older.

The configset in conf/solr now uses the Point field types (with doc values) instead of the Trie field types, LatLonPointSpatialField, CurrencyFieldType instead of CurrencyField, CaffeineCache and luceneMatchVersion 9.12. Use it for new collections. An existing collection can only switch to it by recreating the collection and reindexing the mixed index (SchemaAction.REINDEX), because the field types of existing data can't change.

Solr 9 removed the maxShardsPerNode parameter of the collection creation, and SolrJ 9 no longer offers it. JanusGraph still sends index.[X].solr.max-shards-per-node when it creates a collection on Solr 8, which puts at most 1 replica of a new collection on a node by default. Before JanusGraph creates a collection, it asks Solr for its version. The new option index.[X].solr.major-version (for example 8 or 9) takes precedence over the version Solr reports and saves that request. If neither is known, JanusGraph assumes Solr 8, because Solr 9 ignores the parameter.

index.[X].solr.max-shards-per-node is deprecated, and JanusGraph logs a warning when it is set for Solr 9. After upgrading to Solr 9, remove this GLOBAL_OFFLINE option from the graph's configuration (for example with mgmt.remove("index.search.solr.max-shards-per-node") and mgmt.commit() while only one JanusGraph instance is open) and from the local configuration files.

JanusGraph keeps using the SolrJ clients based on Apache HttpClient (CloudLegacySolrClient for SolrCloud, since SolrJ 9's CloudSolrClient.Builder builds the Jetty based HTTP/2 client), so the Kerberos configuration is unchanged. In the HTTP mode, index.[X].solr.http-connection-timeout (5 seconds by default) now applies to the requests; SolrJ 8's load balancing client connected with its own timeout of 15 seconds. On the Solr server, the Kerberos authentication plugin (org.apache.solr.security.hadoop.KerberosPlugin) is part of Solr 9's hadoop-auth module.

SolrJ 9 is built on Jetty 10, so JanusGraph now manages Jetty 10.0.26 instead of 9.4.58. Applications which embed JanusGraph and use Jetty 9.4 themselves have to align their Jetty version.

The JanusGraph distribution no longer contains noggit-0.8.jar. SolrJ contains its own, newer copy of the org.noggit classes, and when the org.noggit:noggit 0.8 artifact (a dependency of janusgraph-driver for Spatial4j's GeoJSON reader) comes first on the classpath, SolrJ fails with NoSuchMethodError: 'java.lang.Object org.noggit.ObjectBuilder.getValStrict()'. Applications which use janusgraph-solr should exclude org.noggit:noggit as well: Spatial4j works with SolrJ's copy.

Apache Lucene 9

janusgraph-lucene now uses Apache Lucene 9.12.3 instead of 8.11. A JanusGraph installation can contain only one Lucene version, and Solr 9 is built on Lucene 9 (Lucene 10 requires Java 21).

Lucene 9 opens indexes which were created by Lucene 8 (JanusGraph 0.6.0 to 1.1.x), so existing Lucene mixed indexes keep working without a reindex. Once JanusGraph 1.2.0 has written to such an index, earlier JanusGraph versions can't open it anymore. Lucene 9 can't open indexes created by Lucene 7 or older (JanusGraph 0.5.x and older), even if they were written by Lucene 8 later. Delete the directories of such indexes and reindex the mixed indexes.

Minimum and maximum aggregations of Float properties which the Lucene index computes (for example g.V().has("name", "bob").values("weight").max()) now return the right values: they used to read the indexed double values as floats.

Custom analyzers (the string-analyzer and text-analyzer mapping parameters of the Lucene and Solr indexes) are loaded by class name, so they have to exist in Lucene 9. Lucene 9 renamed the lucene-analyzers-common artifact to lucene-analysis-common and moved a few analyzers to other packages, for example ClassicAnalyzer to org.apache.lucene.analysis.classic and UAX29URLEmailAnalyzer to org.apache.lucene.analysis.email.

Zombie instances auto-close during index status update operations

Starting from version 1.2.0 JanusGraph can automatically force-close JanusGraph instances that are unreachable during index status update operations.

To enable this behavior, users can set the following configuration:

graph.management-auto-close-stale-instances=true

By default, an instance is considered stale if it is not reachable for more than 2 minutes. However, this threshold can be adjusted using the following configuration:

graph.management-ack-timeout=240000 ms
This is a breaking change for users who use the JanusGraphIndexStatusUpdate interface.

Faster mixed-index reindex with batched document restores

Mixed-index reindex jobs (SchemaAction.REINDEX against an Elasticsearch, Solr or Lucene index) now flush restored documents to the index backend in bounded batches while the storage scan is still running, instead of buffering a whole scan segment and flushing it once. This reduces the number of bulk requests, transaction commits and management-system reads per reindexed element and bounds worker memory.

This behavior is enabled by default. Two new configuration options control it:

schema.reindex.mixed-index-batch-enabled=true
schema.reindex.mixed-index-batch-size=1000
schema.reindex.mixed-index-batch-size is the number of documents a reindex worker buffers before issuing a single restore (bulk) call. A worker buffers up to this many documents in memory and a reindex runs several workers in parallel, so peak heap grows with mixed-index-batch-size × reindex-threads × average-document-size — raise it only when documents are small and workers have headroom.

Be aware of the following behavior changes when batching is enabled (the default):

  • Restored documents become visible in the index incrementally, as each batch is flushed, rather than once per scan segment. This is safe because a reindex is idempotent and rebuilds index state from the graph as the source of truth; a reindex interrupted partway leaves partially-populated index documents that a re-run deterministically completes.
  • The reindex flush cadence is now driven by mixed-index-batch-size rather than by the storage page size (storage.page-size).

To restore the previous storage-page-sized, flush-once-per-segment behavior, set schema.reindex.mixed-index-batch-enabled=false. See the Elasticsearch reindex tuning guide for tuning the batch size, reindex threads and index.[X].elasticsearch.bulk-refresh together.

Faster OLAP scans (signal-based row hand-off)

The OLAP scan pipeline behind reindex and other scan jobs now hands rows between its internal threads using blocking take() and sentinel markers instead of polling bounded queues on a fixed timer. On fast backends (notably CQL/Cassandra) this removes per-row hand-off latency that could otherwise dominate scan wall-clock time, speeding up single-node reindex and other full scans by a large factor. No configuration or user action is required.

Opt-in parallel token-range scan for CQL full scans

CQL full-table scans (used by reindex and other OLAP jobs) can optionally be split into several token-bounded queries instead of a single coordinator-funneled scan:

storage.cql.parallel-scan-token-ranges=1
The default value 1 preserves the previous single-query behavior. When set above 1 and the Murmur3 partitioner is in use, scan jobs (reindex and other jobs running through the scan-job framework) drain each token range on its own row-collection pipeline — its own data-puller threads and merge thread — so the storage scan runs fully in parallel across ranges and producer throughput scales with the range count until the cluster saturates. Non-scan-job callers of a whole-table scan receive the ranges as token-bounded queries streamed back-to-back in token order (bounded coordinator scans, without extra parallelism). Each range adds concurrent scan queries against the cluster (one per scan-job query, times the number of ranges), so very high values can overload the cluster; a small multiple of the cluster's node count is a sensible starting point. The option is ignored for non-Murmur3 partitioners.

CQL scan-only page size

Full-table scans can use their own CQL page size instead of the OLTP-oriented storage.page-size:

storage.cql.scan-page-size=0
The default 0 keeps using storage.page-size. Since a full scan streams many rows per request, a page size several times larger than the OLTP page size usually cuts scan round trips (and total scan time) substantially; a few thousand rows per page is a reasonable value. Scan pages are additionally fetched with a one-page lookahead, overlapping the network wait of the next page with client-side processing of the current one — this pipelining is always on and needs no configuration.

PER PARTITION LIMIT pushdown for CQL scans

Scan queries now push their per-key entry limit into the CQL query as PER PARTITION LIMIT (enabled by default):

storage.cql.scan-per-partition-limit-enabled=true
Scan jobs issue a grounding (key-existence) query with a per-key limit of 1; previously every cell of every row slice was streamed to the client and discarded there, so the grounding query alone transferred the whole table — including adjacency data irrelevant to the job. With the pushdown the server stops after the per-key limit (one cell per key for the grounding query), which substantially reduces scan time and network transfer on graphs with wide rows (many edges or properties per vertex). Requires PER PARTITION LIMIT support in the backend (Apache Cassandra 3.6+, ScyllaDB). A CQL-compatible service that rejects the clause (e.g. Amazon Keyspaces) automatically falls back to plain scan statements — store open logs a warning and continues instead of failing; set the option to false to skip the attempt entirely.

Scan jobs fail loudly on data-puller errors

Previously, when an internal scan data-puller thread died on a storage error, the scan completed "successfully" with silently missing rows — for a reindex this could ENABLE an incomplete index. A scan job now fails with a TemporaryBackendException describing the failed puller instead of returning a partial result, and it fails fast: the error surfaces as soon as the dead puller's end-of-data marker is observed rather than after the remaining key space has been streamed and discarded.

Scan merge no longer loses data on writes concurrent with the scan

The multi-query scan merge matches each secondary slice query's rows against the grounding (key-existence) stream. A key written while the scan runs can appear only in a secondary stream — the grounding puller had already passed its position — and such a row can never match. Previously it permanently occupied the merge's single pending slot for that query, so every later key was silently merged with empty results for the query: a reindex on a live graph produced documents missing that query's data for the rest of the scan. On backends whose scans iterate keys in their natural order the merge now classifies every row directly against the grounding key (a merge join), so stale rows are dropped outright; on token-ordered backends (such as CQL with the Murmur3 partitioner), where that order is not computable client-side, a bounded per-query buffer of unmatched rows recovers by dropping the rows the grounding stream has provably passed. Keys written during the scan are unaffected either way: they are indexed by normal live-write index maintenance, never by the scan itself. Because the merge join is only sound on a scan that really iterates keys in their natural order, a multi-query scan now also verifies that promise against the stream as it is consumed and fails the scan if the store violates its declared key order, instead of silently dropping rows.

CQL scans under the Murmur3 partitioner use the lossless merge join

Cassandra's Murmur3 ring order — token first, key bytes among equal tokens — is now computed client-side through the driver's token factory (the same code token-aware routing relies on) and declared to the scan framework via the new StoreFeatures#getScanKeyOrder() hook. Multi-query CQL scans, including the split-parallel pipelines of storage.cql.parallel-scan-token-ranges, therefore merge with the lossless merge-join strategy instead of the bounded-buffer strategy, whose recovery from keys written concurrently with the scan is capped (a burst of more than 32 such keys between two matches of one query blanked that query's data for the keys behind the burst). The declared order is verified against every scan as it is consumed — the same tripwire that guards natural-order backends — so a client/server order mismatch fails the scan instead of silently dropping rows. Disabling the driver's token metadata (storage.cql.metadata-token-map-enabled = false) removes the declaration and restores the previous bounded-buffer merge; partitioners whose order the driver cannot compute (RandomPartitioner, Amazon Keyspaces' DefaultPartitioner) keep using it automatically.

Scan progress logging

StandardScannerExecutor now logs a start line (job, query count, processor count, collector type, queue capacity), a progress line every 30 seconds (rows produced by the storage scan, rows processed by workers, current rates, row-queue fill) and a completion summary (total rows, elapsed, average rate). The row-queue fill discriminates the bottleneck at a glance: a near-empty queue means the scan is storage-bound, a near-full queue means processing/index writes are the bottleneck. Per-puller counters are logged at DEBUG level, and mixed-index reindex jobs additionally report bulk-flush count/size/time through the custom scan metrics mixed-index-flushes, mixed-index-flushed-docs and mixed-index-flush-time-ms.

Whole-row deletion on vertex removal (super-node tombstone reduction)

Starting from version 1.2.0, when a vertex is removed JanusGraph deletes its entire storage row in a single operation on backends that support it (CQL/Cassandra issues one partition-level delete instead of one column delete per incident edge), drastically reducing tombstone pressure when removing super-nodes.

This behavior is enabled by default. To restore the previous per-column deletion behavior, set:

storage.drop-whole-row-on-vertex-removal=false
Write-only index state (WRITE_ONLY_ENABLED) and reindexing without automatic enablement

Starting from version 1.2.0 an index can be explicitly enabled for write operations only. A new schema status SchemaStatus.WRITE_ONLY_ENABLED and a new schema action SchemaAction.ENABLE_WRITE_ONLY were added. An index in the WRITE_ONLY_ENABLED state receives updates for all graph mutations but is not used to answer queries, and — in contrast to a REGISTERED index — SchemaAction.REINDEX preserves its state instead of automatically enabling the index. This enables the workflow create index → reindex it (without enabling) → enable it later when necessary as well as demoting an ENABLED index to write-only and re-enabling it later without a reindex. See Index Lifecycle for the full state machine and workflows.

The following behaviors were extended or clarified. None of them change the behavior of existing workflows:

  • Documentation clarification (no behavior change): indexes in the INSTALLED and REGISTERED states have always received updates for graph mutations; queries only ever use ENABLED indexes. Previous documentation incorrectly stated that only ENABLED indexes receive updates. If you relied on the documented (incorrect) behavior and want an index that receives no updates, disable it or use the new create-then-disable workflow below.
  • SchemaAction.DISABLE_INDEX can now also be applied to an INSTALLED or WRITE_ONLY_ENABLED index. Disabling an index within the same management transaction that creates it yields an index definition that is known to the cluster but never receives any writes until activated.
  • SchemaAction.REGISTER_INDEX can now also be applied to a DISABLED index to re-activate it for writes. Reaching the REGISTERED status guarantees that all instances write to the index, which makes a subsequent reindex lossless.
  • A pending index registration no longer overwrites a status change that superseded it. Previously, disabling an index while its registration acknowledgment was still pending could result in the index being moved back to REGISTERED once all acknowledgments arrived.
  • SchemaAction.REINDEX, SchemaAction.ENABLE_INDEX, SchemaAction.DISCARD_INDEX and SchemaAction.MARK_DISCARDED also accept indexes in the WRITE_ONLY_ENABLED state.

Note for downgrades: the new status is persisted in the schema. Schemas containing indexes (or mixed-index keys) in the WRITE_ONLY_ENABLED state cannot be read by JanusGraph versions older than 1.2.0. Only start using SchemaAction.ENABLE_WRITE_ONLY after all JanusGraph instances of the cluster have been upgraded to 1.2.0 or newer.

New action to remove stale index entries (REMOVE_STALE_ENTRIES)

Starting from version 1.2.0 stale index entries — entries which reference graph elements that no longer exist, for example because the element was deleted while an index status change was still propagating through the cluster or while the index was disabled — can be removed with the new schema action SchemaAction.REMOVE_STALE_ENTRIES:

mgmt = graph.openManagement()
mgmt.updateIndex(mgmt.getGraphIndex("myIndex"), SchemaAction.REMOVE_STALE_ENTRIES).get()
mgmt.commit()

The action complements REINDEX: a reindex restores missing entries for existing elements but never removes entries, while REMOVE_STALE_ENTRIES removes entries of deleted elements but never adds entries. The index status is not changed. The returned ScanMetrics report the number of removed entries under the custom metric stale-entries-removed. Previously such entries could only be removed one-by-one with the low-level StaleIndexRecordUtil helper (which required knowing the stale element ids and index record values upfront) or by dropping and rebuilding the whole index.

Supported index types:

  • Composite graph indexes: the internal index store is scanned and every entry whose element no longer exists is deleted.
  • Mixed graph indexes: the documents are enumerated through exists-queries against the index backend. This requires at least one index field whose data type supports exists queries; fields that do not support them are skipped with a warning.
  • Vertex-centric indexes are not supported by this action.

If custom vertex ids are used, avoid deleting and re-creating vertices under the same custom id while a REMOVE_STALE_ENTRIES job is running, because the job could remove the index entries of the re-created vertex (run REINDEX afterwards to restore them). With automatically assigned ids this race cannot occur because ids are never reused.

CDC-based mixed index synchronization

Starting from version 1.2.0, JanusGraph can keep mixed indexes (e.g. ElasticSearch) eventually consistent with the graph via Change-Data-Capture instead of a synchronous index write during the transaction that can diverge on failure (a cause of permanent stale indexes). This is opt-in and disabled by default.

With Apache Cassandra storage, set storage.cql.cdc=true to create the edgestore table with Cassandra CDC enabled, mark the mixed index backend with index.[X].cdc.enabled=true (and index.[X].cdc.synchronous=false for cdc-only mode, which skips synchronous index additions; the few relation-document deletions that change events cannot identify remain synchronous), and run the new janusgraph-cdc worker. Like storage.cql.cdc, the index.[X].cdc.* options are MASKABLE: they can be changed via mgmt.set(...) on a running cluster (instances and the worker pick up the new value when they next open the graph) or overridden in the local configuration of a process. The worker consumes the Cassandra change stream (e.g. via Debezium and Kafka) and reindexes affected elements from their current graph state, which is idempotent and order-independent. See CDC Mixed Index Synchronization for the full setup.

Transient Elasticsearch failures are retried instead of dropping the index mutation

Previously every Elasticsearch failure except an interrupt was reported as a PermanentBackendException. Because index mutations are applied after the storage mutations in a commit they cannot be rolled back, so a transient failure — a rolling restart, a saturated write queue, a socket timeout during a GC pause — dropped the mutation with a single ERROR log line and left the mixed index inconsistent with the graph until a reindex or a transaction-log recovery repaired it. The Elasticsearch client did have a retry mechanism, but it was disabled by default (index.[X].elasticsearch.retry-limit=0 and an empty index.[X].elasticsearch.retry-error-codes), could only act on a failure which produced an HTTP response, and once its attempts ran out the failure was still reported as permanent.

Such a failure is now reattempted at two levels, both enabled by default and both governed by one definition of what counts as transient:

index.[X].elasticsearch.retry-error-codes=429,502,503,504
index.[X].elasticsearch.retry-transport-failures=true
index.[X].elasticsearch.retry-limit=3
  • retry-error-codes lists the HTTP status codes considered transient, whether answered to a request or reported for an individual bulk item. Its default was previously empty.
  • retry-transport-failures (new) covers the failures which produce no HTTP response at all: connection refused, connection reset, socket timeout, a prematurely closed connection and a TLS failure — the last of which is what a rolling restart of a TLS secured cluster produces.
  • retry-limit is the number of attempts the Elasticsearch client makes. Its default was previously 0.

The two levels are:

  1. The Elasticsearch client reattempts the request retry-limit times, with the short waits controlled by retry-initial-wait and retry-max-wait. For a bulk request it resends only the items which failed, so these attempts never resend an item which already succeeded. This level applies to queries as well as writes.
  2. JanusGraph classifies whatever survives those attempts. An index mutation which still failed transiently is reported as a TemporaryBackendException, which the retry loop already present in BackendOperation reattempts with exponential backoff for up to storage.write-time (default 100s) — the same contract that already applies to a temporary storage failure. A bulk request qualifies only when every item which failed did so transiently, since a batch containing a permanently failing item — a mapping conflict, for instance — cannot succeed on a reattempt.

Be aware of the following while the options are enabled:

  • The second level resubmits the mutation without the documents which are known to have applied: every store whose own bulk request had already returned, and in the bulk request which failed every document none of whose items failed or went unsent. A bulk request larger than index.[X].elasticsearch.bulk-chunk-size-limit-bytes is sent in chunks, and the documents of the chunks which were never sent are resubmitted. A failure which reports no item statuses — no response at all, or a status for the bulk request as a whole — leaves it unknown what that request applied, so after one the interrupted bulk is resubmitted whole, the chunks which went through before it included. Resubmission is idempotent for whole-document writes, deletions and SET cardinality properties, but the values of a LIST cardinality property are appended, so a document resubmitted after such a failure can hold them twice. The same holds for a document one of whose items failed while another item of it had succeeded: it is resubmitted whole. Within one bulk request Elasticsearch answers the items of one document, which share a shard request, alike for the statuses it reattempts by default, so this takes either a per-item status which the operator listed as transient, or a document whose items were split between two chunks. Neither can happen to the document of an existing element whose complete indexed content is at hand, which is updated by a single item (see A missing Elasticsearch document is recreated whole below).
  • During an Elasticsearch outage a commit which touches a mixed index now takes up to storage.write-time to report the failure instead of failing fast.
  • While retry-transport-failures is enabled every SSLException is treated as transient, not only an interrupted handshake, so a write against a persistently misconfigured or untrusted certificate is reattempted for the whole write time before it fails.
  • Set index.[X].elasticsearch.retry-error-codes to an empty list and index.[X].elasticsearch.retry-transport-failures=false if LIST values duplicated after such a failure are less acceptable than a dropped mutation; that restores the previous behavior at both levels. Setting retry-limit=0 alone disables only the client level and keeps the JanusGraph level.

Deployments which set none of these options get the new behavior. Deployments which set retry-limit or retry-error-codes explicitly keep their values, and those values now also decide what JanusGraph reattempts.

An interrupt is covered by neither option, and its handling changes regardless. Previously convert recognised an InterruptedException only as the exception it was handed directly, and neither the Elasticsearch client nor JanusGraph's own retry wait ever hands one over unwrapped, so a cancelled index write was in practice reported as a PermanentBackendException. It is now recognised anywhere in the cause chain and reported as a TemporaryBackendException, and the interrupt status of the thread is restored once the InterruptedException has been consumed, so BackendOperation aborts its backoff wait immediately instead of reissuing the request for the whole write time budget.

An Elasticsearch update which failed because the document is missing is now reported

A bulk request reports item level failures inside an otherwise successful HTTP response. JanusGraph previously treated every item answered with HTTP 404 as a success, whatever mutation produced it. That is right for a mutation which only takes content out of the index — a whole document deletion, or the script which deletes fields — because an absent document already satisfies it. It is wrong for a mutation which adds content: there a 404 is a document_missing_exception, and the write did not happen.

A transaction which takes content out of a document and puts other content in — a property removed and a different one set on the same element, or a value of a LIST or SET cardinality property replaced — sends a field deletion and an addition against the same document, and mutate() withholds the upsert from the addition once a mutation has deletions. So if the Elasticsearch document was already missing, both items were answered with 404, both were discarded, and the addition was never indexed. Nothing reported it: the mutation returned normally, so even the <prefix>.indexProvider.<INDEX-NAME>.mutate.exceptions metric stayed at zero. Changing the value of a SINGLE cardinality property is not this case: the deletion of the old value is consolidated away because the same field is added, so that addition carries an upsert and recreates the document, from the changed field alone.

Such an item is now reported. A 404 is not among the transient status codes of index.[X].elasticsearch.retry-error-codes, so the failure is classified permanent and the mutation is dropped rather than reattempted — reattempting cannot recreate a document whose upsert was withheld. What a deployment sees changes on a graph whose mixed index has already diverged. The graph commit itself still returns normally, because JanusGraph commits the storage backend first and never aborts on a mixed index failure, but the commit now logs the dropped mutation at ERROR, counts it in the <prefix>.indexProvider.<INDEX-NAME>.mutate.exceptions metric and, where the transaction log is enabled, records SECONDARY_FAILURE for the transaction, from which transaction log recovery repairs the document. That is the intent: the alternative is that the divergence stays invisible. SchemaAction.REINDEX repairs the affected documents as well.

The same now holds for a removal against an Elasticsearch index which no longer exists. Elasticsearch answers a deletion against a missing index with a 404 carrying index_not_found_exception, which used to be taken for a success like every other 404 of a removal, so the commit passed silently although the index it was meant to update was gone. A removal stays exempt from a 404 which says only that the document is missing, since an absent document is the state it asked for.

Mixed index names on one backing index must now differ in more than case

An index backend derives its own index name from the JanusGraph index name case-insensitively — Elasticsearch lowercases it, Lucene names a directory after it — so two mixed indexes on the same backing index whose names differed only in case, such as byName and byname, shared one backend index: each other's documents and mappings, and a SchemaAction.DISCARD_INDEX of one dropped the other's documents. Creating the second one is now rejected with an IllegalArgumentException which names the existing index. Existing definitions are not touched: a graph which already holds such a pair keeps working as before, and is repaired by discarding one of the two and recreating it under a distinct name. Composite indexes have no backend index and are not affected.

Index restores are reattempted after a transient failure

IndexTransaction.restore, through which SchemaAction.REINDEX, stale entry removal, CDC index updates and transaction log recovery write their documents, called the index provider directly, so a failure the provider classified as transient — for Elasticsearch a 429, a 503 or a dropped connection — failed the batch outright, where the same failure during a commit is reattempted. It now runs through BackendOperation like a commit does: reattempted with backoff for up to storage.write-time, and only then reported. A restore writes whole documents, so a reattempt is idempotent.

Transient Elasticsearch read failures are reattempted

A mixed index query, count or aggregation against Elasticsearch wrapped every failure in a PermanentBackendException, so a throttled or momentarily unreachable cluster failed the traversal outright — although BackendTransaction already runs every index read through BackendOperation, which reattempts a TemporaryBackendException for up to storage.read-time, as it does for a storage read. The read paths now classify a failure the way the write path has since the retry options were unified: a status listed in index.[X].elasticsearch.retry-error-codes (429, 502, 503, 504 by default) and, with index.[X].elasticsearch.retry-transport-failures, a connection or TLS failure are transient and reattempted with backoff within storage.read-time; everything else remains permanent. The pages of a scroll are fetched while the caller consumes the result stream, outside that budget, and are not covered.

A missing Elasticsearch document is recreated whole

When an element exists in the graph but its document is missing from the mixed index — an index write lost earlier, a document removed by hand — the next mutation of the element recreated the document from the fields it touched alone, or, if the mutation also removed content, could not recreate it at all and was reported as a lost write. The transaction now hands the index provider the element's complete indexed content along with every mutation which updates an existing document — a property change on an edge or a vertex property included, although it replaces the relation — and the Elasticsearch provider sends every change of the document, its removals, its collection additions and its single-valued fields, in one script update which carries that content as its upsert. A document which turns out to be missing is therefore recreated whole in the same round trip, whichever shape the mutation has, and in a store with an ingest pipeline it goes through the pipeline, as Elasticsearch runs the pipeline of a bulk request on the upsert of an update whose document is missing. A document which exists is updated as before, except that a single-valued field is assigned rather than merged, which replaces an object value such as a geo shape as a whole. Being one item, the update of a document is also applied or rejected as a whole: a bulk request split into chunks can no longer separate a document's changes, and a reattempt never replays a part of them which had applied. This has two costs. The content is read while the transaction commits, so a commit which updates an existing element covered by a mixed index reads that element's indexed properties, which are usually loaded already. And every update of an existing document carries that content once, so bulk requests grow with the size of the documents they update; an update the content would make larger than index.[X].elasticsearch.bulk-chunk-size-limit-bytes is sent without it, as before this change, rather than failing. A mixed index on a cdc-only backend, whose documents are written asynchronously, is not affected.

Elasticsearch searches open a scroll context only for results larger than a page

Every mixed index query whose limit was at least index.[X].max-result-set-size (50 by default), and every query without a limit, read its result through the scroll API in pages of that size. A query with a limit whose result had at least as many hits also left its scroll context open until index.[X].elasticsearch.scroll-keep-alive expired, because the result stream stopped at the limit before the last page. So a limit(1000) was 20 requests and a 60 s scroll context, and a lookup without a limit was two requests, the scroll search and the one which released its context, both waited for. Elasticsearch also had to count every match of these queries, as a scroll may not switch the total off. JanusGraph now fetches a result whose offset and limit together are within 10,000 hits, Elasticsearch's default index.max_result_window, in one request of exactly that size and without counting the total, and Elasticsearch applies the offset of a direct index query in that request. A query without a limit, or beyond that size, first asks for one hit more than a page after the offset, or for what is left up to the 10,000th hit if that is less; only when the result is larger, or the offset is 10,000 or more, does a scroll take over, which costs such a result one request more than before. The scroll context is released as soon as the result has been read to its end, the limit is reached, or the traversal is closed, which JanusGraph Server does after every request; embedded code should close a traversal it abandons before its end, otherwise the context expires after the keep-alive as before. The release no longer waits for the cluster's answer, and a release which fails no longer fails the query; a release the cluster rejects is logged as a warning once. Deployments which set index.[X].elasticsearch.setup-max-open-scroll-contexts to false, such as Amazon OpenSearch Service, are therefore far less likely to reach the cluster's limit of open scroll contexts.

Relations without properties are parsed once

A loaded relation is deserialized in two steps: its type, direction, id and other end first, and its properties only once something asks for them. The deserialized form is cached on the entry, and whether the cache already held the properties was told from whether it held any. So a relation without properties, which is the common case for a vertex property without meta-properties and for an edge without properties, looked like it had never had its properties parsed, and every read of them parsed the whole entry again: every properties() of a vertex property, every property(key) or has(key) check on an edge, and every vertex which Gremlin Server returns with its properties, since its serializers read the meta-properties of each of them. The cache now records whether the properties were parsed, so each loaded entry of a relation is parsed with its properties at most once. Moreover, the first step now finds out when nothing follows the header, which is the case for a relation without properties whose type has neither a signature nor a sort key, the default, and whose entry carries no timestamp, TTL or visibility metadata. Such an entry is parsed exactly once, where it was parsed twice by the first read of its properties before. A relation without properties whose type has a signature, or which is read through a vertex-centric index, is still parsed twice by its first read, but no longer by every later one. An entry with timestamp, TTL or visibility metadata, which the storage.meta.* options enable (visibility only on a backend with cell-level visibility), is unaffected: the metadata is a property of the relation, so its full parse was already kept. The property map is only allocated for a relation which has properties.

Composite index lookups of several values read their index rows together

A composite index lookup of several values, as has(key, within(values)) makes, or as a composite index on more than one key makes for every combination of the values given for its keys, read each of the index rows this gives only after the previous one had arrived, so each row cost a round trip to the storage backend. Where the backend supports multi-key queries, which the CQL backend (Cassandra, ScyllaDB) and HBase do, the rows are now read together, in calls of up to 1,000 rows which the backend executes concurrently. That happens without a limit, as every row is read then anyway, and with a limit for a unique index, whose rows hold at most one element each, in calls of as many rows as the limit can still take. The limits which JanusGraph adds itself count as limits: those of query.smart-limit and of a query.hard-max-limit below its default, and the one an order the index can't serve brings. A lookup with a limit on an index which isn't unique, and every lookup on a backend without multi-key queries, such as BerkeleyJE, still reads one row after the other and stops at the limit. Either way a lookup asks the storage backend for no row which reading one row after the other would not ask for. query.batch.enabled=false turns reading rows together off, as it turns off batching for traversal steps. Building the condition of a within() or without() also compared each value with every value before it to drop duplicates, so it grew with the square of the number of values; it now takes time in proportion to them.

A vertex which gains relations in a transaction allocates about 2 KB less

Every vertex which gains a relation in a transaction keeps the relations it gained, and since JanusGraph 1.1.0 it sized the set of its edges for 300 edges up front, about 2 KB (4 KB on heaps of 32 GB or more), whether it gained one relation or hundreds. Once one of its existing relations was replaced by a copy, as setting a property of an existing edge or a meta-property of an existing vertex property does, it also sized a table for 330 replaced relations, about 2 KB more. Every transaction allocated such a set as well, read-only ones included. The sets and tables now start small and grow with the relations added. On the inmemory backend, a transaction which adds 1,000 vertices with two properties and an edge each allocates 9% less, one which sets a property of 1,000 existing vertices 24% less, and one which sets a property of an edge of each 22% less.

Storage operations no longer contend for one random generator

Nearly every storage operation, among them every read of a transaction and every mutation its commit writes, runs through a loop which reattempts temporary failures after randomized waits. It drew its first wait before the first attempt, from a random generator which all threads shared, although nearly every operation succeeds at once, and threads which draw from one generator at the same time contend for its seed. The wait is now drawn only once an operation has failed temporarily, from the random generator of its thread, and the waits are the same as before. On the inmemory backend, reading a property of 1,000 vertices on each of eight threads takes 18% less time.

Elasticsearch clients open up to 30 connections to a single host

The Elasticsearch client of an index backend opened at most 10 connections to each Elasticsearch host and 30 to all hosts together, the defaults of the Elasticsearch REST client, which JanusGraph had no option for. A request waits for a free connection, so an index backend had no more than 10 queries and bulk requests in flight at a single host, such as a load balancer or the endpoint of a hosted cluster, however many threads issued them. The new option index.[X].elasticsearch.max-connections sets the total, 30 by default, and index.[X].elasticsearch.max-connections-per-host the connections to each host, by default the total divided evenly among the hosts, but at least 10. So a single host now takes all 30 connections and two hosts 15 each, while three or more hosts keep 10 each, which keeps a host that stops answering from holding every connection. With 32 threads issuing mixed index queries, a single Elasticsearch 9 node on the same machine answers 22% more of them per second, and a host which takes 20 ms to answer three times as many. The new option index.[X].elasticsearch.io-threads sets the number of I/O threads of each client, which is otherwise the number of processors, for every index backend of every graph, and index.[X].elasticsearch.compression compresses requests and responses with gzip.

Conflicting Elasticsearch updates are reattempted

Transactions which change the same element concurrently also update its document in a mixed index concurrently, and Elasticsearch fails an update which finds that the document changed after the update read it, with status 409. JanusGraph treats that status as permanent, so the change was missing from the mixed index unless transaction recovery repaired it. index.[X].elasticsearch.retry_on_conflict now defaults to 3, where it was not sent unless set, so Elasticsearch reattempts such an update against the latest version of the document. Set it to 0 for the previous behavior. The client's own reattempts of a request, after a backoff which starts at index.[X].elasticsearch.retry-initial-wait and grows tenfold up to index.[X].elasticsearch.retry-max-wait, now wait a random time between half of the backoff and all of it, so that requests which failed together don't all come back at the same moment.

CQL reads more properties of a vertex per query

With storage.cql.grouping.slice-allowed, true by default, the CQL backend reads the properties of a vertex whose keys have single cardinality, each one column, together: one query with column1 IN ? for up to storage.cql.grouping.slice-limit of them, 20 by default. Once a vertex had filled such a query, each further property was read with a query of its own, so 50 such properties of a vertex took 31 queries instead of 3. A full query is now followed by another one. On Cassandra 5, reading 50 properties of each of 1,000 vertices in one batch takes half as long.

CQL write batches are marked idempotent

The DataStax driver sends a request again, to the same node or another one, after a closed connection, an overloaded node or a server error, asks its retry policy about a write which timed out, and runs speculative executions, only for statements which are marked idempotent. The query builder with which JanusGraph prepares its statements marks them idempotent, but JanusGraph sends its writes in batches, and a batch isn't idempotent unless it is marked. So every such failure of a write reached JanusGraph, which reattempted a whole chunk of the mutations of a commit after a backoff of at least 25 ms. While graph.assign-timestamp is true, the default, JanusGraph now marks its batches idempotent, as a write sent again then writes the same cells with the same timestamp, and a speculative execution policy configured through storage.cql.internal.* now applies to its writes too; it already applied to its reads. With the new option storage.cql.idempotent-writes=false, or without assigned timestamps, writes keep the driver's default idempotence, basic.request.default-idempotence, as every write did before.

The Gremlin pool of JanusGraph Server can run on virtual threads

JanusGraph Server evaluates requests in the Gremlin pool of Gremlin Server, gremlinPool platform threads, by default one per processor, and a pool thread stays with its request while the request waits for the storage and index backends. The pool therefore evaluated only as many requests per second as its threads could wait for, however idle the processors were, and a larger pool meant as many platform threads. On Java 24 or later, the new JanusGraph Server setting gremlinPoolVirtualThreads: true runs the pool on virtual threads. The pool keeps its bounds, gremlinPool requests evaluated at once, maxWorkQueueSize more waiting and rejection beyond, but a request which waits for a backend releases its platform thread, so gremlinPool can be raised to the wanted concurrency. With Cassandra on the same 18-processor machine, 256 virtual threads answered 20,900 requests per second at 256 requests in flight, against 16,800 for 256 platform threads and 5,480 for the default pool, while the server kept about 90 platform threads instead of 326. The setting is off by default, and JanusGraph Server refuses to start with it on Java versions older than 24, where a virtual thread which waits inside synchronized, as JanusGraph transactions do while they commit, keeps its platform thread.

tx.max-commit-time now defaults to 300 s and is checked against the least a commit may take

tx.max-commit-time is the time after which transaction recovery considers a transaction failed and restores the index documents of the elements it changed from the storage backend, counted from the transaction's first log entry. A commit reattempts its storage write and then each of its index writes in turn, each for up to storage.write-time (100 s by default), so the option has to exceed storage.write-time multiplied by one plus the number of index backends. That is only the minimum: the time a commit spends preparing its writes counts as well, a transaction with more than storage.buffer-size mutations writes storage in several chunks, each reattempted on its own, a storage backend without transaction isolation (Cassandra and HBase, or BerkeleyDB with storage.transactions=false) commits the schema elements a transaction creates in a storage write of their own first, and the last attempt of each write can run past its write time. Yet the option defaulted to 10 s, so recovery could restore the documents of a transaction which was still committing, and the index writes that commit had yet to make then landed on documents which already reflected it — appending the values of a LIST cardinality property twice, or removing an occurrence which should have stayed. The default is now 300 s: the storage write and one index backend at the default write time, with one write time to spare, for instance for a second storage chunk. JanusGraph logs a warning at graph open, when transaction logging is enabled, whenever tx.max-commit-time does not exceed that minimum for the configured index backends; the warning names the minimum and suggests a management system call, ready to paste, which sets one write time more, so a graph with two or more index backends is told what to set, and a graph which commits large transactions should allow more still. The default stops there because the transaction recovery process keeps every transaction it reads, the content of its modifications included, until it has read the log tx.max-commit-time past the transaction (see below): a recovery process of a write-heavy graph now holds thirty times as much in memory as it did with the old default.

After an upgrade a graph which never set the option resolves the new default as soon as its instances are restarted; the value of a GLOBAL option is only stored when it is set explicitly. A graph which did set it keeps its value, gets the warning if that value is too short, and raises it with the management system call the warning suggests, for two index backends at the default write time mgmt.set("tx.max-commit-time", java.time.Duration.parse("PT6M40S")) followed by mgmt.commit(), which a transaction recovery processor picks up when it is next started. The recovery of a transaction which really did fail starts correspondingly later. A transaction recovery process waits about 146 years at most, half the nanoseconds a long holds, so that how far it has read past a transaction always compares correctly with the wait; a longer tx.max-commit-time used to keep it from starting at all and now makes it wait up to that limit.

A transaction which writes a user log (TransactionBuilder.logIdentifier) is exposed for longer: its commit writes the user-log event after the index writes and only then its final status, and a transaction which expires before recovery has read that status gets its user-log event sent again. Such a transaction needs its user-log write (up to log.user.max-write-time when log.user.send-delay is 0; by default the event is sent in the background) inside tx.max-commit-time as well; the final status's own write does not count, since its log entry is timed from before the write (see below). The log identifier is set per transaction, so the warning cannot take it into account.

Transaction recovery no longer gives up on a transaction before reading its final status

Transaction recovery gave up on a transaction tx.max-commit-time after it had read the transaction's first log entry. It reads the transaction log in polls, though, every log.tx.read-interval (5 s by default), each of which can spend up to log.tx.max-read-time on its reads and stops at the end of the 100 s long chunk of the log it is in, so it could read a transaction's final status well after its first entry even when the commit took far less than tx.max-commit-time. Whenever that came too late, recovery took a transaction which had succeeded for a failed one: it restored the index documents of every element the transaction changed, sent its user-log event again, so that user-log consumers received it twice, and counted the transaction once more when the final status arrived. The closer tx.max-commit-time was to the read interval, the likelier this was: with both at 5 s, the least recovery waits, a final status read one poll after the transaction's first entry arrived within milliseconds of the expiry, before or after it.

Recovery now measures the time it gives a transaction by how far it has read the log rather than by how long it has waited: it gives a transaction up once it has read and processed the log up to tx.max-commit-time past the first of the transaction's entries it read, however far apart and however slowly its reads come, so by then it has read everything a commit wrote within tx.max-commit-time and which became visible within log.tx.read-lag-time. A recovery process catching up on a backlog no longer holds each transaction for tx.max-commit-time of waiting either, and gets through the backlog with less in memory. A transaction whose commit does outlast tx.max-commit-time, or whose instance fails between its user-log write and its final status, still gets its user-log event sent again, the repeat carrying the transaction id of the original; the former is counted twice as well, once as failed and once as succeeded when its final status arrives. A partition of the log whose reads fail holds the progress back, and with it every expiry, until they succeed; one whose reads have failed for good, which stops being read, is left out of the progress once its messages are processed, and an error is logged for it. Once that goes for every partition, nothing is read any more, and the progress runs on with the clock from where reading stopped, so that recovery still gives the transactions it has in hand up rather than hold them for good.

The second of the numbers getStatistics() returns, the transactions which failed and whose repair was attempted, now counts a transaction once that attempt has finished rather than before it starts, so the repairs of the transactions it counts are done and the third number already includes those which could not be repaired.

Schema changes committed by ordinary transactions reach the other instances

With schema.constraints=true and a schema maker which creates missing constraints, an ordinary transaction adds property and connection constraints to the schema, and so do addConnection(...) and addProperties(...), on a transaction or on the management system, which sent no eviction either. Other JanusGraph instances used to keep the definitions they had cached, and so created a second copy of a constraint the first time they used it themselves. Such a commit now also sends a cache eviction for the schema elements it changed over the management log, like a management commit does, but as one which nothing waits to see acknowledged: an acknowledged eviction registers a trigger which waits for every instance to acknowledge once its open transactions have closed, and with graph.management-auto-close-stale-instances could get an instance with a long-running transaction force-closed because of an ordinary write. Every instance which reads the eviction, the sender included, expires the elements from its schema cache and its open transactions. A management commit which changes definition edges and sends evictions of its own sends both, and receivers expire the elements twice, at the cost of a re-read. Instances of earlier versions process the eviction as well and acknowledge it; that acknowledgement is ignored, so a rolling upgrade needs no preparation. During the upgrade, each such eviction costs an older instance the thread which waits for its open transactions to close before it acknowledges, for up to a minute, and one with a transaction open for longer than that logs the stale-transaction error it logs for any eviction it waited that long for.

Closing a graph on BerkeleyJE no longer interrupts its log readers

BackgroundThread.close() could interrupt the thread's action() or cleanup(), which its contract rules out, when the thread was between its checks. A log's send thread flushes its last messages in cleanup(), and BerkeleyJE invalidates its whole environment when a thread is interrupted in the middle of a file operation, so every later use of that environment in the JVM failed: that is how BerkeleyGraphTest lost most of its tests on some CI runs. close() now decides to interrupt, and interrupts, under a lock which the thread holds while it turns its interruptibility off; an interrupt pending before action() ends the loop instead of reaching it, and one pending before cleanup() is cleared.

For the same reason, closing a log on a storage backend which does not support interruption, which is BerkeleyJE, no longer interrupts the log's reader threads after a second. This covers graph.close() and the stop of a recurring transaction recovery. The log waits for the readers to finish the pull under way and the messages they have in hand, however long that takes, warning every log.<name>.max-read-time (a second at the least) after the first second, and it does not give that wait up when the closing thread is interrupted. It holds no lock of the log, of its manager or of the log processor framework meanwhile, so a reader which opens a log, registers or unregisters a reader, or adds or removes a log processor while the log closes finishes. A MessageReader which never returns therefore holds graph.close() up on BerkeleyJE, where it used to be interrupted after a second. A log closed from one of its own reader threads cannot wait for them and keeps the second, and so do the logs of other storage backends. The reader threads of every log are now named after it.

Version 1.1.0 (Release Date: November 7, 2024)

<dependency>
    <groupId>org.janusgraph</groupId>
    <artifactId>janusgraph-core</artifactId>
    <version>1.1.0</version>
</dependency>
compile "org.janusgraph:janusgraph-core:1.1.0"

Tested Compatibility:

  • Apache Cassandra 3.11.10, 4.0.6
  • Apache HBase 2.6.0
  • Oracle BerkeleyJE 7.5.11
  • ScyllaDB 6.2.0
  • Elasticsearch 6.0.1, 6.6.0, 7.17.8, 8.15.3
  • Apache Lucene 8.11.1
  • Apache Solr 8.11.1
  • Apache TinkerPop 3.7.3
  • Java 8, 11

Installed versions in the Pre-Packaged Distribution:

  • Cassandra 4.0.6
  • Elasticsearch 7.14.0

Changes

For more information on features and bug fixes in 1.1.0, see the GitHub milestone:

Assets

Upgrade Instructions

Inlining vertex properties into a Composite Index

Inlining vertex properties into a Composite Index structure can offer significant performance and efficiency benefits. See documentation on how to inline vertex properties into a composite index.

Warning

Important Notes on Compatibility. 1. Backward Incompatibility: Once a JanusGraph instance adopts this new schema feature, it cannot be rolled back to a prior version of JanusGraph. The changes in the schema structure are not compatible with earlier versions of the system. 2. Migration Considerations: It is critical that users carefully plan their migration to this new version, as there is no automated or manual rollback process to revert to an older version of JanusGraph once this feature is used.

BerkeleyJE ability to overwrite arbitrary settings applied at EnvironmentConfig creation

The new namespace storage.berkeleyje.ext now allows to set custom configurations which were not directly exposed by JanusGraph. The full list of possible setting is available inside the Java class com.sleepycat.je.EnvironmentConfig. All configurations values should be specified as String and be formated the same as specified in the official sleepycat documentation. Example: storage.berkeleyje.ext.je.lock.timeout=5000 ms

JSON schema initializer

For simplicity JSON schema initialization options has been added into JanusGraph. See documentation to learn more about JSON schema initialization process.

Batched Queries Enhancement: Introduction of JanusGraphNoOpBarrierVertexOnlyStep

In previous versions, when a query that could benefit from batch-query optimization (multi-query) was executed without a user-defined barrier step, JanusGraph would inject a NoOpBarrierStep by default. This approach allowed batching for edges and properties, which do not gain advantages from multi-query optimization.

Starting with JanusGraph 1.1.0, this behavior has been improved. The system now injects a JanusGraphNoOpBarrierVertexOnlyStep instead of the standard NoOpBarrierStep when no barrier steps are detected. This change ensures that batching is applied exclusively to vertices, which do benefit from batch queries, while excluding edges and properties from the batching process.

If a user explicitly defines a .barrier() step in the query, the system will continue to use the NoOpBarrierStep as expected.

Batch Query Optimizations Now Support Traversals Containing the drop() Step

Starting with JanusGraph 1.1.0, batch optimizations for vertex removal have been introduced in the drop() step and are enabled by default. Previously, any batch optimization would be skipped for queries containing at least one drop() step. However, with this update, such queries are now eligible for batch query optimization (multi-query).

Please note that the LazyBarrierStrategy (a TinkerPop strategy) is disabled for any query that includes at least one drop() step.

To disable the drop() step optimization and maintain the previous behavior, users can set the following configuration:

query.batch.drop-step-mode=none

Lazy Loading for Relations

The new transaction configuration option is added lazyLoadRelations() which sets lazy-load for all properties and edges of the vertex. If enabled, then ids and values are deserialized upon demand.
When enabled, it can lead to a performance improvement on large-scale read operations, if only certain types of relations are being read from the vertex.
See performance comparison in GitHub PR #4343.

Text predicates support extended for remote connections

The following text predicates can now be used with remote connections:

  • textNotContains
  • textNotContainsFuzzy
  • textNotContainsPrefix
  • textNotContainsRegex
  • textContainsPhrase
  • textNotContainsPhrase
  • textNotFuzzy
  • textNotPrefix
  • textNotRegex
Vertex mutation optimizations

The following improvements are made for adding new vertex or updating vertex properties:

  • Improved operations to detect vertex changes during transaction-commit from O(N) to O(1);
  • Optimizing properties search by key for newly-created vertex;
  • Optimizing previous property/edge search during vertex update;

For more information see GitHub PR #4292

Version 1.0.1 (Release Date: November 6, 2024)

<dependency>
    <groupId>org.janusgraph</groupId>
    <artifactId>janusgraph-core</artifactId>
    <version>1.0.1</version>
</dependency>
compile "org.janusgraph:janusgraph-core:1.0.1"

Tested Compatibility:

  • Apache Cassandra 3.11.10, 4.0.6
  • Apache HBase 2.5.8
  • Oracle BerkeleyJE 7.5.11
  • ScyllaDB 5.1.4
  • Elasticsearch 6.0.1, 6.6.0, 7.17.8, 8.15.3
  • Apache Lucene 8.11.1
  • Apache Solr 8.11.1
  • Apache TinkerPop 3.7.3
  • Java 8, 11

Installed versions in the Pre-Packaged Distribution:

  • Cassandra 4.0.6
  • Elasticsearch 7.14.0

Changes

For more information on features and bug fixes in 1.0.1, see the GitHub milestone:

Assets

Version 1.0.0 (Release Date: October 21, 2023)

<dependency>
    <groupId>org.janusgraph</groupId>
    <artifactId>janusgraph-core</artifactId>
    <version>1.0.0</version>
</dependency>
compile "org.janusgraph:janusgraph-core:1.0.0"

Tested Compatibility:

  • Apache Cassandra 3.11.10, 4.0.6
  • Apache HBase 2.5.0
  • Oracle BerkeleyJE 7.5.11
  • ScyllaDB 5.1.4
  • Elasticsearch 6.0.1, 6.6.0, 7.17.8, 8.10.4
  • Apache Lucene 8.11.1
  • Apache Solr 8.11.1
  • Apache TinkerPop 3.7.0
  • Java 8, 11

Note

Google Bigtable was removed from this list because there is no automatic testing in place specifically for that backend. Since the adapter for Bigtable is however just using the HBase adapter, it is also covered by the tests for HBase.

We invite anyone who is interested in the Bigtable storage adapter to help with this by contributing so that the tests for HBase are also automatically executed for Bigtable. More information can be found in this GitHub issue: janusgraph/janusgraph#415.

Installed versions in the Pre-Packaged Distribution:

  • Cassandra 4.0.6
  • Elasticsearch 7.14.0

Changes

For more information on features and bug fixes in 1.0.0, see the GitHub milestone:

Assets

Upgrade Instructions

Upgrade TinkerPop from 3.5.x to 3.7.0 (breaking)

Some packages under gremlin-driver are moved to gremlin-util module. This means you need to change some classnames of serializers, e.g. from org.apache.tinkerpop.gremlin.driver.ser.* to org.apache.tinkerpop.gremlin.util.ser.*. See this for more detail.

GraphSON serializers are renamed, e.g. from org.apache.tinkerpop.gremlin.driver.ser.GraphSONMessageSerializerV3d0 to org.apache.tinkerpop.gremlin.util.ser.GraphSONMessageSerializerV3. See this for more detail.

String vertex ID support (breaking change)

Users now can use custom string vertex ids. See Custom Vertex ID documentation. Prior to this change, JanusGraph automatically casts IDs of string type to long type if possible. Now this auto conversion is disabled. If you have a vertex with ID 1234, g.V("1234") would no longer help you find the vertex - you would have to do g.V(1234) now.

This feature brings about a breaking change to GraphBinary serializer. As such, users who use GraphBinary serialization format must update JanusGraph server and all clients at once, since the change is backward incompatible.

Warning

Even if you don't enable string vertex id feature, you are still impacted as long as you use GraphBinary serializer.

Upgrade of log4j to version 2

This change requires a new log4j configuration. You can find an example configuration in conf/log4j2-server.xml. As a result of the changed configuration format, we clean up all configurations. This could lead to unexpected new log lines. Please open an issue, if you see any unwanted log line.

Note

Log4j is only used for standalone server deployments and JanusGraph testing.

Removal of cassandra-all dependency

JanusGraph had a dependency on cassandra-all only for some Hadoop-related classes. We moved these few classes into a new module cassandra-hadoop-util to reduce the dependencies of JanusGraph. If you are running embedded JanusGraph with Cassandra, you have to exclude the cassandra-hadoop-util from janusgraph-cql.

Drop support for HBase 1

We are dropping support for HBase 1.

Drop support for Solr 7

We are dropping support for Solr 7.

Drop support for Gryo MessageSerializer

Support for Gryo MessageSerializer has been dropped in TinkerPop 3.6.0 and we therefore also no longer support it in JanusGraph. GraphBinary is now used as the default MessageSerializer.

Remove support for old serialization format of JanusGraph predicates

We are dropping support for old serialization format of JanusGraph predicates. The old predicates serialization format is only used by client older than 0.6. The change only affects GraphSON.

Allow removal of configuration keys

Users can now remove configuration keys in the ConfiguredGraphFactory's configuration:

ConfiguredGraphFactory.removeConfiguration("<graph_name>", Collections.singleton("<config_key>"))

Or the global configuration:

mgmt = graph.openManagement()
mgmt.remove("<config_key>")
mgmt.commit()

Note that the above commands should be used with care. They cannot be used to drop an external index backend if it has mixed indexes for instance.

New index management

The index management has received an overhaul which enables proper index removal. The schema action REMOVE_INDEX is no longer available and has been replaced by DISCARD_INDEX. See Index Lifecycle documentation for more details.

totals for direct index queries now applies provided offset and limit

Direct index queries which search count for totals (vertexTotals, edgeTotals, propertyTotals and direct execution of IndexProvider.totals) now apply provided limit and offset. Previously provided limit and offset were ignored.
For example, previously the following query would return 500 if there were 500 indexed elements:

gremlin> graph.indexQuery("textIndex", "v.\"text\":fooBar").limit(10).vertexTotals()

Now the above query will return 10 elements because we limited the result to 10 elements only. Same applies to offset. Previously the following query would return 500 if there were 500 indexed elements:

gremlin> graph.indexQuery("textIndex", "v.\"text\":fooBar").offset(10).vertexTotals()

Now the above query will return 490 as a result because we skip count of the first 10 elements.
offset provided with limit will apply offset first and limit last.
For example, the above query will return 10 elements now if there were 500 indexed elements:

gremlin> graph.indexQuery("textIndex", "v.\"text\":fooBar").limit(10).offset(10).vertexTotals()

The new logic is applied similarly to Direct Index Queries vertexTotals(), edgeTotals(), propertyTotals() as well as internal JanusGraph method IndexProvider.totals.

Add support for Java 11

JanusGraph now officially supports Java 11 in addition to Java 8. We encourage everyone to update to Java 11.

Note

The pre-packaged distribution now requires Java 11.

Batch Processing enabled by default. Configuration changes.

query.batch is now a configuration namespace. Thus, previous query.batch configuration is replaced by query.batch.enabled. query.limit-batch-size configuration option is changed to query.batch.limited.

query.batch-property-prefetch was replaced by a better configurable option. In case previous behaviour is desired then use query.batch.has-step-mode = none as replacement for query.batch-property-prefetch = false or use query.batch.has-step-mode = all_properties as replacement for query.batch-property-prefetch = true.

query.fast-property has no influence on values, properties, valueMap, propertyMap, elementMap anymore when query.batch.enabled is true. By default, those steps are configured to fetch only required properties (with separate query per property), but the behaviour can be changed with the configuration query.batch.properties-mode. In case previous behavior is desired, use query.batch.properties-mode = required_properties_only for query.fast-property = false or use query.batch.properties-mode = all_properties for query.fast-property = true.

label step now uses pre-fetching strategy by default. Use query.batch.label-step-mode = none to disable pre-fetching optimization for label step.

Batch processing allows JanusGraph to fetch a batch of vertices from the storage backend together instead of requesting each vertex individually which leads to a high number of backend queries. This was however disabled by default in JanusGraph because these batches could become much larger than what was needed for the traversal and therefore have a negative performance impact for some traversals. That is why an improved batch processing mode was added in JanusGraph 0.6.0 that limits the size of these batches retrieved from the storage backend, called Limited Batch Processing. This mode therefore solves the problem of having potentially unlimited batch sizes. That is why we now enable this mode by default as most users should benefit from this limited batch processing.

If you want to continue using JanusGraph without batch processing, then you have to manually disable it by setting query.batch.enabled to false.

The size of the batches can be limited by using barrier() steps if limited batch processing is used (query.batch.limited set to true). A special strategy exists which already inserts barrier() steps by default for some steps, the LazyBarrierStrategy. A new configuration option query.batch.limited-size exists to configure default barrier step size for batch processing for batch cases when LazyBarrierStrategy not applied .barrier step and no user-provided barrier step exists for batchable query part. Notice, that query.batch.limited-size is only used when query.batch.limited is true (default in this version).

Batch registration for nested batch compatible steps is changed for repeat step

Previously any batch compatible steps like out, in, values, etc. would receive vertices for batch registration from all repeat parent steps, but only for their starts in case of multi-nested repeat steps (skipping their subsequent iterations registration). With JanusGraph 1.0.0 batches registration for the subsequent iterations of multi-nested repeat steps are used as well.

g.V(startVertexId).emit().
    repeat(__.repeat(__.in("connects")).emit()).
    until(__.loops().is(P.gt(2)))
In the example above multi-nested repeat case would not register vertices returned from the inner emit() step for the next outer iteration which would result in sequential calls of in("connects") for next outer iteration. The behaviour is now changed to register these vertices for the next child repeat step start.

The behaviour can be controlled by query.batch.repeat-step-mode configuration option.
In case the old behaviour is preferable then query.batch.repeat-step-mode should be set to starts_only_of_all_repeat_parents.

However, in cases when transaction cache is small and repeat step traverses more than one level deep, it could result for some vertices to be re-fetched again which would mean a waste of operation when it isn't necessary. In such situations closest_repeat_parent mode might be more preferable than all_repeat_parents.
With closest_repeat_parent mode vertices for batch registration will be received from the start of the closest repeat step as well as the end of the closest repeat step (for the next iteration). Any other parent repeat steps will be ignored.

ConfiguredGraphFactory now creates separate indices per graph in Elasticsearch

If the ConfiguredGraphFactory is used together with Elasticsearch as the index backend, then the same Elasticsearch index is used for all graphs (if the same index names were used across different graphs). Now it is possible to let JanusGraph create the index names dynamically by using the graph.graphname if no index.[X].index-name is provided in the template configuration. This is exactly like it was already the case for the CQL keyspace name for example.

Users who don't want to use this feature can simply continue providing the index name via index.[X].index-name in the template configuration.

Mixed index aggregation optimization

A new optimization has been added to compute aggregations (min, max, sum and avg) using mixed index engine (if the aggregation function follows an indexed query). If the index backend is Elasticsearch, a double value is used to hold the result. As a result, aggregations on long numbers greater than 2^53 are approximate. In this case, if the accurate result is essential, the optimization can be disabled by removing the strategy JanusGraphMixedIndexAggStrategy: g.traversal().withoutStrategies(JanusGraphMixedIndexAggStrategy.class).

Add support for ElasticSearch 8

JanusGraph now supports ElasticSearch 8.
Notice, Mapping.PREFIX_TREE mapping is no longer available for Geoshape mappings using new ElasticSearch 8 indices.
Mapping.PREFIX_TREE is still supported in ElasticSearch 6, ElasticSearch 7, Solr, Lucene.
For ElasticSearch the new Geoshape mapping was added Mapping.BKD.
It's recommended to use Mapping.BKD mapping due to better performance characteristics over Mapping.PREFIX_TREE.
The downside of Mapping.BKD is that it doesn't support Circle shapes. Thus, JanusGraph provides BKD Circle processors to convert Circle into other shapes for indexing but use Circle at the storage level. More information about Circle processors available under configuration namespace index.[X].bkd-circle-processor.
ElasticSearch 8 doesn't allow creating new indexes with Mapping.PREFIX_TREE mapping, but the existing indices using Mapping.PREFIX_TREE will work in ElasticSearch 8 after migration. See ElasticSearch 8 migration guide.

Add support for ScyllaDB driver

A new module janusgraph-scylla provides ability to run JanusGraph with ScyllaDB Driver which is an optimized fork version of DataStax Java Driver for ScyllaDB storage backend.
For ScyllaDB storage backend you can use either janusgraph-scylla or janusgraph-cql (which is a general CQL storage driver implementation). That said, it's recommended to use janusgraph-scylla for ScyllaDB due to the provided internal optimizations (more about ScyllaDB driver optimizations can be found here).

Notice that janusgraph-cql and janusgraph-scylla are mutually exclusive. Use only one module at a time and never provide both dependencies in the same classpath. See ScyllaDB Storage Backend documentation for more information about how to make scylla storage.backend options available.

CQL ExecutorService purpose change

Previously CQL ExecutorService was used to control parallelism of both CQL IO operations and results deserialization. Starting from JanusGraph 1.0.0 CQL ExecutorService is now used for CQL results deserialization only. All CQL IO operations are now using internal async approach.
The default pool size is now set to have a value of number of cores multiplied by 2. This ExecutorService is now mandatory and cannot be disabled. The default ExecutorService core pool size is not recommended to be changed as the default value is considered to be optimal unless users want to artificially limit parallelism of CQL results deserialization jobs.

A new multi-query method added into KeyColumnValueStore (affects storage adapters implementations only)

A new method Map<SliceQuery, Map<StaticBuffer, EntryList>> getMultiSlices(MultiKeysQueryGroups<StaticBuffer, SliceQuery> multiKeysQueryGroups, StoreTransaction txh) is added into KeyColumnValueStore which is now preferred for multi-queries (batch queries) with multiple slice queries.
In case a multi-query executes more than one Slice query per multi-query execution then those Slice queries will be grouped per same key sets and the new getMultiSlices query will be executed with groups of different SliceQuery for the same sets of keys.

Notice, if the storage doesn't have multiQuery feature enabled - the method won't be used. Hence, it's not necessary to implement it.

In case storage backend has multiQuery feature enabled, then it is highly recommended to overwrite the default (non-optimized) implementation and optimize this method execution to execute all the slice queries for all the requested keys in the shortest time possible (for example, by using asynchronous programming, slice queries grouping, multi-threaded execution, or any other technique which is efficient for the respective storage adapter).

Added possibility to group multiple slice queries together via CQL storage backend

Starting from JanusGraph 1.0.0 CQL storage implementation now groups queries which fetch properties with Cardinality.SINGLE together into the same CQL query. The behaviour can be disabled by setting configuration storage.cql.grouping.slice-allowed = false.

CQL storage implementation has also ability to group queries of different partition keys together if they belong to the same token range or same replicas set (grouping strategy is available via storage.cql.grouping.keys-class configuration option). This behaviour is disabled by default, but can be enabled via storage.cql.grouping.keys-allowed = true. Please note that enabling the key grouping feature can increase the return size of related queries. This is due to each row in a CQL query, which encompasses grouped keys, needing to return its specific key. Moreover, it could potentially lead to less balanced load on the storage cluster. However, it reduces the amount of CQL queries sent which may positively influence throughput in some cases as well as pricing point of some Serverless deployments. We recommend to benchmark each use-case before enabling keys grouping (storage.cql.grouping.keys-allowed).

Notice, keys grouping (storage.cql.grouping.keys-allowed) can be used only with storage backends which support PER PARTITION LIMIT. As such, this feature can't be used with Amazon Keyspaces because it doesn't support PER PARTITION LIMIT.

Different storage backends may also have restriction set on maximum keys which can be provided via IN operator (which is needed for grouping). It is required to ensure that storage.cql.grouping.keys-limit or storage.cql.grouping.slice-limit is less than or equal to the restriction provided via the storage backend. On ScyllaDB it's possible to configure this restriction using max-partition-key-restrictions-per-query configuration option (default to 100). On AstraDB side it is needed to be asked on AstraDB side to be changed via partition_keys_in_select_failure_threshold and in_select_cartesian_product_failure_threshold threshold configurations (https://docs.datastax.com/en/astra-serverless/docs/plan/planning.html#_cassandra_yaml) which are set to 20 and 25 by default.

See additional properties to control grouping configurations under the namespace storage.cql.grouping.

Removal of deprecated classes/methods/functionalities
Methods
  • JanusGraphIndexQuery.vertices replaced by JanusGraphIndexQuery.vertexStream
  • JanusGraphIndexQuery.edges replaced by JanusGraphIndexQuery.edgeStream
  • JanusGraphIndexQuery.properties replaced by JanusGraphIndexQuery.propertyStream
  • IndexQueryBuilder.vertices replaced by IndexQueryBuilder.vertexStream
  • IndexQueryBuilder.edges replaced by IndexQueryBuilder.edgeStream
  • IndexQueryBuilder.properties replaced by IndexQueryBuilder.propertyStream
  • IndexTransaction.query replaced by IndexTransaction.queryStream
Classes/Interfaces
  • EdgeLabelDefinition class
  • PropertyKeyDefinition class
  • RelationTypeDefinition class
  • SchemaContainer class
  • SchemaElementDefinition class
  • SchemaProvider interface
  • VertexLabelDefinition class
  • JanusGraphId class
  • AllEdgesIterable class
  • AllEdgesIterator class
  • ConcurrentLRUCache class
  • PriorityQueue class
  • RemovableRelationIterable class
  • RemovableRelationIterator class
  • ImmutableConfiguration class

Version 0.6.4 (Release Date: October 14, 2023)

<dependency>
    <groupId>org.janusgraph</groupId>
    <artifactId>janusgraph-core</artifactId>
    <version>0.6.4</version>
</dependency>
compile "org.janusgraph:janusgraph-core:0.6.4"

Tested Compatibility:

  • Apache Cassandra 3.0.14, 3.11.10
  • Apache HBase 1.6.0, 2.2.7
  • Oracle BerkeleyJE 7.5.11
  • Elasticsearch 6.0.1, 6.6.0, 7.14.0
  • Apache Lucene 8.9.0
  • Apache Solr 7.7.2, 8.11.0
  • Apache TinkerPop 3.5.7
  • Java 1.8

Changes

For more information on features and bug fixes in 0.6.4, see the GitHub milestone:

Assets

Upgrade Instructions

Default logging library changed to Reload4j

The default logging library used in the pre-packaged distribution has been changed in version 0.6.3 by accident from Log4j to Logback. While this change meant that some security issues of Log4j were avoided, it was also a breaking change that was not intended. This resulted in only warnings being logged by default and also that a Log4j config file was ignored. To fix this breaking change, we change the default logging library in this release to Reload4j which is completely compatible with Log4j, but fixes the security issues of Log4j. This means that Log4j config files will continue to work with this version.

Note that this only applies to JanusGraph 0.6. JanusGraph 1.0.0 uses Log4j2 by default.

Version 0.6.3 (Release Date: February 18, 2023)

<dependency>
    <groupId>org.janusgraph</groupId>
    <artifactId>janusgraph-core</artifactId>
    <version>0.6.3</version>
</dependency>
compile "org.janusgraph:janusgraph-core:0.6.3"

Tested Compatibility:

  • Apache Cassandra 3.0.14, 3.11.10
  • Apache HBase 1.6.0, 2.2.7
  • Oracle BerkeleyJE 7.5.11
  • Elasticsearch 6.0.1, 6.6.0, 7.14.0
  • Apache Lucene 8.9.0
  • Apache Solr 7.7.2, 8.11.0
  • Apache TinkerPop 3.5.5
  • Java 1.8

Note

Google Bigtable was removed from this list because there is no automatic testing in place specifically for that backend. Since the adapter for Bigtable is however just using the HBase adapter, it is also covered by the tests for HBase.

We invite anyone who is interested in the Bigtable storage adapter to help with this by contributing so that the tests for HBase are also automatically executed for Bigtable. More information can be found in this GitHub issue: janusgraph/janusgraph#415.

Changes

For more information on features and bug fixes in 0.6.3, see the GitHub milestone:

Assets

Version 0.6.2 (Release Date: May 31, 2022)

<dependency>
    <groupId>org.janusgraph</groupId>
    <artifactId>janusgraph-core</artifactId>
    <version>0.6.2</version>
</dependency>
compile "org.janusgraph:janusgraph-core:0.6.2"

Tested Compatibility:

  • Apache Cassandra 3.0.14, 3.11.10
  • Apache HBase 1.6.0, 2.2.7
  • Google Bigtable 1.3.0, 1.4.0, 1.5.0, 1.6.0, 1.7.0, 1.8.0, 1.9.0, 1.10.0, 1.11.0, 1.14.0
  • Oracle BerkeleyJE 7.5.11
  • Elasticsearch 6.0.1, 6.6.0, 7.14.0
  • Apache Lucene 8.9.0
  • Apache Solr 7.7.2, 8.9.0
  • Apache TinkerPop 3.5.3
  • Java 1.8

Changes

For more information on features and bug fixes in 0.6.2, see the GitHub milestone:

Assets

Version 0.6.1 (Release Date: January 18, 2022)

<dependency>
    <groupId>org.janusgraph</groupId>
    <artifactId>janusgraph-core</artifactId>
    <version>0.6.1</version>
</dependency>
compile "org.janusgraph:janusgraph-core:0.6.1"

Tested Compatibility:

  • Apache Cassandra 3.0.14, 3.11.10
  • Apache HBase 1.6.0, 2.2.7
  • Google Bigtable 1.3.0, 1.4.0, 1.5.0, 1.6.0, 1.7.0, 1.8.0, 1.9.0, 1.10.0, 1.11.0, 1.14.0
  • Oracle BerkeleyJE 7.5.11
  • Elasticsearch 6.0.1, 6.6.0, 7.14.0
  • Apache Lucene 8.9.0
  • Apache Solr 7.7.2, 8.9.0
  • Apache TinkerPop 3.5.1
  • Java 1.8

Changes

For more information on features and bug fixes in 0.6.1, see the GitHub milestone:

Assets

Upgrade Instructions

GraphManager changed to JanusGraphManager

A GraphManager is used to instantiate graph instances. JanusGraph Server has used the DefaultGraphManager from TinkerPop for this by default if no other GraphManager was specified in the JanusGraph Server YAML config file. The behavior of this DefaultGraphManager was changed in TinkerPop 3.5.0 which is included in JanusGraph 0.6.0 in how it parses config values, making it impossible to provide comma separated values, e.g., to specify multiple hostnames for the storage backend. The JanusGraphManager does not have this limitation which is why it is now configured as the GraphManager in the JanusGraph Server config files:

[...]
channelizer: org.apache.tinkerpop.gremlin.server.channel.WebSocketChannelizer
graphManager: org.janusgraph.graphdb.management.JanusGraphManager
graphs: {
  graph: conf/janusgraph-berkeleyje-es.properties
}
[...]

If you however want to continue using the DefaultGraphManager, then you can simply remove the setting again or change it to the TinkerPop GraphManager that has been the default before: org.apache.tinkerpop.gremlin.server.util.DefaultGraphManager.

Version 0.6.0 (Release Date: September 3, 2021)

<dependency>
    <groupId>org.janusgraph</groupId>
    <artifactId>janusgraph-core</artifactId>
    <version>0.6.0</version>
</dependency>
compile "org.janusgraph:janusgraph-core:0.6.0"

Tested Compatibility:

  • Apache Cassandra 3.0.14, 3.11.10
  • Apache HBase 1.6.0, 2.2.7
  • Google Bigtable 1.3.0, 1.4.0, 1.5.0, 1.6.0, 1.7.0, 1.8.0, 1.9.0, 1.10.0, 1.11.0, 1.14.0
  • Oracle BerkeleyJE 7.5.11
  • Elasticsearch 6.0.1, 6.6.0, 7.14.0
  • Apache Lucene 8.9.0
  • Apache Solr 7.7.2, 8.9.0
  • Apache TinkerPop 3.5.3
  • Java 1.8

Changes

For more information on features and bug fixes in 0.6.0, see the GitHub milestone:

Assets

Upgrade Instructions

Experimental support for Amazon Keyspaces

Amazon Keyspaces is a serverless managed Apache Cassandra-compatible database service provided by Amazon. See Deploying on Amazon Keyspaces for more details.

Breaking change for Configuration objects

Prior to JanusGraph 0.6.0, Configuration objects were from the Apache commons-configuration library. To comply with the TinkerPop change, JanusGraph now uses the commons-configuration2 library. A typical usage of configuration object is to create configuration using ConfigurationGraphFactory. Now you would need to use the new configuration2 library. Please refer to the commons-configuration 2.0 migration guide for details. Note that this very likely does not affect gremlin console usage, since the new library is auto-imported, and the basic APIs remain the same. For java code usage, you need to import configuration2 library rather than the old configuration library.

Breaking change for gremlin server configs

scriptEvaluationTimeout is renamed to evaluationTimeout. You can refer to conf/gremlin-server/gremlin-server.yaml for example.

Breaking change for gremlin EventStrategy usage

If you are using EventStrategy, please note that now you need to register it every time you start a new transaction. An example is available at ThreadLocalTxLeakTest::eventListenersCanBeReusedAcrossTx See more background of this breaking change in this pull request.

Disable smart-limit by default and change HARD_MAX_LIMIT

Prior to 0.6.0, smart-limit is enabled by default. It tries to guess a small limit for each graph centric query (e.g. g.V().has("prop", "value")) internally, and if more results are required by user, it queries backend again with a larger limit, and repeats until either results are exhausted or user stops the query. However, this is not the same as paging mechanism. All interim results will be fetched again in next round, making the whole query costly. Even worse, if your data backend does not return results in a consistent order, then some entries might be missing in the final results. Until JanusGraph can fully utilize the paging capacity provided by backends (e.g. Elasticsearch scroll), this option is recommended to be turned off. The exception is when you have a large number of results but you only need a few of them, then enabling smart-limit can reduce latency and memory usage. An example would be:

Iterator<Vertex> iter = graph.traversal().V().has("prop", "value");
while (iter.hasNext()) {
    Vertex v = iter.next();
    if (canStop()) break;
}

Prior to 0.6.0, even if smart-limit is disabled, JanusGraph adds a HARD_MAX_LIMIT that is equivalent to 100,000 to avoid fetching too many results at a time. This limit is now configurable, and by default, it's Integer.MAX_VALUE which can be interpreted as no limit.

Add experimental support for Java 11

We started to work on support for Java 11. We would like to get feedback, if everything is working as expected after upgrading to Java 11.

Removal of LoggingSchemaMaker

The schema.default=logging option is not valid anymore. Use schema.default=default and schema.logging=true options together to make application behaviour unaltered, if you are using LoggingSchemaMaker.

Replacing the server startup script is replaced

The gremlin-server.sh is placed by janusgraph-server.sh. The janusgraph-server.sh brings some new functionality such as easy configuration of Java options using the jvm.options file.

The jvm.options file contains some default configurations for JVM based on Cassandra's JVM configurations, Elasticsearch and the old gremlin-server.sh.

Serialization of JanusGraph predicates has changed

The serialization of JanusGraph predicates has changed in this version for both GraphSON and Gryo. The newest version of the JanusGraph Driver requires a JanusGraph Server version of 0.6.0 and above. The server includes a fallback for clients with an older driver to make the upgrade to version 0.6.0 easier. This means that the server can be upgraded first without having to update all clients at the same time. The fallback will however be removed in a future version of JanusGraph so clients should also be upgraded.

GraphBinary is now supported

GraphBinary is a new binary serialization format from TinkerPop that supersedes Gryo and it will eventually also replace GraphSON. GraphBinary is language independent and has a low serialization overhead which results in an improved performance.

If you want to use GraphBinary, you have to add following to the gremlin-server.yaml after the keyword serializers. This will add the support on the server site.

    - { className: org.apache.tinkerpop.gremlin.driver.ser.GraphBinaryMessageSerializerV1, 
        config: { ioRegistries: [org.janusgraph.graphdb.tinkerpop.JanusGraphIoRegistry] }}
    - { className: org.apache.tinkerpop.gremlin.driver.ser.GraphBinaryMessageSerializerV1, 
        config: { serializeResultToString: true }}

Note

The java driver is the only driver that currently supports GraphBinary, see Connecting to JanusGraph using Java.

Note

Version 1.0.0 moves everything under org.apache.tinkerpop.gremlin.driver.ser package to org.apache.tinkerpop.gremlin.util.ser package.

Note

Version 1.0.0 adds a breaking change to GraphBinary for Geoshape serialization, see the 1.0.0 changelog for more information.

New index selection algorithm

In version 0.6.0, the index selection algorithm has changed. If the number of possible indexes for a query is small enough, the new algorithm will perform an exhaustive search to minimize the number of indexes which need to be queried. The default limit is set to 10. In order to maintain the old selection algorithm regardless of the available indexes, set the key query.index-select-threshold to 0. For more information, see Configuration Reference

Removal of Cassandra Thrift support

Thrift will be completely removed in Cassandra 4. All deprecated Cassandra Thrift backends were removed in JanusGraph 0.6.0. We already added support for CQL in JanusGraph 0.2.0 and we have been encouraging users to switch from Thrift to CQL since version 0.2.1.

This means that the following backends were removed: cassandrathrift, cassandra, astyanax, and embeddedcassandra. Users who still use one of these Thrift backends should migrate to CQL. Our migration guide explains the necessary steps for this. The option to run Cassandra embedded in the same JVM as JanusGraph is however no longer supported with CQL.

Note

The source code for the Thrift backends will be moved into a dedicated repository. While we do not support them any more, users can still use them if they for some reason cannot migrate to CQL.

Drop support for Cassandra 2

With the release of Cassandra 4, the support of Cassandra 2 will be dropped. Therefore, you should upgrade to Cassandra 3 or higher.

Note

Cassandra 3 and higher doesn't support compact storage. If you have activated or never changed the value of storage.cql.storage-compact=true, during the upgrade process you have to ensure your data is correctly migrated.

Introduction of a JanusGraph Server startup class as a replacement for Gremlin Server startup

The gremlin-server.sh and the janusgraph.sh are configured to use the new JanusGraph startup class. This new class introduces a default set of TinkerPop Serializers if no serializers are configured in the gremlin-server.yaml. Furthermore, JanusGraph will log the version of JanusGraph and TinkerPop after a shiny new JanusGraph header.

Note

If you have a custom script to startup JanusGraph, you propably would like to replace the Gremlin Server class with JanusGraph Server class:

org.apache.tinkerpop.gremlin.server.GremlinServer => org.janusgraph.graphdb.server.JanusGraphServer

Drop support for Ganglia metrics

We are dropping Ganglia as we are using dropwizard for metrics. Dropwizard did drop Ganglia in the newest major version.

DataStax cassandra driver upgrade from 3.9.0 to 4.13.0

All DataStax cassandra driver metrics are now disabled by default. To enable DataStax driver metrics you need to provide a list of Session level metrics and / or Node level metrics you want to enable. To provide a list of enabled metrics, you can use the next configuration options: storage.cql.metrics.session-enabled and storage.cql.metrics.node-enabled. Notice, DataStax metrics are enabled only when basic metrics are enabled (i.e. metrics.enabled = true). See configuration references storage.cql.metrics for additional DataStax metrics configuration.

An example configuration which enables some CQL Session level and Node level metrics reporting by JMX:

metrics.enabled=true
metrics.jmx.enabled=true
metrics.jmx.domain=com.datastax.oss.driver
metrics.jmx.agentid=agent
storage.cql.metrics.session-enabled=bytes-sent,bytes-received,connected-nodes,cql-requests,throttling.delay
storage.cql.metrics.node-enabled=pool.open-connections,pool.available-streams,bytes-sent,cql-messages

See advanced.metrics.session.enabled and advanced.metrics.node.enabled sections in DataStax Metrics Configuration for a complete list of available Session level and Node level metrics.

Due to driver upgrade the next cql configuration options have been removed:

  • local-core-connections-per-host
  • remote-core-connections-per-host
  • local-max-requests-per-connection
  • remote-max-requests-per-connection
  • cluster-name

storage.connection-timeout is now used to control initial connection timeout to CQL storage and not request timeouts. Please, use storage.cql.request-timeout to configure request timeouts instead.

New cql configuration options should be used for upgrade:

  • max-requests-per-connection
  • session-name

storage.cql.local-datacenter is mandatory now and defaults to datacenter1.

See more new cql configuration options in configuration references under storage.cql section.

Automatic configurations of dynamic graph binding

If the JanusGraphManager is configured, dynamic graph binding will be setup automatically, see Dynamic Graphs.

Note

Breaking changes in the config of the gremlin-server.yaml.

Following, classes are removed and have to be replaced by tinkerpop equivalent:

removed class replacement class
org.janusgraph.channelizers.JanusGraphWebSocketChannelizer org.apache.tinkerpop.gremlin.server.channel.WebSocketChannelizer
org.janusgraph.channelizers.JanusGraphHttpChannelizer org.apache.tinkerpop.gremlin.server.channel.HttpChannelizer
org.janusgraph.channelizers.JanusGraphNioChannelizer org.apache.tinkerpop.gremlin.server.channel.NioChannelizer
org.janusgraph.channelizers.JanusGraphWsAndHttpChannelizer org.apache.tinkerpop.gremlin.server.channel.WsAndHttpChannelizer
Breaking change Lucene and Solr fuzzy predicates

The text predicates text.textFuzzy and text.textContainsFuzzy have been updated in both the Lucene and Solr indexing backends to align with JanusGraph and Elastic. These predicates now inspect the query length to determine the Levenshtein distance, where previously they used the backend's default max distance of 2:

  • 0 for strings of one or two characters (exact match)
  • 1 for strings of three, four or five characters
  • 2 for strings of more than five characters

Change Matrix:

text query previous result new result
ah ah true true
ah ai true false
hop hop true true
hop hap true true
hop hoop true true
hop hooop true false
surprises surprises true true
surprises surprizes true true
surprises surpprises true true
surprises surpprisess false false

Version 0.5.3 (Release Date: December 24, 2020)

<dependency>
    <groupId>org.janusgraph</groupId>
    <artifactId>janusgraph-core</artifactId>
    <version>0.5.3</version>
</dependency>
compile "org.janusgraph:janusgraph-core:0.5.3"

Tested Compatibility:

  • Apache Cassandra 2.2.10, 3.0.14, 3.11.0
  • Apache HBase 1.2.6, 1.3.1, 1.4.10, 2.1.5
  • Google Bigtable 1.3.0, 1.4.0, 1.5.0, 1.6.0, 1.7.0, 1.8.0, 1.9.0, 1.10.0, 1.11.0, 1.14.0
  • Oracle BerkeleyJE 7.5.11
  • Elasticsearch 6.0.1, 6.6.0, 7.6.2
  • Apache Lucene 7.0.0
  • Apache Solr 7.0.0
  • Apache TinkerPop 3.4.6
  • Java 1.8

Changes

For more information on features and bug fixes in 0.5.3, see the GitHub milestone:

Assets

Version 0.5.2 (Release Date: May 3, 2020)

<dependency>
    <groupId>org.janusgraph</groupId>
    <artifactId>janusgraph-core</artifactId>
    <version>0.5.2</version>
</dependency>
compile "org.janusgraph:janusgraph-core:0.5.2"

Tested Compatibility:

  • Apache Cassandra 2.2.10, 3.0.14, 3.11.0
  • Apache HBase 1.2.6, 1.3.1, 1.4.10, 2.1.5
  • Google Bigtable 1.3.0, 1.4.0, 1.5.0, 1.6.0, 1.7.0, 1.8.0, 1.9.0, 1.10.0, 1.11.0, 1.14.0
  • Oracle BerkeleyJE 7.5.11
  • Elasticsearch 6.0.1, 6.6.0, 7.6.2
  • Apache Lucene 7.0.0
  • Apache Solr 7.0.0
  • Apache TinkerPop 3.4.6
  • Java 1.8

For more information on features and bug fixes in 0.5.2, see the GitHub milestone:

Upgrade Instructions

ElasticSearch index store names cache now enabled for any amount of indexes per store

In JanusGraph version 0.5.0 and 0.5.1 all ElasticSearch index store names are cached for efficient index store name retrieval and the cache is disabled if there are more than 50000 indexes available per index store. From JanusGraph version 0.5.2 index store names cache isn't limited to 50000 but instead can be disabled by using a new added parameter enable_index_names_cache. It is still recommended to disable index store names cache if more than 50000 indexes are used per index store.

Version 0.5.1 (Release Date: March 25, 2020)

<dependency>
    <groupId>org.janusgraph</groupId>
    <artifactId>janusgraph-core</artifactId>
    <version>0.5.1</version>
</dependency>
compile "org.janusgraph:janusgraph-core:0.5.1"

Tested Compatibility:

  • Apache Cassandra 2.2.10, 3.0.14, 3.11.0
  • Apache HBase 1.2.6, 1.3.1, 1.4.10, 2.1.5
  • Google Bigtable 1.3.0, 1.4.0, 1.5.0, 1.6.0, 1.7.0, 1.8.0, 1.9.0, 1.10.0, 1.11.0, 1.14.0
  • Oracle BerkeleyJE 7.5.11
  • Elasticsearch 6.0.1, 6.6.0, 7.6.1
  • Apache Lucene 7.0.0
  • Apache Solr 7.0.0
  • Apache TinkerPop 3.4.6
  • Java 1.8

For more information on features and bug fixes in 0.5.1, see the GitHub milestone:

Upgrade Instructions

Two Distributed package is splitted into two version

The default version of the distribution package does no longer contain the janusgraph.sh. This includes a packaged version of cassandra and elasticsearch. If you want to have janusgraph.sh, you have to download distribution with the suffix -full.

Gremlin Server distributed with the release uses inmemory storage backend and no search backend by default

Gremlin Server is by default configured for the inmemory storage backend and no search backend when started with bin/gremlin-server.sh.
You can provide configuration for another storage backend and/or search backend by providing a path to the appropriate configuration as a second parameter (./bin/gremlin-server.sh ./conf/gremlin-server/[...].yaml).

Version 0.5.0 (Release Date: March 10, 2020)

<dependency>
    <groupId>org.janusgraph</groupId>
    <artifactId>janusgraph-core</artifactId>
    <version>0.5.0</version>
</dependency>
compile "org.janusgraph:janusgraph-core:0.5.0"

Tested Compatibility:

  • Apache Cassandra 2.2.10, 3.0.14, 3.11.0
  • Apache HBase 1.2.6, 1.3.1, 1.4.10, 2.1.5
  • Google Bigtable 1.3.0, 1.4.0, 1.5.0, 1.6.0, 1.7.0, 1.8.0, 1.9.0, 1.10.0, 1.11.0
  • Oracle BerkeleyJE 7.5.11
  • Elasticsearch 6.0.1, 6.6.0, 7.6.1
  • Apache Lucene 7.0.0
  • Apache Solr 7.0.0
  • Apache TinkerPop 3.4.6
  • Java 1.8

For more information on features and bug fixes in 0.5.0, see the GitHub milestone:

Upgrade Instructions

Distributed package is renamed

The distribution has no longer the suffix -hadoop2.

Reorder dependency of Hadoop

Hadoop is now a dependency of supported backends. Therefore, MapReduceIndexJobs is now split up into different classes:

Old Function New Function
MapReduceIndexJobs.cassandraRepair CassandraMapReduceIndexJobsUtils.repair
MapReduceIndexJobs.cassandraRemove CassandraMapReduceIndexJobsUtils.remove
MapReduceIndexJobs.cqlRepair CqlMapReduceIndexJobsUtils.repair
MapReduceIndexJobs.cqlRemove CqlMapReduceIndexJobsUtils.remove
MapReduceIndexJobs.hbaseRepair HBaseMapReduceIndexJobsUtils.repair
MapReduceIndexJobs.hbaseRemove HBaseMapReduceIndexJobsUtils.remove

Note

Now, you can easily support for any backend.

Warning

Cassandra3InputFormat is replaced by CqlInputFormat

ElasticSearch: Upgrade from 6.6.0 to 7.6.1 and drop support for 5.x version

The ElasticSearch version has been changed to 7.6.1 which removes support for max-retry-timeout option. That is why this option no longer available in JanusGraph. Users should be aware that by default JanusGraph setups maximum open scroll contexts to maximum value of 2147483647 with the parameter setup-max-open-scroll-contexts for ElasticSearch 7.y. This option can be disabled and updated manually in ElasticSearch but you should be aware that ElasticSearch starting from version 7 has a default limit of 500 opened contexts which most likely be reached by the normal usage of JanusGraph with ElasticSearch. By default deprecated mappings are disabled in ElasticSearch version 7. If you are upgrading your ElasticSearch index backend to version 7 from lower versions, it is recommended to reindex your JanusGraph indices to not use mappings. If you are unable to reindex your indices you may setup parameter use-mapping-for-es7 to true which will tell JanusGraph to use mapping types for ElasticSearch version 7. Due to the drop of support for 5.x version, deprecated multi-type indices are no more supported. Parameter use-deprecated-multitype-index is no more supported by JanusGraph.

BerkeleyDB

BerkeleyDB storage configured with SHARED_CACHE for better memory usage.

Default logging location has changed

If you are using janusgraph.sh to start your instance, the default logging has been changed from log to logs

In-Memory backend moved into dedicated module

The built-in in-memory backend has been moved into a dedicated module. Users who use it for instance in tests, have to explicitly declare it as a dependency:

<dependency>
    <groupId>org.janusgraph</groupId>
    <artifactId>janusgraph-inmemory</artifactId>
    <scope>test</scope>
    <version>0.5.0</version>
</dependency>
implementation 'org.janusgraph:janusgraph-inmemory:0.5.0'

Version 0.4.1 (Release Date: January 14, 2020)

<dependency>
    <groupId>org.janusgraph</groupId>
    <artifactId>janusgraph-core</artifactId>
    <version>0.4.1</version>
</dependency>
compile "org.janusgraph:janusgraph-core:0.4.1"

Tested Compatibility:

  • Apache Cassandra 2.2.10, 3.0.14, 3.11.0
  • Apache HBase 1.2.6, 1.3.1, 1.4.10, 2.1.5
  • Google Bigtable 1.3.0, 1.4.0, 1.5.0, 1.6.0, 1.7.0, 1.8.0, 1.9.0, 1.10.0, 1.11.0
  • Oracle BerkeleyJE 7.5.11
  • Elasticsearch 5.6.14, 6.0.1, 6.6.0
  • Apache Lucene 7.0.0
  • Apache Solr 7.0.0
  • Apache TinkerPop 3.4.4
  • Java 1.8

For more information on features and bug fixes in 0.4.1, see the GitHub milestone:

Upgrade Instructions

TinkerPop: Upgrade from 3.4.1 to 3.4.4

Adding multiple values in the same query to a new vertex property without explicitly defined type (i.e. using Automatic Schema Maker to create a property type) requires explicit usage of VertexProperty.Cardinality for each call (only for the first query which defines a property) if the VertexProperty.Cardinality is different than VertexProperty.Cardinality.single.

Version 0.4.0 (Release Date: July 1, 2019)

Legacy documentation: https://old-docs.janusgraph.org/0.4.0/index.html

<dependency>
    <groupId>org.janusgraph</groupId>
    <artifactId>janusgraph-core</artifactId>
    <version>0.4.0</version>
</dependency>
compile "org.janusgraph:janusgraph-core:0.4.0"

Tested Compatibility:

  • Apache Cassandra 2.2.10, 3.0.14, 3.11.0
  • Apache HBase 1.2.6, 1.3.1, 1.4.10, 2.1.5
  • Google Bigtable 1.3.0, 1.4.0, 1.5.0, 1.6.0, 1.7.0, 1.8.0, 1.9.0, 1.10.0, 1.11.0
  • Oracle BerkeleyJE 7.5.11
  • Elasticsearch 5.6.14, 6.0.1, 6.6.0
  • Apache Lucene 7.0.0
  • Apache Solr 7.0.0
  • Apache TinkerPop 3.4.1
  • Java 1.8

For more information on features and bug fixes in 0.4.0, see the GitHub milestone:

Upgrade Instructions

HBase: Upgrade from 1.2 to 2.1

The version of HBase that is included in the distribution of JanusGraph was upgraded from 1.2.6 to 2.1.5. HBase 2.x client is not fully backward compatible with HBase 1.x server. Users who operate their own HBase version 1.x cluster may need to upgrade their cluster to version 2.x. Optionally users may build their own distribution of JanusGraph which includes HBase 1.x from source with the maven flags -Dhbase.profile -Phbase1.

Cassandra: Upgrade from 2.1 to 2.2

The version of Cassandra that is included in the distribution of JanusGraph was upgraded from 2.1.20 to 2.2.13. Refer to the upgrade documentation of Cassandra for detailed instructions to perform this upgrade. Users who operate their own Cassandra cluster instead of using Cassandra distributed together with JanusGraph are not affected by this upgrade. This also does not change the different versions of Cassandra that are supported by JanusGraph (see <> for a detailed list of the supported versions).

BerkeleyDB : Upgrade from 7.4 to 7.5

The BerkeleyDB version has been updated, and it contains changes to the file format stored on disk (see the BerkeleyDB changelog for reference). This file format change is forward compatible with previous versions of BerkeleyDB, so existing graph data stored with JanusGraph can be read in. However, once the data has been read in with the newer version of BerkeleyDB, those files can no longer be read by the older version. Users are encouraged to backup the BerkeleyDB storage directory before attempting to use it with the JanusGraph release.

Solr: Compatible Lucene version changed from 5.0.0 to 7.0.0 in distributed config

The JanusGraph distribution contains a solrconfig.xml file that can be used to configure Solr. The value luceneMatchVersion in this config that tells Solr to behave according to that Lucene version was changed from 5.0.0 to 7.0.0 as that is the default version currently used by JanusGraph. Users should generally set this value to the version of their Solr installation. If the config distributed by JanusGraph is used for an existing Solr installation that used a lower version before (like 5.0.0 from a previous versions of this file), it is highly recommended that a re-indexing is performed.

Version 0.3.3 (Release Date: January 11, 2020)

<dependency>
    <groupId>org.janusgraph</groupId>
    <artifactId>janusgraph-core</artifactId>
    <version>0.3.3</version>
</dependency>
compile "org.janusgraph:janusgraph-core:0.3.3"

Tested Compatibility:

  • Apache Cassandra 2.1.20, 2.2.10, 3.0.14, 3.11.0
  • Apache HBase 1.2.6, 1.3.1, 1.4.4
  • Google Bigtable 1.0.0, 1.1.2, 1.2.0, 1.3.0, 1.4.0
  • Oracle BerkeleyJE 7.4.5
  • Elasticsearch 1.7.6, 2.4.6, 5.6.5, 6.0.1
  • Apache Lucene 7.0.0
  • Apache Solr 5.5.4, 6.6.1, 7.0.0
  • Apache TinkerPop 3.3.3
  • Java 1.8

For more information on features and bug fixes in 0.3.3, see the GitHub milestone:

Version 0.3.2 (Release Date: June 16, 2019)

Legacy documentation: https://old-docs.janusgraph.org/0.3.2/index.html

<dependency>
    <groupId>org.janusgraph</groupId>
    <artifactId>janusgraph-core</artifactId>
    <version>0.3.2</version>
</dependency>
compile "org.janusgraph:janusgraph-core:0.3.2"

Tested Compatibility:

  • Apache Cassandra 2.1.20, 2.2.10, 3.0.14, 3.11.0
  • Apache HBase 1.2.6, 1.3.1, 1.4.4
  • Google Bigtable 1.0.0, 1.1.2, 1.2.0, 1.3.0, 1.4.0
  • Oracle BerkeleyJE 7.4.5
  • Elasticsearch 1.7.6, 2.4.6, 5.6.5, 6.0.1
  • Apache Lucene 7.0.0
  • Apache Solr 5.5.4, 6.6.1, 7.0.0
  • Apache TinkerPop 3.3.3
  • Java 1.8

For more information on features and bug fixes in 0.3.2, see the GitHub milestone:

Version 0.3.1 (Release Date: October 2, 2018)

Legacy documentation: https://old-docs.janusgraph.org/0.3.1/index.html

<dependency>
    <groupId>org.janusgraph</groupId>
    <artifactId>janusgraph-core</artifactId>
    <version>0.3.1</version>
</dependency>
compile "org.janusgraph:janusgraph-core:0.3.1"

Tested Compatibility:

  • Apache Cassandra 2.1.20, 2.2.10, 3.0.14, 3.11.0
  • Apache HBase 1.2.6, 1.3.1, 1.4.4
  • Google Bigtable 1.0.0, 1.1.2, 1.2.0, 1.3.0, 1.4.0
  • Oracle BerkeleyJE 7.4.5
  • Elasticsearch 1.7.6, 2.4.6, 5.6.5, 6.0.1
  • Apache Lucene 7.0.0
  • Apache Solr 5.5.4, 6.6.1, 7.0.0
  • Apache TinkerPop 3.3.3
  • Java 1.8

For more information on features and bug fixes in 0.3.1, see the GitHub milestone:

Version 0.3.0 (Release Date: July 31, 2018)

Legacy documentation: https://old-docs.janusgraph.org/0.3.0/index.html

<dependency>
    <groupId>org.janusgraph</groupId>
    <artifactId>janusgraph-core</artifactId>
    <version>0.3.0</version>
</dependency>
compile "org.janusgraph:janusgraph-core:0.3.0"

Tested Compatibility:

  • Apache Cassandra 2.1.20, 2.2.10, 3.0.14, 3.11.0
  • Apache HBase 1.2.6, 1.3.1, 1.4.4
  • Google Bigtable 1.0.0, 1.1.2, 1.2.0, 1.3.0, 1.4.0
  • Oracle BerkeleyJE 7.4.5
  • Elasticsearch 1.7.6, 2.4.6, 5.6.5, 6.0.1
  • Apache Lucene 7.0.0
  • Apache Solr 5.5.4, 6.6.1, 7.0.0
  • Apache TinkerPop 3.3.3
  • Java 1.8

For more information on features and bug fixes in 0.3.0, see the GitHub milestone:

Upgrade Instructions

Important

You should back-up your data prior to attempting an upgrade! Also please note that once an upgrade has been completed you will no longer be able to connect to your graph with client versions prior to 0.3.0.

JanusGraph 0.3.0 implements Schema Constraints which made it necessary to also introduce the concept of a schema version. There is a check to prevent client connections that either expect a different schema version or have no concept of a schema version. To perform an upgrade, the configuration option graph.allow-upgrade=true must be set on each graph you wish to upgrade. The graph must be opened with a 0.3.0 or greater version of JanusGraph since older versions have no concept of graph.storage-version and will not allow for it to be set.

Example excerpt from janusgraph.properties file

# JanusGraph configuration sample: Cassandra over a socket
#
# This file connects to a Cassandra daemon running on localhost via
# Thrift.  Cassandra must already be started before starting JanusGraph
# with this file.

# This option should be removed as soon as the upgrade is complete. Otherwise if this file
# is used in the future to connect to a different graph it could cause an unintended upgrade.
graph.allow-upgrade=true

gremlin.graph=org.janusgraph.core.JanusGraphFactory

# The primary persistence provider used by JanusGraph.  This is required.
# It should be set one of JanusGraph's built-in shorthand names for its
# standard storage backends (shorthands: berkeleyje, cassandrathrift,
# cassandra, astyanax, embeddedcassandra, cql, hbase, inmemory) or to the
# full package and classname of a custom/third-party StoreManager
# implementation.
#
# Default:    (no default value)
# Data Type:  String
# Mutability: LOCAL
storage.backend=cassandrathrift

# The hostname or comma-separated list of hostnames of storage backend
# servers.  This is only applicable to some storage backends, such as
# cassandra and hbase.
#
# Default:    127.0.0.1
# Data Type:  class java.lang.String[]
# Mutability: LOCAL
storage.hostname=127.0.0.1

If graph.allow-upgrade is set to true on a graph graph.storage-version and graph.janusgraph-version will automatically be upgraded to match the version level of the server, or local client, that is opening the graph. You can verify the upgrade was successful by opening the management API and validating the values of graph.storage-version and graph.janusgraph-version.

Once the storage version has been set you should remove graph.allow-upgrade=true from your properties file and reopen your graph to ensure that the upgrade was successful.

Version 0.2.3 (Release Date: May 21, 2019)

Legacy documentation: https://old-docs.janusgraph.org/0.2.3/index.html

<dependency>
    <groupId>org.janusgraph</groupId>
    <artifactId>janusgraph-core</artifactId>
    <version>0.2.3</version>
</dependency>
compile "org.janusgraph:janusgraph-core:0.2.3"

Tested Compatibility:

  • Apache Cassandra 2.1.20, 2.2.10, 3.0.14, 3.11.0
  • Apache HBase 0.98.24-hadoop2, 1.2.6, 1.3.1
  • Google Bigtable 1.0.0
  • Oracle BerkeleyJE 7.3.7
  • Elasticsearch 1.7.6, 2.4.6, 5.6.5, 6.0.1
  • Apache Lucene 7.0.0
  • Apache Solr 5.5.4, 6.6.1, 7.0.0
  • Apache TinkerPop 3.2.9
  • Java 1.8

For more information on features and bug fixes in 0.2.3, see the GitHub milestone:

Version 0.2.2 (Release Date: October 9, 2018)

Legacy documentation: https://old-docs.janusgraph.org/0.2.2/index.html

<dependency>
    <groupId>org.janusgraph</groupId>
    <artifactId>janusgraph-core</artifactId>
    <version>0.2.2</version>
</dependency>
compile "org.janusgraph:janusgraph-core:0.2.2"

Tested Compatibility:

  • Apache Cassandra 2.1.20, 2.2.10, 3.0.14, 3.11.0
  • Apache HBase 0.98.24-hadoop2, 1.2.6, 1.3.1
  • Google Bigtable 1.0.0
  • Oracle BerkeleyJE 7.3.7
  • Elasticsearch 1.7.6, 2.4.6, 5.6.5, 6.0.1
  • Apache Lucene 7.0.0
  • Apache Solr 5.5.4, 6.6.1, 7.0.0
  • Apache TinkerPop 3.2.9
  • Java 1.8

For more information on features and bug fixes in 0.2.2, see the GitHub milestone:

Version 0.2.1 (Release Date: July 9, 2018)

Legacy documentation: https://old-docs.janusgraph.org/0.2.1/index.html

<dependency>
    <groupId>org.janusgraph</groupId>
    <artifactId>janusgraph-core</artifactId>
    <version>0.2.1</version>
</dependency>
compile "org.janusgraph:janusgraph-core:0.2.1"

Tested Compatibility:

  • Apache Cassandra 2.1.20, 2.2.10, 3.0.14, 3.11.0
  • Apache HBase 0.98.24-hadoop2, 1.2.6, 1.3.1
  • Google Bigtable 1.0.0
  • Oracle BerkeleyJE 7.3.7
  • Elasticsearch 1.7.6, 2.4.6, 5.6.5, 6.0.1
  • Apache Lucene 7.0.0
  • Apache Solr 5.5.4, 6.6.1, 7.0.0
  • Apache TinkerPop 3.2.9
  • Java 1.8

For more information on features and bug fixes in 0.2.1, see the GitHub milestone:

Upgrade Instructions

HBase TTL

In JanusGraph 0.2.0, time-to-live (TTL) support was added for HBase storage backend. In order to utilize the TTL capability on HBase, the graph timestamps need to be MILLI. If the graph.timestamps property is not explicitly set to MILLI, the default is MICRO in JanusGraph 0.2.0, which does not work for HBase TTL. Since the graph.timestamps property is FIXED, a new graph needs to be created to make any change of the graph.timestamps property effective.

Version 0.2.0 (Release Date: October 11, 2017)

Legacy documentation: https://old-docs.janusgraph.org/0.2.0/index.html

<dependency>
    <groupId>org.janusgraph</groupId>
    <artifactId>janusgraph-core</artifactId>
    <version>0.2.0</version>
</dependency>
compile "org.janusgraph:janusgraph-core:0.2.0"

Tested Compatibility:

  • Apache Cassandra 2.1.18, 2.2.10, 3.0.14, 3.11.0
  • Apache HBase 0.98.24-hadoop2, 1.2.6, 1.3.1
  • Google Bigtable 1.0.0-pre3
  • Oracle BerkeleyJE 7.3.7
  • Elasticsearch 1.7.6, 2.4.6, 5.6.2, 6.0.0-rc1
  • Apache Lucene 7.0.0
  • Apache Solr 5.5.4, 6.6.1, 7.0.0
  • Apache TinkerPop 3.2.6
  • Java 1.8

For more information on features and bug fixes in 0.2.0, see the GitHub milestone:

Upgrade Instructions

Elasticsearch

JanusGraph 0.1.z is compatible with Elasticsearch 1.5.z. There were several configuration options available, including transport client, node client, and legacy configuration track. JanusGraph 0.2.0 is compatible with Elasticsearch versions from 1.y through 6.y, however it offers only a single configuration option using the REST client.

Transport client

The TRANSPORT_CLIENT interface has been replaced with REST_CLIENT. When migrating an existing graph to JanusGraph 0.2.0, the interface property must be set when connecting to the graph:

index.search.backend=elasticsearch
index.search.elasticsearch.interface=REST_CLIENT
index.search.hostname=127.0.0.1

After connecting to the graph, the property update can be made permanent by making the change with JanusGraphManagement:

mgmt = graph.openManagement()
mgmt.set("index.search.elasticsearch.interface", "REST_CLIENT")
mgmt.commit()

Node client

A node client with JanusGraph can be configured in a few ways. If the node client was configured as a client-only or non-data node, follow the steps from the transport client section to connect to the existing cluster using the REST_CLIENT instead. If the node client was a data node (local-mode), then convert it into a standalone Elasticsearch node, running in a separate JVM from your application process. This can be done by using the node’s configuration from the JanusGraph configuration to start a standalone Elasticsearch 1.5.z node. For example, we start with these JanusGraph 0.1.z properties:

index.search.backend=elasticsearch
index.search.elasticsearch.interface=NODE
index.search.conf-file=es-client.yml
index.search.elasticsearch.ext.node.name=alice

where the configuration file es-client.yml has properties:

node.data: true
path.data: /var/lib/elasticsearch/data
path.work: /var/lib/elasticsearch/work
path.logs: /var/log/elasticsearch

The properties found in the configuration file es-client.yml and the index.search.elasticsearch.ext.* properties can be inserted into $ES_HOME/config/elasticsearch.yml so that a standalone Elasticsearch 1.5.z node can be started with the same properties. Keep in mind that if any path locations have relative paths, those values may need to be updated appropriately. Once the standalone Elasticsearch node is started, follow the directions in the transport client section to complete the migration to the REST_CLIENT interface. Note that the index.search.conf-file and index.search.elasticsearch.ext.* properties are not used by the REST_CLIENT interface, so they can be removed from the configuration properties.

Legacy configuration

The legacy configuration track was not recommended in JanusGraph 0.1.z and is no longer supported in JanusGraph 0.2.0. Users should refer to the previous sections and migrate to the REST_CLIENT.

Version 0.1.1 (Release Date: May 11, 2017)

Documentation: https://old-docs.janusgraph.org/0.1.1/index.html

<dependency>
    <groupId>org.janusgraph</groupId>
    <artifactId>janusgraph-core</artifactId>
    <version>0.1.1</version>
</dependency>
compile "org.janusgraph:janusgraph-core:0.1.1"

Tested Compatibility:

  • Apache Cassandra 2.1.9
  • Apache HBase 0.98.8-hadoop2, 1.0.3, 1.1.8, 1.2.4
  • Google Bigtable 0.9.5.1
  • Oracle BerkeleyJE 7.3.7
  • Elasticsearch 1.5.1
  • Apache Lucene 4.10.4
  • Apache Solr 5.2.1
  • Apache TinkerPop 3.2.3
  • Java 1.8

For more information on features and bug fixes in 0.1.1, see the GitHub milestone:

Version 0.1.0 (Release Date: April 11, 2017)

Documentation: https://old-docs.janusgraph.org/0.1.0/index.html

<dependency>
    <groupId>org.janusgraph</groupId>
    <artifactId>janusgraph-core</artifactId>
    <version>0.1.0</version>
</dependency>
compile "org.janusgraph:janusgraph-core:0.1.0"

Tested Compatibility:

  • Apache Cassandra 2.1.9
  • Apache HBase 0.98.8-hadoop2, 1.0.3, 1.1.8, 1.2.4
  • Google Bigtable 0.9.5.1
  • Oracle BerkeleyJE 7.3.7
  • Elasticsearch 1.5.1
  • Apache Lucene 4.10.4
  • Apache Solr 5.2.1
  • Apache TinkerPop 3.2.3
  • Java 1.8

Features added since version Titan 1.0.0:

  • TinkerPop 3.2.3 compatibility

    • Includes update to Spark 1.6.1
  • Query optimizations: JanusGraphStep folds in HasId and HasContainers can be folded in even mid-traversal

  • Support Google Cloud Bigtable as a backend over the HBase interface

  • Compatibility with newer versions of backend and index stores

    • HBase 1.2

    • BerkeleyJE 7.3.7

  • Includes a number of bug fixes and optimizations

For more information on features and bug fixes in 0.1.0, see the GitHub milestone:

Upgrade Instructions

JanusGraph is based on the latest commit to the titan11 branch of Titan repo.

JanusGraph has made the following changes to Titan, so you will need to adjust your code and configuration accordingly:

  1. module names: titan-* are now janusgraph-*

  2. package names: com.thinkaurelius.titan are now org.janusgraph

  3. class names: Titan* are now JanusGraph* except in cases where this would duplicate a word, e.g., TitanGraph is simply JanusGraph rather than JanusGraphGraph

For more information on how to configure JanusGraph to read data which had previously been written by Titan refer to Migration from titan.