Slater – Low-memory graphdb designed for read-heavy graphs
Posted by rickkjp 3 days ago
Comments
Comment by rickkjp 3 days ago
Slater starts with the premise of a fixed memory budget applied to an LRU cache of what's on disk, then uses ISAM blocks and DiskANN/Vamana/PQ to allow paging the contents of the graphs and vectors you need into that cache. It's designed around read-heavy-write-light cases, and can be backed either by local disk or by S3/GCS buckets with an optional sized local disk L2 cache as well.
Speaks standard Bolt, and is GDPR friendly (encryption at-rest and in-transit). Multi-user-multi-graph with ACLs, Rust-with-forbid-unsafe and NFS-friendly too (no mmaps). It's Apache licensed.
Please give it a try if you get a chance. Would love any suggestions or feedback.
Full disclosure: yes, it was authored by Claude Code, although I provided the storage model and design it used, along with code samples and influences from other open source projects like FalkorDB and Memgraph for Bolt wire-compatibility.
Thanks
Comment by j-pb 3 days ago
I still always want to know the numbers per edge. Because there'a a reaon why these graph DBs try to hold the stuff in memory, everything else is slow as heck. We have ~30bytes per edge in our succinct indexes, which makes the 1.5gb wiki dataset clock in at ~50gb.
If I understand correctly you shard the graph and then content adress each chunk:
- How do you decide on the sharding? I'd expect finding good connected subsets with nice memory locality to be extremely difficult computationally?
- How do you canonicalise your graph. Which has also been an extremely difficult problem in RDF land for example. Although that one is at least efficiently solvable.
Comment by rickkjp 2 days ago
The wikidata image I was using (sourced from a huggingface dataset Arun Sharma kindly uploaded for LadybugDB) is about 10GB zipped I think, and when imported into Slater takes up about 20GB of disk space. It's about 133GB of raw Cypher.
So in bytes/edge, that's probably about 14.
In terms of how to shard (which I'm taking to mean allocation of nodes to ISAM blocks) during an initial build or during a consolidation step, it will try to pack neighbours into the same blocks along with metadata about neighbouring segments, with some additional heuristics to avoid superhubs causing over-filling issues. For on the fly writes it just uses a normal WAL that it builds tables out of until theres a segment merge, at which point it then uses the same logic as for the consolidation.
In terms of canonicalising, it asks for a pk property during import, and then uses a (label, key property, value) tuple for uniqueness. So in a sense it doesn't really do that itself, it relies on that being handled outside the import.
Hope that helps.
PS as a follow-up: one really important memory-saving trick I used was Elias-Fano encoding on the adjacency matrices. Each node keeps an array of nodes it's joined to, and Elias-Fano does an amazing job of compression on those, both on disk and in RAM.
Comment by j-pb 2 days ago
So you're also using succinct datastructures with Elias-Fano, you might want to check out this paper https://aidanhogan.com/docs/ring-graph-wco.pdf on WCO joins over succinct indexes.
Comment by FrustratedMonky 3 days ago
Since it is a much in demand feature. Why do you think they have not done it themselves already, versus what you have done?
Curious if this is breakthrough? Or there are some negatives that prevent Neo4J from doing it also?
Comment by rickkjp 3 days ago
Comment by FrustratedMonky 2 days ago
A lot of DB design is around this subject. And so it seems like if you found a way to allow a graph DB data to expand beyond in-memory by allowing it to reside on disk is a really big deal. But, since it would be a big deal, I'd think the big guys with a lot of resources would offer it as an option, or feature. It would be worth investing in.
Comment by packetlost 3 days ago
LLMs have no concept of focus when it comes to docs, so they spew out way more information than is necessary or helpful to a human reading the document.
Comment by rickkjp 3 days ago
Comment by hankbond 3 days ago
Comment by UltraSane 3 days ago
Comment by valentynkit 2 days ago
Comment by pbronez 3 days ago
Comment by rickkjp 3 days ago
Comment by maxdemarzi 3 days ago