So in
the previous post, we got introduced to a dataset that I have been wanting to get into Neo4j for a long time: a Supply Chain Management dataset.
Read up about it over here, but the long and short of it is that we got ourselves into the situation where we have an up and running Neo4j database with 38 different multi-echelon supply chains. Result!
As a quick reminder, here's what the data model looked like after the import:

Or visually:
Data validation and profiling
The first thing to do when you have a new shiny dataset like that, is of course to get a bit of a feel for the data. In this case, it really helps to understand the nature of the different SupplyChains - as we know from the original Excel file that they are quite different between the 38 of them. So let's do some profiling:
match (n) return distinct labels(n), count(*)
Read more »Labels: apoc, database, google spreadsheet, graph, graphdb, multiechelon supply chain, neo4j, refactoring, scm, supply chain management