Showing posts with label graph databases. Show all posts
Showing posts with label graph databases. Show all posts

Tuesday, 14 January 2020

Graphistania 2.0 - Episode 3 - This Month in Neo4j

Happy new year everyone - although it actually seem like the holidays are already very far behind us! But great times were had, at least in my family, and so I feel super energised to make 2020 another great start to a decade of graphs :) ... Here's to that!

It also means that we are continuing to see all these awesome community stories pop up left right and center in the Neo4j "This week in Neo4j" developer newsletter. And so on our Graphistania podcast, we are going to continue talking about these on a monthly basis. So that's what we're doing - and I have again invited my friend and colleague Stefan Wendin to join me.

From the newsletter, we always select a few stories that we think will be more interesting and/or meaningful to discuss. This month, we found a number of them, and the interesting thing was that the graph-stories seemed to play at very different scales... The Personal, Corporate, and Society levels. Here are some of the ones we liked:

At the Personal scale
At the Corporate scale
At the Society scale, we saw some amazing posts:
So I think you agree that we had plenty of stuff to talk about. Let's get into that!

Wednesday, 3 July 2019

Finally: someone interviewed me on their podcast

This is worth a small celebration. Super conversation (in Dutch) with Jurjen Helmus, Walter van der Scheer and wingman Ron van Weverwijk on the Dataloog podcast about the use of graph databases and Neo4j. Listen to it over here over here - or find it on itunes/spotify.


Lots of fun to do - hope you enjoy it as much too!

Cheers

Rik

Thursday, 5 July 2018

Podcast Interview with Matt Casters, Neo4j & Kettle

A couple of years ago, I got to know another Belgian data aficionado that was doing quite a bit of work in the open source community, called Bart Maertens. For a while, we actually met at Antwerp Airport when we were both "commuting" to London City Airport for business - and we got a conversation going. Bart was organising a Pentaho Community Meeting in Antwerp, less than 500m from my home, and invited me to come along and talk a bit about my favourite subjects: beer and graphs :) ... 
So one thing lead to another, and Bart started to do some interesting work integrating his data integration tools with Neo4j. He wrote the code, and blogged about it in some detail

Fast forward to early 2018. Neo4j is more and more in the Enterprise market, with very large organisations seeing the value of graph databases and the platform around it. But most of these environments are NOT greenfield environments - they almost always require some kind of data integration work to make the tools work effectively. So it became very natural for us to start look for architects and experts that could help us... and that's effectively what brought my next Podcast guest to the Graph: Matt Casters has worked together with many other Neo4j people in a previous life, and is now the Chief Solutions Architect in our professional services organisation. 

Here's my chat with Matt:



Monday, 6 March 2017

Podcast Interview with Kristof Van Tomme, Pronovix

Last month I had one of those cool encounters of the graph kind at the Belgian Beerfest that we have been organising a couple of times in the the last few years at the occasion of Fosdem - the amazing open source conference that's taking place in Brussels every year. This year, I got talking to a fellow countryman that has been doing some amazing work on integrating the Drupal content management system with Neo4j - something that has a lot of potential in a lot of areas, I think. So - we just HAD TO have a chat :) ...


Here's the transcript of our conversation:
RVB: 00:03.346 Hello, everyone. My name is Rik, Rik Van Bruggen from Neo Technology. And here I am again the third time in two days, this is wonderful, I'm on a roll here, recording another podcast for our Neo4j Graphistania podcast. And today I have a fellow Belgian on the other side of this Skype call, and that's Kristof Van Tomme from Pronovix. Hi, Kristof. 
KVT: 00:27.466 Good morning Rik. How are you? 
RVB: 00:29.593 I'm really well, and I hope the Skype gods bear with us, because we've had some trouble in the past couple of minutes, but I'm sure it will fine. Hey, Kristof, we met each other at the FOSDEM conference, which was a great experience, and I loved the Beer Fest afterwards [laughter]. But yeah, you told me about some really great stuff that you guys are doing with graph databases. So, first of all, let's start from the beginning, who are you, what do you do and what's your relationship to the wonderful world of graphs? 
KVT: 01:07.202 So I'm a bit of a weird duck because I'm actually a bioengineer who ended up in IT through a biotech startup that did research in schizophrenia. It's a whole other life. But I got involved in the Drupal community a little over 10 years ago when we started making websites for biotech companies. 
RVB: 01:35.332 Okay. Drupal is like a content management system, right? 
KVT: 01:38.557 Yes, Drupal the open source content management system. The other really good Belgian product after beer and chocolates [laughter]. And I got really strongly involved in that community 10 years ago. I helped organise one of the big European conferences, and then we built a consultancy around that. Then, about five years ago, I got really excited about documentation, and reuse of documentation specifically, and how to deliver it and reuse bits and pieces so that you could build deliverables that can easily reuse between different channels. And that's how I got excited about graph databases, and Neo in specifically. 
RVB: 02:32.949 When you say documentation, you mean technical recommendation for software, right? 
KVT: 02:35.667 Yes. Yes, I do. The thing that everybody's like, "Ooh, documentation." 
RVB: 02:41.017 Ah, damn it. Yeah, exactly. 
KVT: 02:44.417 So that's how I got involved in-- because we had one of our colleagues, a long time ago, I think six years ago or something, started playing with graph databases, and actually, he built a first connector for Drupal for Neo. And he's like, "Kristof, I did this thing, and I'm really excited about graph databases, and I think it's cool. Can we do something with this?" And I was like, "I have no idea." So that was the first connector for Neo for Drupal, and then that kind of died because there was-- technically it was there, but then there were no further implementations, and I was not sold, and people didn't figure out how to use it. But then because of the documentation thing, I actually started seeing what you would use a graph database for and that's when I got really excited. 
RVB: 03:46.370 Super cool. Because documentation, I don't know if you notice, but this is where Neo4J started as well, as an open source project, 15 years ago, Viking hackers in a garage. They were all about content management at the time as well because they were working for a media company that was managing digital assets. So it's funny that there's this convergence or link between the two worlds, right? What is the use case all about? How does it work? 
KVT: 04:20.696 So I've been thinking-- I've got this DITA, which is another of those words. It's a standard that's fairly popular in the technical writing community for writing reusable documentation. It's like an XML standard. Some people scratch their heads when they hear about it, and other people are raving mad about it. So in the DITA community, I've been doing talks about consult management systems and open source and things like that. I think two years ago, I started thinking about personalisation and embedding information. What I dream about is this; instead of having a manual that the documentation system knows who you are and serves you the right information when you need it. I did a talk about that at the DITA conference here, I think it was in Europe, and I was thinking, "So how would you do that?" And then I started thinking yeah, actually, probably it wouldn't really work with a relational database because you need to start collecting a whole lot of information and start analysing for patterns. And that's how I started thinking about Neo and graph databases more in general. 
RVB: 05:48.382 So as a personalisation engine for documentation, right? So you wouldn't need to search for documentation as much, but you would have a recommended set of documentations that would be served to you semi-automatically. 
KVT: 06:04.195 Yeah. So it's the idea that, for example, you're in an application, you're in a web app, and you can't find that one damn button that you know is somewhere-- 
RVB: 06:16.996 We've all been there.
KVT: 06:18.043 Yeah, we've all been there. So you're clicking around, and you're going through settings, and I don't know, connections, so you keep going circles and circles and circles because you can't find the damn button. And at that point, the system would say, "This looks a lot like what people do when they're looking for this thing," and then you would get a little pop-up saying, "Are you maybe looking for this?" And similarly, if you're using a certain feature and you're doing something really weird and other people have done that, and then they went through the documentation and found some other feature, then you could shortcut that and skip a few jumps in that graph and immediately serve them the information that they're looking for. So it's kind of like analysing patterns of behaviour that people have inside of a web application and then serving them-- that's patterns of behaviour that they normally do just before going to documentation sites and then serving them that documentation that people normally will find when they go to documentation site after they've done a certain thing, and then serving that information to them. So that's one of the really cool things that I would like to do. 
RVB: 07:32.834 Yeah, I understand. So why is that such a good use case for a graph database? Is that because of the pattern recognition, or what's the secret sauce? 
KVT: 07:44.973 So it's the pattern recognition. So I think CMSs are really good at storing data in a-- storing similarly structured information because most of CMSs use SQL databases and they're pretty good at that, just building up a content model and then reusing that over and over again. But being able to recognise behaviour-- well, that's not something that we are normally doing in the CMS space. We have some very basic things, like there's some recommendation based on the content and shared keywords and things like that, but behaviour analysis is not one of the things that you normally find in the CMS. So for that, we need different technology because in a SQL database you would have to do so many joints to even figure out what's going on, yeah, that I don't think that it would make sense to do it that way. And ideally, it would be a system that you don't have to program everything but that it can start looking for patterns on its own eventually. And that you build this graph of interactions and content and kind of like a graph that combines those two to do things with that. So yeah. 
RVB: 09:04.602 So where are you guys with this? How far along that path are you? I know you've done some prototyping already, right? 
KVT: 09:11.567 Yeah. So we are very, very early. So our main business right now is developer portals. So two years ago we started working-- well, a year and a half ago we started working with APG, that's now part of Google, and they have a developer portal that we are customising for their customers. And we built this whole business around documentation, specifically about APIs, so that's where our core focus is right now. And so the AI and personalised documentation is something that we're doing research on. So the thing we've done currently is we've built a connector for Drupal for Neo - I did a talk about that at FOSDEM - and that was-- 
RVB: 10:02.991 I went to that one, yeah. 
KVT: 10:04.291 Yeah. So that talk was not just about this use case. It was about what could you do if you combine a CMS and a graph database and looking at it from an added-value perspective, rather than a replacement perspective. Because I know that in the DO community people are like, "Just get rid of the stupid SQL databases [laughter]." They're worthless and graph databases can do everything so much better. I think--
RVB: 10:37.056 That's a pipe dream in my opinion. 
KVT: 10:38.368 Probably. You could build a CMS graph database, and I think that could work. But I think that there's so much existing technology already where it's a large amount of extensions and huge communities that it would make more sense to create an add-on instead of a replacement because if you replace it then you have to rewrite everything. 
RVB: 11:05.120 I couldn't agree more. 
KVT: 11:06.124 Yeah. So that's why I think that's their sweet spot for Neo in the CMS community but I think there's two stress facts to this. One is the sweet spot for neo in the CMS community, and that could be recommendation and pattern finding. But then there's also the inverse that you could think about and that's what if you were to put an open source CMS like Drupal in front of a graph database and we use it as an interface to manipulate the graph and to add, maybe, some structured objects into your graph? And then use the CMS to build reports about those objects and the graph to find out which ones you're going to put into your reports. So that was my talk about. 
RVB: 11:57.922 Well, you've already touched on my last question, which is what does the future hold [laughter]? What could we do in the future? And I know that we'll be doing some meet-ups together and I'm really looking forward to those, but where does this go, Kristof? What's in your crystal ball? 
KVT: 12:21.174 So I love thinking about a future. I really love Kevin Kelly's book, The Inevitable. And in that book, he talked about-- I think this is the basic pattern that got me thinking about this, also. He talks about flowing and it's a very, very interesting concept that we're moving from an Internet where we used to have documents to an Internet where we have pages today, where we'll have flows of information tomorrow. And this idea of going from having an object that's structures and it has a context-- has a manual context, or a book context, or a document's context where you put all the information in context of the rest of the book into a very rigid structure. That's how we used to do things. That's how books and manuals were built, even when printing press-- even before the printing press was invited. And what the Internet has been doing, and what search engines have been doing, is that we've been moving towards pages where you can just dive into any object-- sorry, any document, any book, and just find out one page where a certain concept is explained. So you can just jump in. You don't have to read the whole book to be able to understand something. And that's where we are today. But I think that's the next step in this process, and it's also what Kevin Kelly talks about is flows, where you have a flow of information that's much more personalised, and we're just constantly dipping in and out of these information flows around us that are serving us the documentation that we need at a certain time to be able to do what we need to do and that are aware of our contexts so that we don't have to adjust to the context of the documentation, but the documentation adjusts to our own personal context, and I think-- yeah? 
RVB: 14:31.872 So what I'm hearing is you see this graph database integration and everything that you guys are building as a means to that end, to get there somewhere, somehow, to get closer to it. 

KVT: 14:44.560 Yeah. So we have a first customer where I've been talking about this concept, and-- they're an SaaS company. So what I imagine is that we could track users, the administrators as interacting with the software, and then basically serve them the contents this way where you look at their whole experience inside of your tool, and then you serve them the information they need to be able to interact better and get more value out of your system. So it's kind of like the idea-- the way that I describe it going from the context of the manual to the context of the one, like the one person, one single user and how they are interacting with the system. This is very, very-- there's a lot of work to get here [laughter]. But I think that we can take baby steps, start with first implementation. Start with building a graph of the behaviour and how people interact with documentation and with the tools that are documented by the documentation and then use that to start recommending content. And yeah, I'm really excited about it. We started a mailing list about it at one of the meet-ups where I was presenting. We actually had one of the people that worked on the Clippy years and years ago at Microsoft who was also really excited about the idea. Because I think this is actually what Clippy wanted to do, or wanted to be, but it was not possible. And I think that graph databases could be the piece of technology that enables the dream of Clippy [laughter]. 
RVB: 16:40.452 Well, I think on that bombshell [laughter], I think that's a great time to kind of wrap up this podcast. Thank you so much for coming online, Kristof, and we'll be publishing some more details around your work and also the talks that you've been doing with the transcription of the podcast so people can read up about it. And I look forward to seeing you at one of our meet-ups, right? Because we'll be doing some community work together in the next couple of months as well. So really looking forward to that. 
KVT: 17:12.653 Likewise. 
RVB: 17:13.552 Thank you so much. Have a nice day, Kristof. 
KVT: 17:16.253 Yeah, you too. 
RVB: 17:16.990 Bye. 
KVT: 17:17.383 Bye.
Subscribing to the podcast is easy: just add the rss feed or add us in iTunes! Hope you'll enjoy it!

All the best

Rik

Friday, 6 March 2015

Starting a Podcast about Graph Databases and Neo4j

Been wanting to do this for months, and last week at Qcon I decided to take the bull by the horns, and just try it out. I want to create a podcast about Graph Databases and the Neo4j ecosystem, featuring a weekly short post about what is going on in the wonderful world of Graph Databases. We'll try to publish a session every week - or perhaps more often if people like it a lot.

You can subscribe to the podcast on this feed url, this iTunes link or on the Soundcloud playlist. I will also embed it here.

The first session is an interview with my friend and colleague Michael Hunger. General lovely guy, but also one of the driving forces in the Neo4j community. Someone to listen to:
You can find the transcript of the podcast below:

RVB: Hi everyone. My name is Rik Van Bruggen, I work for Neo Technology, and this is a test of a series of podcasts that we want to be doing in the next couple of months, and I have invited someone that I would love to ask a couple of questions to about graph databases in general and where this whole thing is going. So with me today is Michael Hunger, Chief Evangelist at Neo Technology. Michael, do you mind introducing yourself?

MH: Yeah. I am Michael Hunger, I have been with Neo Technology now for almost 5 years, and I love to help people be succesful with Neo4j and make them happy - that’s what I like to do.

RVB: That’s fantastic. So how have you been with Neo, and how did you get to Neo?

MH: Actually I met Emil Eifrem, the CEO of Neo Technology, in 2008 on a Geek Cruise, where he talked about this Graph Database thing and he got me pretty much excited, I tried it out, we stayed in contact and in 2010 I joined the company - and have been active and happy ever since.

RVB: So awesome. Really I have only two questions for you, Michael. The first thing I wanted to ask you was - what do you love about Graph Databases? And then I would also like to talk to you about what you see in the future? So let’s start with the beginning: what do you love about graph databases, and why do you think they are the best thing since sliced bread?

MH: So the best thing about graph databases is this intuitive and flexible data model. So you can just grab your domain, take a whiteboard, your business expert, and just sit down and draw what you want to talk about, what you want to ask - and then take this data and these relationships and put them in your database without having to destroy it or change the structure … you kind of stay in the same model all the time. Graph databases and Neo4j specifically are really great at keeping this rich connected dataset and making managing and accessing it super fast and super efficient.

RVB: so do you think this would be something that anyone could use, that the average developer could be using? Or are we going to need a PhD. before we can start using it?

MH: fortunately you don’t need to have a PhD. It’s actually as easy as using a relational database. Just get Neo4j dowloaded, start it on your server machine and get started directly because it comes with this really intuitive trial image out of the box that is super easy to learn and super easy to use. You can get up and running within a few minutes, and be productive after this already.

RVB: so do you have any favourite use cases that you think are super applicable for it?

MH: there are a lot of cool use cases out there, but my personal favourite is software analytics, actually - because software is a graph and there is a lot of connected information in a software project - and I love to use Graph Databases for that.

RVB: so that means like analysing dependencies and all of those types of things?

MH: exactly. There are so many connections - there are static class and method structures, information about heaps and database connections, and lots of other things.

RVB: Super. So one more question for you. Where do you see this going? Where do you see it and where do you want it to go in 2-3-5 years from now for graph databases like Neo4j.

MH: So first of all I would love for people being addicted to graphs because I think it is a really universal data model and there are so many interesting domains that you could use it for, So I want to see it more broadly adopted, in general. And then of course I want more interesting use cases. From the feedback in the user community I want Neo4j to grow in many dimensions. From versatility to handling new data types, like spatial and time/versioning better. That would enable new applications. But then also integrated with other technologies: so to be able to take all of the infrastructure that you have, integrate it with Neo4j and have a really nice roundtrip experience.

RVB: So thanks so much Michael. We are going to keep it at that - I think it was a super interesting conversation. If you want to know more about Neo4j, just go to neo4j.com, reach out to us on Twitter @neo4j or just come and see us at one of our events near you. Thanks for listening - talk to you soon.

Hope you like it. Feedback welcome.

Cheers

Rik

Thursday, 22 January 2015

Innovation Pitch

Some companies are interesting. I mean, I myself have been working in startup environments for (djeez! has it been so long!) decades now, but some large organisations are equally interesting - especially in today's "information age". Every so often I get to meet fascinating people that are working in an industry that is literally being thrown upside down because of the modern technology swell of connectedness, mobile information, demanding customers, and innovative applications. After years, centuries sometimes, of successful business ventures in the "good old days", they find themselves in a place where they are sitting on wonderful assets, with real value, but also facing a growing need to re-assess how it all fits in this new age of digitalism. They need to innovate.

Innovation is a hard nut to crack. I am not an expert, but when I read the "Innovator's dilemma" a few year's ago, it became blatantly clear to me that innovation does not come natural to a large organisation. It simply doesn't. There's all kinds of internal and external forces that actually make it tremendously hard for large organisations to truly innovate.

That's probably why I personally find Startup organisations more my cup of tea, but it's also why I am truly impressed and greatly sympathetic when I see large organisations make a truly consolidated effort to innovate.


Yesterday, I was part of such an effort. Wolters Kluwer, global publishing powerhouse with a long standing history, headquartered in the Netherlands, organised an Innovation Pitch event for their executive team. Almost all of their board members and execs were there, and I had 10 minutes to "pitch Neo4j". Interesting.

I thought about this a bit - and I decided to go for the "high road". The pitch was not meant to sell product, not meant to position Neo4j even - but really was geared to getting these top-level international execs to think differently - to open up their minds to the wonderful world of graphs. I used the example of "How Wolves Change Rivers" to help illustrate that - as seen over here, or in the GraphGist over here.



I recorded the pitch at home - see below. The actual presentation included some Q&A and took a bit longer in total - but it was pretty much like this:



Slides are over here:



Probably a ton of other things that I could have said - but my main goal was to be remembered and get a conversation going with Wolters Kluwer. I would love to get your feedback, if any.

Cheers

Rik




Friday, 19 September 2014

Graphs for HR Analytics

Yesterday, I had the pleasure of doing a talk at the Brussels Data Science meetup. Some really cool people there, with interesting things to say. My talk was about how graph databases like Neo4j can contribute to HR Analytics. Here are the slides of the talk:

I truly had a lot of fun delivering the talk, but probably even more preparing for it.

My basic points that I wanted to get across where these:
  • the HR function could really benefit from a more real world understanding of how information flows in its organization. Information flows through the *real* social network of people in your organization - independent of your "official" hierarchical / matrix-shaped org chart. Therefore it follows logically that it would really benefit the HR function to understand and analyse this information flow, through social network analysis.
  • In recruitment, there is a lot to be said to integrate social network information into your recruitment process. This is logical: the social network will tell us something about the social, friendly ties between people - and that will tell us something about how likely they are to form good, performing teams. Several online recruitment platforms are starting to use this - eg. Glassdoor uses Neo4j to store more than 70% of the Facebook sociogram - to really differentiate themselves. They want to suggest and recommend the jobs that people really want.
  • In competence management, large organizations can gain a lot by accurately understanding the different competencies that people have / want to have. When putting together multi-disciplinary, often times global teams, this can be a huge time-saver for the project offices chartered to do this. 
For all of these 3 points, a graph database like Neo4j can really help. So I put together a sample dataset that should explain this. Broadly speaking, these queries are in three categories:
  1. "Deep queries": these are the types of queries that perform complex pattern matches on the graph. As an example, that would something like: "Find me a friend-of-a-friend of Mike that has the same competencies as Mike, has worked or is working at the same company as Mike, but is currently not working together with Mike." In Neo4j cypher, that would something like this
 match (p1:Person {first_name:"Mike"})-[:HAS_COMPETENCY]->(c:Competency)<-[:HAS_COMPETENCY]-(p2:Person),  
 (p1)-[:WORKED_FOR|:WORKS_FOR]->(co:Company)<-[:WORKED_FOR]-(p2)  
 where not((p1)-[:WORKS_FOR]->(co)<-[:WORKS_FOR]-(p2))  
 with p1,p2,c,co  
 match (p1)-[:FRIEND_OF*2..2]-(p2)  
 return p1.first_name+' '+p1.last_name as Person1, p2.first_name+' '+p2.last_name as Person2, collect(distinct c.name), collect(distinct co.name) as Company;  

  1. "Pathfinding queries": this allows you to explore the paths from a certain person to other people - and see how they are connected to eachother. For example, if I wanted to find paths between two people, I could do
 match p=AllShortestPaths((n:Person {first_name:"Mike"})-[*]-(m:Person {first_name:"Brandi"}))  
 return p;  

and get this:
Which is a truly interesting and meaningful representation in many cases.
  1. Graph Analysis queries: these are queries that look at some really interesting graph metrics that could help us better understand our HR network. There are some really interesting measures out there, like for example degree centrality, betweenness centrality, pagerank, and triadic closures. Below are some of the queries that implement these (note that I have done some of these also for the Dolphin Social Network). Please be aware that these queries are often times "graph global" queries that can consume quite a bit of time and resources. I would not do this on truly large datasets - but in the HR domain the datasets are often quite limited anyway, and we can consider them as valid examples.
 //Degree centrality  
 match (n:Person)-[r:FRIEND_OF]-(m:Person)  
 return n.first_name, n.last_name, count(r) as DegreeScore  
 order by DegreeScore desc  
 limit 10;  
   
 //Betweenness centrality  
 MATCH p=allShortestPaths((source:Person)-[:FRIEND_OF*]-(target:Person))  
 WHERE id(source) < id(target) and length(p) > 1  
 UNWIND nodes(p)[1..-1] as n  
 RETURN n.first_name, n.last_name, count(*) as betweenness  
 ORDER BY betweenness DESC  
   
 //Missing triadic closures  
 MATCH path1=(p1:Person)-[:FRIEND_OF*2..2]-(p2:Person)  
 where not((p1)-[:FRIEND_OF]-(p2))  
 return path1  
 limit 50;  
   
 //Calculate the pagerank  
 UNWIND range(1,10) AS round  
 MATCH (n:Person)  
 WHERE rand() < 0.1 // 10% probability  
 MATCH (n:Person)-[:FRIEND_OF*..10]->(m:Person)  
 SET m.rank = coalesce(m.rank,0) + 1;  

I am sure you could come up with plenty of other examples. Just to make the point clear, I also made a short movie about it:

The queries for this entire demonstration are on Github. Hope you like it, and that everyone understands that Graph Databases can truly add value in an HR Analytics contect.

Feedback, as always, much appreciated.

Rik

Wednesday, 4 June 2014

Graph Local queries - revisited

Recently I had some feedback from Chuck Daniels on my blogpost on Graph Local queries. In the comment, Chuck challenged me in saying that really I should be comparing apples to apples, and that I should make sure that I was applying the same kind of locality.

Basically his question was as follows: once we go from the small dataset to the large one, shouldn't the query that we test performance for also include a richer, more complex pattern. Essentially he wanted me to compare these two queries:

match 
(eqt:EQUIPMENT_TYPE)<-[:IS_OF_TYPE]-(eq:EQUIPMENT)-[:LOCATED_AT]->(ol:OBSERVATION_LOCATION)<-[:OBSERVED_AT_LOCATION]-(o:OBSERVATION {id:1001})
return eqt.name, ol.name, o.id; 

and

match 
(eqt:EQUIPMENT_TYPE)<-[:IS_OF_TYPE]-(eq:EQUIPMENT)-[:LOCATED_AT]->(ol:OBSERVATION_LOCATION)<-[:OBSERVED_AT_LOCATION]-(o:OBSERVATION {id:1001}),
(eq)-[:USED_FOR]->(ot:OBSERVATION_TYPE)<-[:IS_OF_TYPE]-(o)
return eqt.name, ol.name, o.id; 

The second query obviously being a bit more complex, as it adds a few more "hops" to the traversal. 

So lets test this out

I have created a little gist for you to try this yourself. Look at this one for loading the data, and testing the results yourself. We start with an empty database, of course, and load the small dataset first.

Then we run the sample queries (both of them: the easier one AND the more complex one) and create the index on the OBSERVATIONS.

Next up is adding the larger dataset, and running the queries again:

And surprise surprise, the principle of graph locality and index-free adjacency still survived - the queries are still lighting fast. 

Hope this clarifies the point that Chuck raised - and reinforces the fact that graph local queries are GREAT for many different use cases!

All the best

Rik

Tuesday, 14 January 2014

Cool graph events in the next couple of weeks!

Waw. I just realised that there are a TON of very cool graph events coming up that I am going to be fondly participating in.

Hoping to see many of you there!



Friday, 20 December 2013

Graphs for Everyone!

Here's some of my thoughts on how to best promote innovative, wonderful, and new technology like neo4j 2.0 - and get it to be used ubiquitously. These are just my own thoughts - but I was hoping they would be useful to our thousands of devs and architects out there that are struggling to sell graph database technology to their peers, their bosses, their business.

So turn up your sound (the Prezi has voice-over - a great new feature!)



Let me know if you have any feedback - would love to hear your thoughts.

Hope this is useful.

Cheers

Rik

Thursday, 12 December 2013

Saint Nicolas brought me a new Batch Importer!!!

After my previous blogpost about import strategies, the inimitable Michael Hunger decided to take my pros/cons to heart and created a new version of the batch importer - which is now even updated to the very last GA version of neo4j 2.0. Previously you actually needed to use Maven to build the importer - which I did not have/know, and therefore never used it. But now, it's supposed to be as easy as download zip-file, unzip, run - so I of course HAD to test it out. Here's what happened.

Yet another dataset

First: I wanted to create a "large-ish" dataset (Michael actually calls it "tiny") with 1 millions nodes and 1 million relationships. So what do you do? MS Excel to the rescue. I created an Excel file with two worksheets, one for nodes and one relationships. The "nodes sheet" has nodes arranged in the following model of persons and animals that are each-other's friends (thanks Alistair again for the Arrows):

Creating the nodes sheet was easy, creating the relationship sheets I actually used a randomization function to create random relationships:

=RANDBETWEEN(nodes!A$2;nodes!A$1048576)

The Excel file that I made is over here. By doing that I actually get a fairly random graph structure - if I would manage to import it into neo4j. In order to do so with the batch importer, I simply had to export the file to two .csv files: one for nodes, one for relationships. And then there was one more step: I had to replace the semi-colons with tabs in order for the batch importer to like the files (I probably could have done it without this step, by editing the batch.properties file as in these instructions). Easy enough in any text editor - done in 2 seconds.

Drumroll: would it work?

So I downloaded the zip file, unzipped, and went

./import.sh graph.db nodes.csv rels.csv

Then I wait 20 seconds (apparently this is going to get a lot faster very soon - but I was already impressed!) and: TADAAH!

Job done!! 

All I had to do then was to copy the graph.db directory (download the zipped version from over here) to my shiny new 2.0 GA instance directory, fire up the server, and all was fun and games. Look at the queries in the neo4j browser, and you see a beautiful random animal-person social network. So cool!


What did I learn?

Thanks to Michael's work, the import process for large-ish datasets is now really easy. If I can do it you can. But. There was a but.

Turns out that the default neo4j install on my machine (with an outdated version of Java7, I must admit) actually ran painfully slow after a few queries. But as soon as I changed one little setting (the size of the initial/maximum Java Heap size = 4096, on my 8GB RAM machine) it was absolutely smoking hot fast.  Look for the neo4j-wrapper.conf file in your conf directory of the neo4j install.
I guess I just never played around with larger datasets in the past - this definitely made a HUGE difference on my machine.

UPDATE: I just updated my Java Virtual Machine to the latest version, and this problem has now gone away. You don't need the above step if you are on the latest version - just leave it with the default settings and it will work like a charm!

So: THANK YOU SAINT NICOLAS for bringing me these shiny new toys - I will try to continue to be a good boy!

Hope this was useful.

Rik

Friday, 22 November 2013

Meet this "Tubular" graph!

Many of us know London. Those of us that have visited London will know "the Tube", "the Underground" - simply the fastest and most efficient way to get around (although I must admit that Hailo has been quite a contender lately...). Beautiful city, lovely place to work, and since I started working for Neo, it feels a bit like my home away from home.

The Tube: A Great Graph

As you can easily imagine, or just plainly see from looking at any of the maps of the tube, the Underground really is a very sophisticated system, and can only be described as a very sophisticated graph. We always refer to it - in our Neo4j presentations - as the perfect example of how one-page-graphs can easily represent and provide *insight* into complex system ... without having to have a PhD in maths. Literally: almost everyone can use the tube - almost everyone can use a graph.

Finding a nice "tubular" dataset

Since we talk about this example all the time, and since I am indeed an avid, non-native tube-user, I thought it would be interesting to look at how I could fit the Tube system into a neo4j database. It took me a while, but of course the data is out there: this page links to this spreadsheet that has a very nice starting point. It contains the Line, the Direction, the Stations, the Distance between stations, and then 3 different time measurements between the stations.


Importing this into a neo4j database is really, really easy.

Creating a neo4j Tube database

First things first: from the above spreadsheet, we would probably be best off to transform it into a .csv file. Easy peasy in Excel: the result is over here. Once we have that, we can use the ever so awesome neo4j-shell-tools (the 2.0 version is over here, in case you can't find it!) to import the data into a nice little graph model:

Kudos to Alistair Jones for making Arrows - it's actually very useable these days :)) ...

In other words: Stations have to be unique, are connected by one or more "Lines" in two directions, and the "Lines" have a "Direction" property (east, west, north, south...), a "Time" between stations property (which can be different in opposite directions!), and a "Distance" between stations property.

The import script for the .csv file is quite simple, as it completely leverages the new neo4j 2.0RC1 way of working:
  • it uses a schema constraint to ensure that the stations are unique
  • it uses the new Match-syntax (with property-matching in the pattern instead of in a where clause)

All in all it is very simple and effective. The resulting graph.db directory is over here.

Exploring the tube in the neo4j browser

Ever since it's introduction at GraphConnect San Francisco, the neo4j browser has become my favourite place to play around with neo4j and cypher. One of it's coolest features is the ability to apply stylesheets to your graph visualisations. So I wanted to apply this to my new tube-graph, and use the "official" tube-line colours in the browser. 


LINETRUE HEXADECIMALWEB SAFE HEXADECIMAL
Bakerloo
#B36305
#996633
Central
#E32017
#CC3333
Circle
#FFD300
#FFCC00
District
#00782A
#006633
Hammersmith and City
#F3A9BB
#CC9999
Jubilee
#A0A5A9
#868F98
Metropolitan
#9B0056
#660066
Northern
#000000
#000000
Piccadilly
#003688
#000099
Victoria
#0098D4
#0099CC
Waterloo and City
#95CDBA
#66CCCC
So then all I had to do was to download the .grass file from the browser, and start editing the "relationship" sections. In the example below, the .Circle and .Central are the names of the relationship types "Circle" and "Central". Logical.

You can download the full .grass file that I created from over here.

Nice: if I start exploring the surrounding tube network for "London Bridge" station, I quickly get a feel for the network:


But of course the real fun begins with the queries.

Exploring the Tube with Cypher

Obviously I don't have the technical skills - at all - to develop anything like a route planner for the London Underground. But: using the dataset that we just created, it's quite easy to see how it would be very doable to create something like that. Let's look at some of the queries that I created:

Show the different underground lines:


Show the most densely connected underground station

With "densely connected" meaning the most different underground lines passing through it.

And then you can drill into this really easily and explore some more:

And finally: pathfinding

Of course we can do some rudimentary pathfinding in Cypher. But it's rudimentary - and just included for fun. Let's say that I would want to go from Tower Hill to Southwark (one of the most tedious tube connections that I would take sometimes to get to our London office).


Anyone a bit familiar with London knows that this is "b*ll*cks", and noone would ever do that. The right thing (I think) to do is to take the district/circle lines from Tower Hill to Blackfriars - and then just walk across the bridge to the office. Easy.


I have included some other pathfinding queries in the gist - but I am pretty sure that they would need work :) ...

That's about it for now. I think I have demonstrated how easy it is make the Great Tube Graph even greater by putting it into a graph database like neo4j - and how you could easily use something like the neo4j browser to find your way around one of the world's most complicated networks. 

Hope you enjoy!

Rik

Tuesday, 12 November 2013

Presenting: Neo4j!

Last week, I had the great pleasure of spending some time working on a BI-centric presentation of neo4j. Tereza subtly re-introduced me to Prezi, a presentation format that I had already used a couple of years ago when I was still working for Imprivata. I found that a great tool had only gotten better, and that it was actually quite a lot of fun to store and present neo4j this way. After all - when you think of presentations in general, and prezi more specifically - it's actually quite easy to represent any presentation/process as a graph.

So I had this weird idea: what if I would create a prezi-presentation about neo4j, and a simple but complete little neo4j database that would essentially contain the same information as the prezi - so that people could explore neo4j - in neo4j. Neo4j presenting itself in a neo4j database. I know - it's a bit of a joke. But here goes anyway.

Creating the prezi

I spent a bit of time acquainting myself with the new prezi interface, and created this prezi:



I actually quite like the result. It's a nice overview presentation - and I must say I am particularly proud of the "elevator pitches" that I created.

The Elevator Pitches

An elevator pitch is supposed to explain a concept to another person, while you're in the elevator together. Depending on the size of your typical high-rises, that means that you would have between 10-20 seconds. Short. So how do you explain neo4j in that time? Well, I think the trick is to know - or at least make some assumptions about - who you are talking to, and tune the story to the audience.

Last week, in just a few days time, I had multiple "graph database virgins" (people that had no idea what it was, and that sometimes also did not have a lot of technical baggage) ask me "what neo4j was". And even today, I still found it challenging. Here's what I came up with to explain neo4j to "mom & pop":

There's a couple of other "pitches" in the prezi, covering other audiences like developers, architects, project managers, CIOs and business managers. I am sure they are not perfect -  I would love some feedback on these if you feel like it - but hey, Elevator Pitches are not meant to be perfect. They are meant to cause interest, so that the conversation can continue.

But then I wanted to have some fun and put the prezi into neo4j.

Creating the neo4j database presenting... neo4j

I know this is silly - but when you think about it really isn't that stupid. Prezi assumes that there is a "path" to present the prezi. Meaning: I, the presenter will determine how I take you through the information presented. And of course, that process is a bit arbitrary for the attentive listener: every listener has his/her own personal background, knows more or less about technology, and of course, about graph databases in general and neo4j specifically. So it actually could make sense for someone to want to present the information in the prezi in a "freer" format that could be explored randomly by the audience. And that format would be: a neo4j database :) ...

I ended up creating the database with the spreadsheet method that I used before: take a look at the sheet over here. Run the cypher queries to create the nodes in a neo4j 2.0 instance, create the index, and then connect them up with some cypher queries to create the relationships. Or: just download the graph.db directory from over here, and copy it onto your neo4j server. Fire up the awesome neo4j browser, and you will soon be looking at something like this:
It's essentially the same thing as the prezi - just nicer :) ... Neo4j explaining itself to you - with neo4j! How cool is that?



I have also created a little graphgist that you can take a look at. Download the gist from over here.

That's about it. Hope you like it - as always feel free to comment or ping me if you want.

Cheers

Rik