Showing posts with label graphaware. Show all posts
Showing posts with label graphaware. Show all posts

Tuesday, 15 November 2022

A 2nd, better way to WorldCupGraph

Hours after publishing my previous blogpost about the WorldCup Graph, I actually found a better, and more up to date dataset that contained all the data of the actual squads that are going to play in the actual World Cup in Qatar. I found it on this wikipedia page, which lists all the tables with the actual squads, some player details, coaches etc. as they were announced on 10th/11th of November.

So: I figured it would be nice to revisit the WorldcupGraph, and show a simpler and faster way to achieve the results of the previous exercise. So: I have actually put this data in this spreadsheet, and then downloaded a .csv version:

These two files are super nice and simple, and therefore we can actually use the Neo4j Data Importer toolset to import these really easily.

Tuesday, 26 March 2019

Podcast Interview with János Szendi-Varga, GraphAware

Finally - managed to get out another episode of our podcast. Early February, our friends of GraphAware published a great blogpost about their view of the "Graph Technology Landscape". The main author János Szendi-Varga, was kind enough to spend some time with me on the phone talking about our industry and where it's going. Hope you enjoy it as much as I did!


Here's the transcript of our conversation:
RVB: 00:00:03.184 Hello everyone. My name is Rik, Rik Van Bruggen from Neo4j, and here we are again recording another episode of the Graphistania, Neo4j graph database podcast. And today I have a special guest on the podcast all the way from the UAE in the Middle East and that's a very, very interesting person that did some amazing work recently with the graph database landscape. It's János Szendi-Varga. Hi János.

Thursday, 7 February 2019

The Graph Technology Landscape Graph

Our friends at Graphaware, and specifically Janos Szendi-Varga created a fantastic overview of the Graph Technology Landscape in 2019. It's actually pretty cool:
There's lots of interesting data in there, which Janos put into a CSV file over here. I thought that was really cool, and took the data for a spin in a Google Sheet, and then modified it a little bit from there.

Next thing for me was to create a

Which I have now made available as an import script and a graphgist over here. You can explore the data in the deployed graphgist yourself, and figure out some of the hidden and not-so-hidden connections in the Graph Technology Landscape. 

Hope you will have as much fun with this as I had.

Cheers

Rik

Monday, 13 November 2017

Podcast Interview with Nicolas Mervaillie, GraphAware

Here's another great interview with a long time Graphista that has done a lot of really interesting work in our French community, and is now having lots of graph-fun at GraphAware: Nicolas Mervaillie. Nicolas has been and still is working on some really cool stuff - so a chat was long overdue! Here's our recording - including a fancy new jingle from PremiumBeat ("Fantastic Voyage", by Olive Musique):

Here's the transcript of our conversation:
RVB: 00:00:03.275 Hello everyone, my name is Rik, Rik Van Bruggen from Neo and here I am again on a Skype call, recording the next episode in our Graphistania podcast. And today I have a-- I would say an oldtimer in our Neo4j community on the other side of this Skype call, all the way from Lille in France, and that's Nicolas Mervaillie. Hey Nicolas, how are you? 
NM: 00:00:26.083 Hey, good morning Rik. Thanks for inviting me.  

Tuesday, 27 October 2015

Podcast Interview with Luanne Misquitta, GraphAware

A few weeks ago I had a great conversation with Luanne Misquitta from GraphAware. Luanne has been part of our community for a long time, has written a lot on the Neo4j blog, and is just a generally lovely person to talk to. So we spent some time on a Skype call, and this is what that conversation went like:

Of course, there's also a transcript of our conversation:
RVB: 00:01 Hello everyone. My name is Rik - Rik van Bruggen - from Neo Technology,  and here I am recording a podcast with someone from a very different part of the world - all the way from India, in Mumbai - Luanne Misquitta. Hi Luanne. 
LM: 00:17 Hi Rik, how are you? 
RVB: 00:18 I'm very well. Thanks for joining us. It's great to have someone from the Eastern part of the world joining us in the podcast [chuckles]. 
LM: 00:27 I'm glad to be here. 
RVB: 00:29 Yeah, super. Luanne, I think we've known each other for a couple of years. I first got to know you when you were writing about Flavor Networks, which I did a couple of blog posts about, as well. Actually, they were one of my most popular blog posts. How did you get into graphs, Luanne? Could you explain it to us? 
LM: 00:52 I've been working with graphs for about six years. I think that's close to the time there was this Neo4j challenge running to write an application on Heroku using Neo4j, and Flavorwocky was the application that I submitted, that eventually I won the challenge. Some time before that, I got into graphs because I was working for a company that had trouble managing the profile of a person. This was a people management company so dealt with all kinds of aspects - hiring, talent management, compensation, learning, all kinds of things. And central to this entire system was the concept of the profile of a person. Up until the time that I found Neo4j, it was modelled in a relational database, and at some point it was the table that stored data for the person profile.  It was huge. It had over 200 columns, and most of them were null, and it was normalized and denormalized very frequently, and it just didn't make sense any more. So, I was looking for something to solve this problem and that's when I came across this thing called a graph database, and of course, that was Neo4j. At that time, there wasn't any Cypher, there wasn't the Neo4j server either, it was just good old embedded. 
RVB: 02:27 Good old embedded, yeah [chuckles]. 
LM: 02:30 And immediately I knew that was the thing to model this, because it was schema-free, there were direct connections between a person and everything he's done. So one of the problems we have is that when you're collecting information about a person's profile, there are no two profiles that are the same in terms of the kinds of attributes they capture on a person. Of course, there are the usual standard ones - your name and where you live - but there are things like gaps in employment history, there are gaps in their learning. There are different types of learning. You do your talent manage-- your talent section looks quite different depending on where you are, and what you are doing, in which company you are in. So these kinds of things are very hard to fit into rows and columns, and that's where I thought Neo4j fit in. Once I started with it, I have never looked back. It's really been the solution for most things. 
RVB: 03:27 Yeah, it's quite a common thing, right? You've got a really complex domain that really is so difficult to model and store, query in traditional relational systems and then you take it to the graph world and a-- 
LM: 03:42 That's right. 
RVB: 03:42 --whole range of problems go away. 
LM: 03:44 That's right. 
RVB: 03:46 It's very, very cool. And how did you do the-- 
LM: 03:48 It wasn't just modelling the person. I'm sorry. 
RVB: 03:52 No, go for it. 
LM: 03:54 It wasn't just modelling the person, but just the fact that we could model it this way  exposed a whole lot of insights that we did not think about earlier, and for that kind of-- for the company that I was in that was really important. So, you weren't just interested in a model of a person and then spitting it back out to whoever is looking at your profile, but you wanted insights into potential jobs that this person might be good for, or if he's got a high flight risk, then who is going to replace him, what sales does he match. Is he in the wrong job at this time? Is he really sitting in your company working in a job that he hates, and he's blogging about other things? And the last one I mentioned it was real use case. When I did a quick POC in about two hours one night, and I just picked up random data from one of our internal company social networks, and the first thing that surfaced was this guy who was sitting and doing java development, there quite miserable, but he was blogging about and doing sample applications on iOS. At the very same time, the very same company was looking for people to work in their new iOS team. And here he was sitting under their noses, and no one knew about him. So these were the kind of things you could source immediately, which staring at a table-- 
RVB: 05:19 You would never find it. 
LM: 05:20 Impossible. Almost impossible. 
RVB: 05:22 Totally right, very cool. So I'm hearing the modelling advantages but what did the flaring-- the flavor example have to do with that? That was an interesting one as well? How did you get into that one? Are you a chef? 
LM: 05:39 Well, I love food. No, I'm nowhere near a chef but I love food, I love cooking. And I was thinking of what am I going to submit to this contest? Most of the enterprise-y things were done and I was quite bored with enterprises at that time.
Then I thought about food, I like food. Something that I had been reading about recently, actually a book review called the Flavor Bible, which lists ingredients and what they pair with. The rest of which from the classical triangle. So, two ingredients  paired with a third and you can then combine them and produce something that tastes really good. So I thought that is a good graph problem, and it was really simple, but it was a domain that I liked and I thought it was fun to work with. That's how Flavorwocky came about. 
RVB: 06:40 Luanne, recently - more recently, at least - you actually started working full-time with Neo projects at GraphAware, right? You're part of the GraphAware team-- 
LM: 06:49 That is right, I work for GraphAware now. Yeah. 
RVB: 06:53 How is that going for you [chuckles]? 
LM: 06:56 Very great. There can't be anything better than  working with Neo4j all day. At the moment, I am part of the Spring Data Neo4j team, and as you know, we've just released our Spring Data Neo4j 4 in September. So it's going really good. We've produced some great stuff and-- 
RVB: 07:18 Very cool. I am a big fan. Last night I met up with Michal in London and we were doing some Czech beer tasting [chuckles]. It was an eventful evening. Very cool. 
LM: 07:33  What do you know? 
RVB: 07:35 No, don't get started on that one [chuckles]. That's going to be a very long podcast then. No. Let's talk about the future, Luanne. What does the future hold? You've been involved in some really exciting projects like, for example, Spring Data Neo4j. Where do you see this going? Can you share some of your perspectives? 
LM: 08:00 Well, I'll share with you what I've been noticing over the past couple of years and-- and as you know, I do the trainings for Neo4j in India as well, so I've been looking at various kinds of audiences coming in through for at least two years now. And there has been a definite shift in attitude towards graph databases from the early days where it was a real struggle to get people to understand why it's important. Although they got it, it was still something that they would really have to fight very hard to use in their organization. That has changed significantly over the last, well, even the last six months really, where you have people who already know about graph databases, they can immediately see where it's going to fit into solving their problems. I think with the kind of-- the reputation that Neo4j has , it's now becoming easier to get Neo4j are used in these kinds of companies, and I'm not talking about startups which pick up Neo4j very quickly, the other larger, mid sized company or enterprises. I think what I would like to see at some point and fairly soon is when graph databases become something that you just use. If you're planning a new project, it's very common to say, "Hey, we need a database," and no one will challenge you. You'll go and you'll pick up a database and you'll use it, typically an RDBMS. If you were to say, "Hey, we need a graph database," then it should be exactly that. Pick it up and use it. You shouldn't have to be debating for months over whether it's better than sequel or not and whether it fits or not. I'm really looking at that day as a defining moment for graph databases  where they are just used, and people know when to use them and why to use them, and there aren't any questions about should I or should I not, unless it's a really stupid use case. That's what I think would be the ultimate future for Neo4j and graph databases. 
RVB: 10:20 I'm looking forward to that day as well. And I think lots of us work towards that goal and it would be a great thing. What about the Spring Data stuff? Is that ready for prime time right now? Are you guys planning new stuff there? 
LM: 10:37 Yes, it is. The Spring Data for that been released in September is, of course, ready. We are always planning new things. We have support for-- as you know, Spring Data 4 supports currently the remote Neo4j server mode only, and it was written from ground-up to actually support that, so it's  really fast. We've broken it up into-- it's not only Spring Data Neo4j 4. It actually depends on a new library called the Neo4j Object Graph Mapper. So, if you want really fast OGM, then-- and not Spring, you can use the OGM directly. If you're a Spring person, then, of course, Spring Data Neo4j really uses the OGM under the covers, so there are a lot of features planned for both of those two. Very shortly, one of them is support for embedded Neo4j, and a whole lot of the new protocol, which is [crosstalk] coming up in Neo4j. Support for those, and as well as-- there is so much to do, but we are continuing to work on that and you should see some releases out pretty soon I hope. 
RVB: 11:51 Super. And I'm hoping that I'll see you at GraphConnect in San Francisco? 
LM: 11:55 I think you will, yeah. 
RVB: 11:57 Yeah, well, looking forward. That's super, great. All right, thank you for coming on the podcast, Luanne. It's been a great chat and I really appreciate it. 
LM: 12:06 Thank you Rik - it was great talking to you again. 
RVB: 12:08 Same here, and I look forward to seeing you on the West Coast. 
LM: 12:12 Yeah, soon. 
RVB: 12:14 Cheers, bye. 
LM: 12:15 Bye-bye.
Subscribing to the podcast is easy: just add the rss feed or add us in iTunes! Hope you'll enjoy it!

All the best

Rik

Thursday, 21 May 2015

Cycling Tweets Part 4: Ranking the Nodes

In the previous couple of blogposts in this series (here's part 1, part 2 and part 3 for you), I have explained how I got into the Cycling Twitterverse, how I imported data from a mix of sources (CQ Ranking, TwitterExport, and a Python script talking to the Twitter API), and thereby constructed a really interesting graph around Cycling.

There's so many more things to do with this dataset. But in this post, I want to explore something that I have been wanting to experiment with for a while: The GraphAware Framework. Michal and his team have been doing some real cool stuff with us in the past couple of years, not in the least the creation of a couple of very nice add-ons/plugins to the Neo4j server.

One of these modules is the "NodeRank" module. This implements the famous "PageRank" algorithm that made Google what it is today.
It does this in a very smart way - and also very unintrusively, utilising only excess capacity on your Neo4j server. It's really easy to use. All you need to do is

  • drop the runtimes in the Neo4j ./plugins directory
  • activate the runtimes in the Neo4j.properties file that you find in your Neo4j ./conf directory. 
Here's what I added to my server (also available on github):

//Add this to the <your neo4j directory>/conf/neo4j.properties after adding //graphaware-noderank-2.2.1.30.2.jar and //graphaware-server-enterprise-all-2.2.1.30.jar //to <your neo4j directory>/plugins directory   com.graphaware.runtime.enabled=true  #NR becomes the module ID: com.graphaware.module.NR.1=com.graphaware.module.noderank.NodeRankModuleBootstrapper   #optional number of top ranked nodes to remember, the default is 10 com.graphaware.module.NR.maxTopRankNodes=50   #optional damping factor, which is a number p such that a random node will be selected at any step of the algorithm #with the probability 1-p (as opposed to following a random relationship). The default is 0.85 com.graphaware.module.NR.dampingFactor=0.85   #optional key of the property that gets written to the ranked nodes, default is "nodeRank" com.graphaware.module.NR.propertyKey=nodeRank   #optionally specify nodes to rank using an expression-based node inclusion policy, default is all business (i.e. non-framework-internal) nodes com.graphaware.module.NR.node=hasLabel('Handle')   #optionally specify relationships to follow using an expression-based relationship inclusion policy, default is all business (i.e. non-framework-internal) relationships com.graphaware.module.NR.relationship=isType('FOLLOWS') #NR becomes the module ID: com.graphaware.module.TR.2=com.graphaware.module.noderank.NodeRankModuleBootstrapper   #optional number of top ranked nodes to remember, the default is 10 com.graphaware.module.TR.maxTopRankNodes=50   #optional damping factor, which is a number p such that a random node will be selected at any step of the algorithm #with the probability 1-p (as opposed to following a random relationship). The default is 0.85 com.graphaware.module.TR.dampingFactor=0.85   #optional key of the property that gets written to the ranked nodes, default is "nodeRank" com.graphaware.module.TR.propertyKey=topicRank   #optionally specify nodes to rank using an expression-based node inclusion policy, default is all business (i.e. non-framework-internal) nodes com.graphaware.module.TR.node=hasLabel('Hashtag')   #optionally specify relationships to follow using an expression-based relationship inclusion policy, default is all business (i.e. non-framework-internal) relationships com.graphaware.module.TR.relationship=isType('MENTIONED_IN')
As you can see from the above, I have two instances of the NodeRank module active. 
  1. The first attempts to get a feel for the importance of "Nodes" (in this case, the nodes with label "Handle") by calculating the nodeRank along the "FOLLOWS" relationships. After just half an hour of "ranking" we get a pretty good feel:

    This seems to be confirming - in my humble opinion - some of the more successful riders in April, for sure. But also confirms that the "big names" (Contador, Froome, Cancellara) are attracting their share of Twitter activity no matter what.
  2. The second does the same for the "Topics" (in this case, the nodes with the label "Hashtag") along the the "MENTIONED_IN" relationships.

    The classic races are clearly "top of mind" in the Twitterverse! But upon investigation I have also found that there are a lot of confusing #hashtags out there that make it difficult to understand the really important ones. Would love to investigate a bit more there.
Like I said before, the GraphAware framework is really interesting. It gives you the opportunity to make stuff that you could also do in Cypher more easily, faster, and more consistently. I really liked my experience with it.

Hope this was useful for you - as always feedback is very very welcome.

Cheers

Rik

Friday, 24 April 2015

Podcast interview with Michal Bachman, GraphAware

Here's another great conversation for you in our Neo4j Graph Database podcast series. I met up with Michal Bachman of GraphAware, one of the awesome Neo4j partners out there. Michal and I have been working together on different projects, presentations and beer-tastings - and I am happy to say that his capabilities, visions and strategies when it comes to Graphs and Neo4j are WAY better than when it comes to his taste of beer :) ... Listen to the interview and find out why:

Here's the transcript of our conversation:
RVB: Hello, everyone. This is Rik. Welcome, again, to one of our Neo4j graph data-base podcasts. It's another remote session. I'm joined today by Michal Bachman of GraphAware. Hi Michal. 
MB: Hi Rik, thanks very much for inviting me. 
RVB: Yeah, absolutely. It's great to have you on the podcast. Michal maybe people don't know you yet, so why don't you introduce yourself? Who are you? 
MB: Sure. My name is Michal Bachman and I'm the founder and managing director of a company called GraphAware, which is a London based company dedicated to Neo4j consultancy, training and development. Being based in London we are in a great position to travel around the whole world pretty much and help people succeed with Neo4j. That's what we do for a living. 
RVB: Absolutely. I can hear the London Police in the background [laughter]. That's absolutely great. Thanks, Michal. How did you get to graphs and how did you get to Neo4j? Tell us a little bit about that and what attracted you. What do you love about graphs? 
MB: I started with Neo as a user, pretty much, about four or five years ago. I was involved in a few projects, in fact, that used Neo. One was a recommendation engine, and another one was an impact analysis solution for one of the large telcos. And I really liked the experience as a user, and I then went on and took a bit of a break, and did a master's degree at Imperial College, London where I wrote a thesis on graph databases. 
RVB: Oh, yeah? 
MB: Yeah. Quite inspired by Jim Webber and his idea. That was great, and I loved it. I loved the experience as a user. I loved doing research about it, so the natural next step was to start my own company that will focus only on Neo4j. That's how I pretty much started, and it's been everyday [chuckles]. 
RVB: Absolutely. What attracted you most? What did you like most about working with graphs and Neo4j, specifically? 
MB: The actual thing that I liked the most is, surprisingly not a technical thing. It's the fact that when you introduce people to graphs, and we are doing that every day, you can see the moment - the "Ah" moment - in their eyes. 
RVB: The lights come on [chuckles]. 
MB: Yeah. The lights come on, and then they're like, why haven't I used this before? This is not just like another 10% better way of storing data. This is a complete game changer, and people seemed to get it immediately, and it's applicable to every domain out there. There's a huge potential, and I just liked the fact that you know when people get it. They just fall in love with it. 
RVB: Just to follow onto that, one of my Dutch community members, or community members in the Dutch graph database community, once told me, "Once you start working with graphs, relational databases feel like a youthful sin," [chuckles]. 
MB: [laughter] Yeah. And it makes so much sense if you think about it. Most people work with object-oriented languages, and objects are graphs. Everything is a graph, so it just feels so natural after you've made that transition. 
RVB: Tell me a little bit more about GraphAware now. You guys have a wonderful graph framework these days, right? The GraphAware Framework. Tell me a little bit more about it. 
MB: We've doing two things, really. We've been doing consultancy, as you know. We are involved in projects, very hands-on, helping customers develop software with Neo4j. And as we are gaining more experience about what the use cases are and what people need, we're distilling some of those ideas and experiences into open source extensions for Neo4j. That's the two things, and the third one of course is training. We're running also community events, but also we run public trainings. In the future, we're seeing doing more of the actual extension development and open source software built on top of Neo as the way to go. 
RVB: What are some of the functionality of the framework, in just two minutes? 
MB: One that we recently released, and that's getting quite popular, and running meetups around it as well, is a recommendation engine extension that allows people to build quite complex high-performance engines on top of Neo. That's one. And the other ones are quite technical. We've got modules for representing time as time series data in Neo4j as a tree and easy creating, and there's loads of other modules to main specific use cases. 
RVB: I'll put a link to the repo on the blogpost to go with the podcast so people can take a look at that. Let's maybe move on a little bit. Where is it going, Michal? Where are you guys going as GraphAware, but also where do you see the industry going? Any perspectives that you want to share? 
MB: Absolutely, I think we're going to see, and we are going to see quite soon, this technology being adopted by large enterprises in a massive scale. And as that's happening, I'm seeing some enterprise features, more of the enterprises features being developed, whether the part of the core product or as extensions, so that companies like banks and trans companies, and so on, find it easier to use. I'm talking about security, auditing and things like that. And I see people starting to build whole platforms around the graph use cases, include graph-compute engines in them, include other great software to build whole data analytic platforms, where the graphics is the center of the game. And extensions for impact analysis, fraud detection, recommendations, complete solutions, I think it's what we're going to be seeing in the near future. 
RVB: Very cool. Okay. One more question for you, and it's the most important one. What do you prefer best, Belgian beer or Czech beer? 
MB: [laughter] I have to be honest with you, I prefer Czech beer [laughter]. 
RVB: Oh my God, I can't believe that! All right, thank you so much for coming on the podcast Michal, it was great having you. 
MB: Thank you, Rik, for inviting me. I want to say one last thing. We're of course going to be present at the Graph Connect. We're sponsoring the conference, 7th of May, we're going to be there, so if anyone's interested in having a chat with us, please come to Graph Connect in London, and we'll see you there. 
RVB: Yeah. Absolutely. Thanks a lot, Michal. Talk to you soon, man. 
MB: Thanks, Rik. Bye-bye.
Subscribing to the podcast is easy: just add the rss feed or add us in iTunes! Hope you'll enjoy it!

All the best

Rik