Showing posts with label rdf. Show all posts
Showing posts with label rdf. Show all posts

Monday, 15 March 2021

Graphistania 2.0: The Lockdown Anniversary session

This weekend marked the 1 year anniversary of the "Covid era" - the time when many of us have been hunkered down at home, or close to home at least, to deal with the raging pandemic. It has been a strange couple of months, and while I personally have been able to deal with it quite comfortably, I must say that my heart has been going out to all the friends and families that have had a far worse time with this. I personally got to experience it first hand in the first couple of months of 2021, loosing a very dear friend to the virus and its awful disease, and will not easily forget this period of our lives.

But we do want to keep the fantastic drive and atmosphere of our Neo4j community up - it would be a terrible shame if we lost that, too, to the virus. So that's why me and Stefan are going to continue making these podcast episodes, at least for the foreseeable future. Not in the least because they are a ton of fun to make :) ... 

So here's our latest chat. We have found a true treasure trove of great use cases and such in the Twin4j newsletter, which we will try to highlight:

Hope you will find our chat interesting! Here it goes:


RVB: 00:00:00.768 Hello, everyone. My name is Rik, Rik Van Bruggen from Neo4j, and yep, it's that time again. Yippee! Yeah. We have another podcast recording day, and for that, I have my dear friend Stefan on the other side of this Zoom call. Hey, Stefan.

SW: 00:00:18.130 Hello, Rik. Nice to be back here with you.

RVB: 00:00:20.830 Hey, there.

SW: 00:00:21.733 Coming back from a week of vacation, so very nice to be back, and what better way to get your graph mind up and running than to hang out with you here in this lovely podcast?

RVB: 00:00:33.018 I hope so too. Thanks for being here, and I hope you had a good holiday. And it's been two months, Stefan, so we really need to get our act together. It's been a very busy--

SW: 00:00:43.185 Holy crap.

RVB: 00:00:44.057 --couple of months, but. We try to do these things once a month, but that didn't work in February, and so we've got a lot to talk about, actually. And maybe I'll just kind of frame it for you. I went through all of the Twin4j this week, the Neo4j newsletters over the past couple of weeks, and I found some really interesting themes. And maybe we can talk about those for a bit. There's three of them. Is that okay?

SW: 00:01:16.179 Yeah, yeah. Let's go. I think that should be super fun.

RVB: 00:01:20.885 Yeah. So the first one I wanted to talk about is, basically, use cases. I mean, we see this all the time in our community, really, right? That there's these unbelievable interesting use cases popping up. There's a couple of them that popped up in the newsletters. The first one is actually dear to my heart. It's something that I started working on back in 2012 when I first joined Neo4j. It's protein interaction networks. Have you taken a look at that one?

SW: 00:01:50.023 Yeah. That was an amazing one. As always with those kind of things, when there is-- this was written by Tomaz, right? So I started--

RVB: 00:02:01.558 Tomaz, [crosstalk].

SW: 00:02:02.277 --reading the article, but then halfway through the article, I was like, "Oh, but I better just try it out." That's how I learn. So I ended up running this, and it is so neat that this is, in some sort of way, so accessible. I think for me, that is super cool in that. So I find it extremely interesting and also because of the simplicity of it and how complex it is, even if it's just such a simple thing as a protein that interacts with another protein, basically, at the foundation, and it's still amazing. I think it's kind of mind boggling in that sense, so.

RVB: 00:02:44.678 A little story from my side: when I first started working with Neo4j, we did some work with the University of Ghent here in Belgium, and they were working on a topic called metaproteomics, which is exactly this, interactions of proteins. And they struck a nerve with me because one of their most important research customers was, of course - drumroll - a brewery.

SW: 00:03:13.135 Of course. I was wondering, "Where is this going to go?"

RVB: 00:03:17.022 Yes.

SW: 00:03:17.797 There it was. Brewery time again.

RVB: 00:03:18.825 There it is. Yeah.

SW: 00:03:20.664 There it is.

RVB: 00:03:21.182 It was a beer brewery, and they were basically saying yeasts that are being added to brewing systems, they create these protein interactions, and if you better manage those protein interactions, you can actually influence the brewing and the taste of the brew by doing so. So that was an interesting one. I had a good time exploring that. There was another one, Stefan, around asset management. This is a very well-known one as well, right? Things like configuration management databases, building information management. It's all about managing assets, isn't it? It's a very networked problem.

SW: 00:04:06.005 Yeah, yeah. And also, this is one of the use cases that seems to pop up more and more often in customer or prospect interactions, I think, so I think it's going to be very helpful for people. So go check it out if you are interested in that. I think that would be very neat to do that.

RVB: 00:04:29.195 But I know you want to skip to the third one that we had [crosstalk].

SW: 00:04:32.095 Yeah, the big one. This is what I'm kind of waiting for, like, "How can we get to this point faster?" Ha ha!

RVB: 00:04:37.534 Yeah. "How can we get to that third point faster?" which is, of course, the use case of getting to Mars and NASA. It was all over the news in the past couple of weeks, obviously, but yeah, that story of how David Meza from NASA-- he's the - how do you call it? - chief knowledge management architect or something like that with NASA, and he worked on that Lessons Learned database in Neo4j. Super cool, right? It's so good.

SW: 00:05:12.489 Yeah. It's super good. I think the use case is good. The interview is also great. David is also very relaxed, I think, also. Ashley - or what is the name of the girl interviewing? - is also doing a great job. There's very good chemistry in there, really enjoyed it. And again, of course, anyone that was dreaming about going to space as a kid, imagine working with such a thing. I mean, it can't get better. This is the moment when you go to work, and you go, "Holy crap. I'm so proud now." But I think it's an interesting thing on how much this actually speed up time for them, right? How much is saved not only time, in that sense, but also taxpayers' money, right?

RVB: 00:05:59.053 Of course.

SW: 00:05:59.211 I think that's what I keep coming back to in talking about use cases. So you can do a lot of things with a lot of technologies, right? So very often, people ask me, "How can I use Neo?" Right? But then I say, "You can do it for this, but you can, of course, do this with your old technology, in theory. However, if that theory takes you two years, maybe you can't really do it in practice." So I keep coming back to think about that, and I think this is such a good showcaser on that. It's a great YouTube clip here, so for those that have a hard time reading or just want to listen while-- I was saying commuting to work. Apparently, commuting to your working room should be better now, in these times. But I really enjoyed it. Happy to see it, so yeah, hope you like it as well.

RVB: 00:06:49.604 Cool. Yeah. It was super nice. And there's another, I mean, theme to the newsletters in the past couple of months. I've seen so many-- it's kind of like a use case, but it's also a technology foundation of people that are using graphs in combination with natural language processing, right? I saw a number of posts from Jesús, our colleague, who was talking about RDF-related work, WordNet, those types of things, but there's also people that are doing really interesting work on extracting new knowledge from existing documents, right? Were you able to make any sense of that?

SW: 00:07:39.354 Yeah. And I think it's like this is also one of those kind of super untapped-- and I think it's also a perfect bridge from NASA, right? Because literally, that was what they were doing. They had the answers; they just couldn't see it, right? So it's a classical, "You can't see the forest because of all the trees," right? So I think that is really interesting. And I think also, there is a great post about-- I think it was called From Text to Knowledge: The Information Extraction Pipeline or something, basically where Tomaz then explained why he see a combination of NLP and graphs as one path to explainable AI, right? And I think this is also one of those topics that are super important from a lot of things to kind of understand but also compliance and a lot of things, right? So I think this is also one of those kind of areas, use cases, or whatever you want to call it that literally are exploding on all different kind of verticals, you may almost use as a word there. But yeah, I think that, again, our--

RVB: 00:08:44.103 Yeah. Well, it's been very popular in domains where there's a lot of documentation, right? So academics, pharmaceuticals, patterns, legal texts, all those things have been really a great showcase for this type of work, I think.

SW: 00:09:04.354 Yeah. No, but as you're saying, I can see anything from academia, patterns-- I mean, I don't know how many of these kind of works that we have done with prospects and really kind of tapping into this kind of super kind of deep knowledge, but you can't see the new perspective because it is just a lot of deep silos, right? So in that sense, it's super graphy, and I think when people start to see it, this is also where they get so excited so were almost screaming, "Take my money," and I was like, "Calm down. Behave good. What is it that you're thinking of answering, or what is the thing that would help you," right? It doesn't have to be the money query, but just don't throw technology at the problem. So be mindful of what you want. I mean, we can see it a lot. I think one of the interesting parts is also looking upon the entire web, scraping information and making sense of it and treating that as a kind of a knowledge grab itself. It's also one of a neat couple of projects that I'm working on personally. Yeah. So there's a lot of funny things to do with the NLP and knowledge graphing combination, I think.

RVB: 00:10:19.330 Yep, totally. Well, what strikes me there-- and this is a perfect segue to our third theme. What strikes me is that it's becoming so much easier to do this, right? So NLP a couple of years ago, that was just so exotic and difficult to use. You basically had to have some kind of a computer science degree or a PhD to be able to use it, but these days, the tools that we have to implement some of these techniques are super accessible. I mean, even a lost sales guy like me can use it. Do you know what I mean? It's pretty usable, even in its basic form. So I wanted to talk about some of the really interesting tools that we see emerging in our community. The one that I was so happy about, to finally see it fully released, is the Arrows app. I think you've used it for a long time already when it was still in an alpha or beta stage, and now it's--

SW: 00:11:26.893 Exactly.

RVB: 00:11:27.278 --actually been released. Alistair Jones' pet project is now finally out there in the wild. Great, though, right? I mean, really great.

SW: 00:11:37.420 Yeah. No, but I think it's so useful in so many ways, of course, for graph modelling, where it was intended, right? And I have a couple of memories working with a lot of C-level people trying to help them understand the power of connected data and so on. So they're not going to write any code, but what we tend to do is do some modelling practices and basic cipher, and for that, I use Arrows. And one of the constant feedback that I get because of the simplicity of the tool and how it naturally kind of lends your thinking to this kind of graph thinking idea is that, every single one of these sessions, these CEOs, COOs are coming back to me like, "Oh, this is a really good way of thinking of the business, the different domain, and how it's connected. It has given me tons of new ideas." So in that sense, that kind of doubled down as a ideation kind of tool almost, if it makes sense, but I think that's such a cool kind of thing, right? If you're going to build an app or if you have a business logic or anything, that really helps you to kind of map it out in that sense. So it's also kind of neat to see that and to see those people, also, stepping into the graph arena. But yeah.

RVB: 00:12:53.804 Yep. That frame also. Yeah. So Arrows is one of those really amazing tools, but there's other things that are coming up, right? We've all known about the GRANDstack, the development framework that's been around for quite some time. A new release for that and new features, capabilities there, and some examples, also, for people to use and to abuse, I would say.

SW: 00:13:18.057 Use and abuse. That's a perfect way to do it.

RVB: 00:13:19.407 Use and abuse, yeah.

SW: 00:13:21.935 Just dive in there and try.

RVB: 00:13:22.283 And then another one that I wanted to mention was I've really enjoyed using this tool that Niels created called NeoDash. It's a graph app that plugs into the Neo4j browser and allows you to put together dashboards on top of Neo4j. Really no-code development, that type of thing, super easy to use. I was very impressed by that, how accessible it's made everything.

SW: 00:13:52.996 Yeah. And I think that's such a great part because I think just the ability to do something without that no code. Of course, that is in the topic with the cloud itself, one of those kind of things that you can really see how accessible these are. But seeing people with no previous knowledge just trying and fiddling around there, they kind of stumble upon the solution almost. That simple. I think the NeoDash is such a great kind of application for just exploring or visualising the data that you have in a graph that you normally would not use. So we tend to try it out with a lot of the nontechnical people when we work, and it works like a charm every single time. There is literally almost no studying kind of to get started, so I again encourage you, as with all articles, use and abuse, right? Dive in, try around because actually, you're going to get kind of far by just doing that. And that is, I think, a common theme for all of these. You can really see how this whole paradigm of connected data is changing, which is, of course, super cool.

RVB: 00:15:17.502 Yeah, absolutely. Well, I mean, so many other things to talk about. What we'll do is when we get to the blog post created together with this recording, we'll also put all the links to the amazing Twin4j newsletter items and the blog posts and everything all together, and then people can have a look, have a play, use and abuse, and that should set them on their path for even more graph adoption, right? So that's the [crosstalk] idea.

SW: 00:15:50.313 Even more graph adoption. Yeah. And I'm also going to squeeze in a last one because we also had the GDS 1.5 release, right? And there's a great piece on it from Amy and Alicia, two of my most inspiring colleagues. They have taught me a lot of things and are super nice as well. So it's about the new supervised machine learning workflows in Neo4j, so imagine that being even accessible to just try. I can't even think of this. If I would have guessed this five years ago, I'd be like, "Nah, that's not going to happen."

RVB: 00:16:29.543 It was impossible. Yeah, exactly.

SW: 00:16:31.365 "That's impossible. You can't do that on your computer at home in your sofa." But I think that's so cool. So that's a great article by Amy and Alicia; go check it out as well. I'm going to push that in there, but we're going to, of course, post the links, as I said.

RVB: 00:16:48.591 We will, for sure. Well, Stefan, thank you so much for taking the time running through this with me and making a little bit more sense out of it. It was great talking to you. You know that we want to keep these things shortish, at least, so we're going to wrap it up for now, and we're going to try to have this one published soon and then do another one in April, right? We should [crosstalk].

SW: 00:17:14.412 Yes, of course, 1st of April. I will record it from my new podcast studio in the barn in [inaudible] in southern Sweden.

RVB: 00:17:22.097 Ooh.

SW: 00:17:23.410 Ooh, yes. And then--

RVB: 00:17:25.195 I look forward to that one.

SW: 00:17:26.657 --as soon as we get to travel, this is a standing invitation for you to come join me and also for any of our listeners. Bear in mind--

RVB: 00:17:37.641 Absolutely [crosstalk].

SW: 00:17:37.964 --COVID restrictions has to be better before that, so I am not encouraging any anti-vaccine behaviour here, but as soon as we are allowed to travel, come join us. It's going to be a great talk about graphs.

RVB: 00:17:51.294 Fantastic. Thank you, Stefan. It was great talking to you, and I'll talk to you soon.

SW: 00:17:55.905 Likewise. Bye.

RVB: 00:17:57.591 Bye.

Hope you enjoyed that as much as we did. If you have any comments or questions, just reach out!

All the best

Rik & Stefan

Tuesday, 6 October 2015

Podcast Interview with Jesus Barrasa, Neo Technology

Over the past couple of months, we have been doing a lot of work at Neo4j to try to better explain the value of Neo4j to our prospects and customers. This has been a true team effort, with lots of engineers, marketeers and sales folks participating in articulating how complex technology can be used to add true value to business processes. And as we did that, we added some really talented people to our team. One of them is the person that I am interviewing in this particular podcast episode: Jesus Barrasa. Jesus has a lot of experience with graphs and even (a particular kind of) graph databases - so going in I knew it was going to be an interesting chat. And guess what: it was. Listen on:

Here's the transcript of our conversation:
RVB: 00:00 Hello everyone. My name is Rik, Rik Van Bruggen from Neo and here I am again recording a Neo4j graph database podcast and my guest today is all the way in the UK. Jesus, hi Jesus. 
JB: 00:13 Hi Rik. How are you? 
RVB: 00:14 Hey. I'm always scared of pronouncing your name in the wrong way. I'm sorry. 
JB: 00:19 You did great. You did great [laughter]. I've heard much worse than that so that was brilliant. 
RVB: 00:24 Okay [chuckles]. Okay, Jesus, you just joined Neo a couple of months ago as a pre-sales engineer but you have a long-standing history with graphs. You did a PhD on the subject if I'm not mistaken, right? 
JB: 00:37 That's correct, yeah. It all started probably more than ten, nearly 15 years ago, so quite a while yeah. 
RVB: 00:43 Yeah. 
JB: 00:43 And you're right. It all started in the semantic technology space, in the RDF space, so that was the first time I was exposed to modeling data as graphs and yeah. That's been a long story. After the PhD I did work for a company in London called Ontology where we did use graphs to best resolve problems in companies, mostly in the telecommunications sector and yeah. As you say, two months ago I joined the field engineering team in the London office. 
RVB: 01:18 That's super. What was your PhD about exactly then? 
JB: 01:21 Well, at the time… I did model-- I mean, I formalized a way of translating relational schemas into ontologies. Ontologies is the way you represent metadata in the RDF model. I don't know if we will have time to talk about that but yeah, it's actually an automated mapping between relational schemas and ontologies. That's what I-- 
RVB: 01:45 There's a lot of people looking at that I think, still today, I get that question quite often. 
JB: 01:50 Absolutely. 
RVB: 01:51 So another question because you've been working so long in the RDF world, what's the difference between the RDF semantic technology space and the property graph model of Neo4j - what's the key difference for you? 
JB: 02:04 All right. I'd say you can answer this question from two perspectives because obviously RDF is a presentation paradigm and there's the different implementations, the triple stores that you can find in the market but as a model I think the main thing they have in common is the fact that they both represent data as a graph. And that makes them very, very close to each other. The difference I would say is RDF is simple as it can be. It's only based on the notion of URIs to identify resources or nodes if you want to establish the properties with a property graph but there's this element called triple, subject, predicate, object. And that's all the constructs that you have to model your domain. And of course I think that's the biggest difference because in the property graph model, you have nodes with properties, you have relationships with the properties and they have this brilliant thing that's this white-board friendliness, this excellent thing that's the way you conceive it, the way you model it in your head, in your whiteboard, is exactly the way it's represented and stored physically, whereas in RDF there's still this gap where you have to translate that into triples, things that may sound simple like giving a weight to a relationship so there's a connection between Rik and Jesus because we work together. You want to give weight to this relationship. That's something that's completely natural in the property graph but it's not something you can do directly on a triple, so you have to model this relationship probably as an intermediate resource. There's a bit of a gap and I think that makes it sometimes less intuitive, sometimes a bit less humane if you want. So that's the difference. 
RVB: 04:06 Well [crosstalk] actually that was sort of what I was hoping you would say because I've always been told that the difference is really on the predicate, the fact that it's so difficult to qualify a predicate in the RDF model. 
JB: 04:20 That's [crosstalk] exactly. Then there are other interesting things in RDF, the whole idea of being able to use the model itself to represent the metadata, the ontology and that gives you in certain cases some interesting powerful things you can do like querying all the data and the metadata at the same time. But I would say the biggest thing is more what they have in common and it's the fact that they look at data as a graph, a set of connected nodes. 
RVB: 04:49 What attracted you to the graph in the first place? Why did you get into this field in the first place if you don't mind me asking? 
JB: 04:55 Well, yeah, sure. I think there's two things. One was this incredible flexibility. At the time when I started we didn't talk about NOSQL. That's a concept that was coined later. It's the idea of being able to start storing data without having to model it up front. I thought that was brilliant and that gives you an incredible flexibility, this thing of schemaless model or at least implicit schema depends on how you talk about it but this flexibility was one of the things and the other one, I'd say the possibility of infering new knowledge based on the information on your graph, you identify patterns and you can enrich your graph with new knowledge based on the data elements that you have. So I think probably these two things attracted me to this world and I find them unbelievably useful and powerful. 
RVB: 05:51 You know, as your couple months at Neo, what do you think is the most interesting use-cases that you've seen so far? Anything that jumps out? 
JB: 06:01 Sure. Well, I'm amazed-- the thing is the great thing about Neo is how you can use and update your graph in real time, at what speed you can not only read it and query itbut also keep it up to date and how it's possible to identify fraud rings for example is one of the cases that I like the most, being able to pick up the status of your accounts, your users, their information, the transactions that they're carrying out, and at the same time be able to pick up, detect the patterns that identify a fraud ring is one of the ones that I enjoy the most. 
RVB: 06:43 And do that in real time you mean? 
JB: 06:44 Exactly, the real time is the key thing and that's pretty impressive yeah. 
RVB: 06:48 It's not like with my experience a couple of weeks ago with my Amex card, I get a fraudulent transaction on Friday and I get a call on Monday [laughter]. 
JB: 06:57 Oh yeah, just in time right. 
RVB: 06:59 Just in time [laughter]. Didn't really work. Very cool. So where do you see this going, Jesus, where's the industry taking us do you think? 
JB: 07:09 Well, I think adoption is growing. It's amazing the number of different organizations and companies that we talk to. I can't think of a single vertical, a single sector that would not benefit from scenarios where modeling data as a graph adds incredible value, so I think adoption will definitely grow and that's one of the things, and then the other one that I think is going to be key as well is about integrating the graph with the rest of the data architecture. These days there are so many alternatives to represent data and some of them adequate for certain scenarios. I'm all for peaceful co-existence with all the other approaches. So integration I think is going to be the other important one and I think being able to expose the graphing in ways that make it easy to inter-operate with other stores, sometimes not all of the information is going to be in your graph but the extremely valuable information in your graph will need to be combined with some external information. That's one case where you will want to visualize it in different ways using BI tools, well you name it, there's so many elements now in data architectures that integration I think is going to be the other important aspect that we'll see developing in the next few years. 
RVB: 08:22 There was one thing that I wanted to ask you and I obviously forgot. I'm so good at this podcasting thing [chuckles] is actually you've done a lot of work on virtualization of data, right and then [crosstalk] integrate and that links to that integration story I suppose. 
JB: 08:36 Exactly. Exactly. I did work in the data integration space with a data virtualization company in the couple of years between Ontology and Neo and yes, I'm particularly interested in that and it's a way of integration data virtualization that's based on this idea of wrapping the sources and make them look as if they were relational even though they're not so they're not copying the data into centralized stores. You leave the data where it is and you define the logic to extract it and combine it and I think that was a powerful paradigm for new ways of representing data like the graph and make it easy to integrate them with other technologies and other types of stores and yes. 
RVB: 09:27 So in a case like that the Neo4j would be one of the sources of virtualization? Is that what I'm--? 
JB: 09:32 That's correct [crosstalk]. One essential one, that's the thing because the importance in the end is what value is there in your source? Neo can be perfectly, for example, in an MDM scenario. It can be the core. It can be your master model. And you can link it with the different rest of it and provided the detailed information about the entities but exactly, it would be one of the sources and the data virtualization will expose it in a way that's easily consumable by say for example BI tools in analytic scenarios or that's one of the-- 
RVB: 10:05 Are there any examples of that yet? You know, like open-source virtualization tools that integrate with Neo [crosstalk]? 
JB: 10:10 Well, there's not much to be honest. There's one quite limited community edition of one of the vendors called Denodo which is the one I used to work for. There's another one from JBoss but I'd say there's not much, Rik, available in there. I mean, JBoss would be the obvious option. I'm actually now trying to work a little bit with it and try to build some integrations with Neo and yeah. That's what you can find. 
RVB: 10:40 I think this is kind of like community call for help, you know [crosstalk]. 
JB: 10:45 Yeah [crosstalk]. I definitely really hope to be publishing something soon, at least in some idea, some small examples that can inspire people to look at these. 
RVB: 10:55 That would be great. Cool. I think we're going to wrap up here. I think we like to keep these podcastsquite short and snappy but thanks a lot for sharing your perspective. I think that was very interesting although because of my limited presentation skills a little bit chaotic [laughter]. 
JB: 11:13 Right. No, it was great [chuckles]. Great to have this chat with you Rik. 
RVB: 11:15 Thank you, Jesus. And I'll see you soon yeah. 
JB: 11:18 Lovely. Cheers. 
RVB: 11:19 Bye. 
JB: 11:20 Bye now.
Subscribing to the podcast is easy: just add the rss feed or add us in iTunes! Hope you'll enjoy it!

All the best

Rik