About a year and a half ago, at some meeting in Malmö at Neo's Swedish HQ, I bumped into a new colleague there who was... a little different than most Swedes. He was a bit louder, a bit more outspoken, unapologetically sarcastic, VERY funny - and a big fan of all kinds of (Belgian and other) beers. So we started talking - over a beer. I think it's fair to say we hitted of - and I got to know this Canadian-gone-Swedish guy a bit more, talked more, drank more, ... and decided to invite him to talk about some of the interesting stuff that he works and worked on for our podcast. His name? Señor Baker. Senior Baker. Srbaker. Also know as Steven Baker - working for Neo. Here's our chat:
Showing posts with label neo technology. Show all posts
Showing posts with label neo technology. Show all posts
Thursday, 1 June 2017
Tuesday, 2 May 2017
Podcast Interview with Andrew Bowman, Neo Technology
BY FAR the most annoying thing about working for Neo4j, is that there are so many, MANY cool things to do. And that means that sometimes cool things fall through the crack. Like for example this podcast episode, which dates from March already - a great conversation with Andrew Bowman about his work in the Neo4j community. As it so happens, Andrew just recently joined our "Customer Success" team, and is now not just an active community member - but he can actually live and breathe Neo4j 24/7 now :)) ... Here's our chat:
Here's the transcript of our conversation:
Here's the transcript of our conversation:
RVB: 00:03.249 Hello, everyone. My name is Rik, Rik Van Bruggen from Neo Technology, and here I am again, recording another podcast for the Graphistania podcast, and this time I've got another introduction of my dear friend Michael Hunger on the other side of this Skype call and that's Andrew Bowman. Hi, Andrew.
Friday, 17 March 2017
(Another) Podcast Interview with Alistair Jones, Neo Technology
Just before ending the week, I thought I would publish another great episode on our Graphistania podcast. Ever since the launch of Neo4j 3.1, I had been wanting to do an episode about the new Neo4j clustering architecture. It's so innovative, new and a great piece of engineering - we just had to sing its praise :) ... So who better to invite back to the podcast than Alistair Jones, who was one of the lead engineers at Neo Technology to pursue the effort. Here's our chat:
Here's the transcript of our conversation:
All the best
Rik
Here's the transcript of our conversation:
RVB: 00:02.563 Hello everyone. My name is Rik, Rik Van Bruggen from Neo Technology. And here I am again recording the second podcast of this year. I know it's only two months into the year, so I've been slacking, but I [laughter]--AJ: 00:15.346 You've been picking up the pace again.RVB: 00:16.082 Yeah, picking up the pace again. And for the second episode, I have invited a returning guest to our podcast, and that's my friend and colleague, Alistair Jones, from the Neo Technology engineering team. Hi, Alistair.AJ: 00:28.380 Hi, Rik.RVB: 00:29.089 Hey, thank you for making the time. I know you're a busy man these days, so thanks for taking the time. Alistair, the reason why I invited you back is because I know you've been hard at work in the engineering team, on some of the really big, new features in Neo4j. 3.1 was released at GraphConnect San Francisco last year. Or, no, it was actually announced and was released a little bit later, but one of the biggest new features in Neo4j 3.1 was the new clustering architecture, right?AJ: 01:03.036 Yep.RVB: 01:03.119 And that was what you and your team were working on?AJ: 01:05.054 Yeah, it was a really big thing for us, actually. So I've been working on this area for nearly two years, actually, on this new clustering architecture. And as you know, Neo4J is a clustered database designed to run over multiple servers. And we've had clustering in place for six or seven years in Neo. This is the biggest change we've ever made, by miles. It's a huge, huge upgrade of all the technology around the clustering.RVB: 01:40.760 Wow. I remember like in version 1.8 it was like Zookeeper that was doing some of the work.AJ: 01:46.629 Yeah, we had a small change in the 1.9 release back in the day.RVB: 01:54.312 Back in the day, yes.AJ: 01:55.544 This is a much bigger release in 3.1.RVB: 02:00.387 So what's it all about?AJ: 02:01.882 So the first part of it is getting up to date. So the world around us has moved on, and one of the great things about Neo is that we can take research from academia and actually apply it. So reasonably recent stuff that, if you read the academic papers and blog about it, we read all those, and some of those things we can put them fairly quickly into the product. So, for us, this time, it was doing the Raft protocol, which is a consensus algorithm. So what that means is getting agreement between participants, so computers in this case. We--
NOTE: Diego reacted to this part of the podcast with a super-cool tweet:RVB: 02:52.316 Members of the cluster, right?AJ: 02:53.396 Yeah. So, in this case, it can be different services in the class are getting consensus between those servers, when the servers themselves and the communication between servers is potentially unreliable. So you need to account for the unreliability in the design. Now, we know a little bit about consensus algorithms because previously, back in that 1.9 release, we implemented Paxos. And, at the time, that was the kind-of state-of-the-art thing to do. Raft, you could argue, at some theoretical level, is the same thing, but it's much more clearly structured. Raft is--RVB: 03:38.960 You mean a consensus protocol, right?AJ: 03:40.209 Yeah, yeah, exactly. So it's from Diego Ongaro, who's the lead researcher in this area, and it's really impressive how it's described.
@rvanbruggen @apcj nice. I found an error though: I'm not "the leading" anything.— Diego Ongaro (@ongardie) March 17, 2017
Subscribing to the podcast is easy: just add the rss feed or add us in iTunes! Hope you'll enjoy it!It's actually aimed to be simple to understand and to explain. And that makes it really good to implement because you can be very clear about what you've done. You can see the direction that you've gone in. So we've changed from one consensus algorithm to another.RVB: 04:13.411 Yep. Which is a big change [crosstalk].AJ: 04:14.871 Which is a big change, but architecturally it's totally different, because previously we were using Paxos to agree on membership of the cluster. So actually a very small amount of data. Not that many servers. They don't go that often. Now what we're doing is we're using Raft, and we're using it for every single transaction in the database. So every single node, relationship, property you create in the database it goes through the Raft protocol. You've got consensus across the cluster. And what that means is that every single change is agreed to by a majority of the cluster, so no matter what happens in terms of loss of connectivity or failure of the minority of the servers, still, the cluster as a whole agrees on what the state is as you move forward, so--RVB: 05:06.995 Sounds a bit like open heart surgery to me.AJ: 05:09.230 Yeah, it's quite a major change, but it's actually really nice. Once you've got that super solid foundation, you can build a whole load of things on top of it. So it's extremely solid for-- it's like the most reliable we could make it, and it stores every single transaction in this replicated log across all members of the cluster. And also as the membership changes, that's agreed to with protocol as well, so you know every time who the people were, who the servers were. People were allowed to [inaudible] transactions and to get them committed. So the whole thing's very tightly integrated into the core of the clustering.RVB: 05:52.520 So I would never claim that I understand everything about it, but what I've read is that it's very different architecturally in terms of-- previously we had masters and slaves, now we talk about cores and edges, right?AJ: 06:05.636 Yeah. The second part of this is that we were aiming to have much larger clusters than people had previously been running in Neo. Neo's been around for a long time. And, previously, people used to think of having 3, 5, 10 servers being a large database cluster. Now people want to run hundreds of servers, and we have customers and users running 200 servers in a database cluster. We want to be able to get higher than that, and the consensus algorithm that we were using before, the design of it, or perhaps the membership, yeah, it had a sort of limit on the-- or do we say kind of--? It was hard to get to that scale.AJ: 06:55.530 And the reason is that all of the servers had to be aware of each other and what they were doing at any stage to basically make sure that they hadn't disappeared. So that led to heartbeats going from every server to every other server, and that ultimately gets very expensive when you have a large number of servers. It also gets very difficult when you're committing across the majority of the servers because you have to wait for a large number of them to come back before you can say, "Yes, this is now safely committed."AJ: 07:30.229 So just having one huge cluster of Raft servers is not a good design for that kind of hundreds of servers or thousands of servers. So we came up with a new architecture. And what we do now is we divide the cluster into two groups. We mark some of the servers as being in what we call a call. Call servers participate in Raft and they are about safety. They're about storing your data durably. Secondly, we have a lot of potentially much larger group of read replicas. And these are servers that are for running your queries on, and--RVB: 08:14.895 Read queries, not write queries.AJ: 08:16.298 Yeah, yeah, read queries. So you don't have to worry about safety here, and the idea is these are about-- they're disposable, where you can scale them up and down; when your web traffic is high a certain time of day, have more and more of them.RVB: 08:30.725 Just have more of them, yep.AJ: 08:30.721 [inaudible] your cloud instances when it's quieter, and you can adapt to the shape of your traffic with the read replicas. What's interesting is that the name read is that we're doing more service than reading. Why does that make sense in a--? How does that help you in a database, have more read only things? Surely you need them more so to write. Well, that's because of the shape of graph data. It's because, actually, when we look at the-- I'll show you, because it's a nice slide [laughter] with audio only. You're looking at a slide that shows kind of how we see people do stuff with graphs, and what you notice is that the right [inaudible] updates tend to be quite small.RVB: 09:17.172 Local [crosstalk]?AJ: 09:17.605 Yeah, very, very local. Like, two or three nodes in relationships, up to maybe 100 things in a transaction, whereas on the read side - the whole point of graphs is to really fast, and people go a long way - they traverse along the graph in a read transaction. So they're doing hundreds of thousands of relationships in one transaction. Now, that's very fast, but it still takes resources. It takes memory bandwidths, it takes CPU to run these queries. And that's what people are really hammering their graph with, thousands of these, each very big, queries. And that's an enormous amount of computational load. We want to spread that across a lot of servers, and this is a way to do it - have loads of re-replicas that can handle that traffic for you. So it is really helping you in the kind of [inaudible] applications. It's a very specific architecture to the type of system that we're building.RVB: 10:13.028 Pretty cool. And so, as I understand it, the core is-- so they're all about the safety, and about writing to the graphs, and the age servers are all about reading. Is there any downside to this? Is this good news show all the way around, or are there some things that we should take care with?AJ: 10:34.656 So there's one thing that's just like-- a challenge here for people when they're deploying these type of applications, is that the transaction's being pushed out from the core, out to the B replicates, and there's some delay in that happening. It's very small, but there is some delay. So people call this eventual consistency, and this is something that we're aware of. And lots of modern sort of web systems that you get into this kind of eventual consistency situation. An example of this that could kind of catch you out is, say you're a user, you create an account, or you make a booking, that's a right transaction. It updates the graph. Then when you come to refresh your page, you try another operation and it's a read only operation, maybe you hit a read replica that hasn't quite seen your update, so, as a user, it almost appears like the thing you just did has disappeared, like you've gone back in time. There's a bit of a--RVB: 11:45.015 It's [crosstalk] read your own writes problem.AJ: 11:46.385 Yeah, I can't really-- so what we did at the same time as this, is we actually added a whole new feature that became the name of the whole clustering architecture. So this is what I like to call causal clustering, because we added in a feature of causal consistency.RVB: 12:08.446 Tell me more about that, because I don't know what that means [laughter].AJ: 12:10.836 Okay, Rik. So causal consistency. So it's actually something that's been-- again, from research, there's some academic and industry research in this area, but it's not very commonly implemented. There are only a handful of other implementations out there, and what it's about is trying to represent what causally has happened in the user's application. So the cause and effects of the changes that you've made.AJ: 12:45.088 Practically, it's very easy to use. What happens is that when you update the graph or when you touch the graph in any way, the database can give you a bookmark. And this bookmark represents the latest thing that you've changed or the latest thing that you've seen in the database. And then when you make another request to any other server in the cluster, you can supply that bookmark that's saying bookmark, and the database will make sure that it has at least as up-to-date a state as the bookmark represents. So the bookmark is just a little string and it comes back to your database driver into your application code. You can store it in your application server, or you can hold onto it temporarily while you make another inquiry, or you can send it all the way back to the client. You can send it back to your web browser or your mobile device, and route it back, ultimately, to the database.RVB: 13:46.373 So that basically assures that the client of the database always takes into consideration everything that it calls [crosstalk]?AJ: 13:54.484 Yeah, it prevents you from going back in time--RVB: 13:56.751 Ah, yeah, that's it.AJ: 13:56.890 --is what it does. And it supports a totally stateless architecture - everything between the user and the database. The database is storing state. Why should you need to store it anywhere else? So this is [inaudible]. Your sessions, you don't need to worry about sophisticated routing. Just have stateless application servers, pass your bookmark around, and you get causal consistency. That's the idea.RVB: 14:28.017 Wow.AJ: 14:28.699 And we've tried to make this even easier to use by building some of the primitives. The kind of passing backwards and forwards keeping track of things is built into the database drivers. So in 3.0, we introduced--RVB: 14:42.424 BOLT drivers, right?AJ: 14:42.490 Yeah, the BOLT drivers. So they initially supported native language drivers in your [crosstalk]--RVB: 14:48.561 Right. And so the new version of the driver supports this bookmarking--AJ: 14:51.515 Exactly, yeah.RVB: 14:52.721 --and that gives us the causal consistency.AJ: 14:54.338 The causal consistency, yeah. Exactly.RVB: 14:56.476 So let's talk a little bit about the future. What's coming up? What are you working on now, and what keeps you up at night, and [laughter]---?AJ: 15:03.404 Yeah. Well, [crosstalk]. I mean, it's kind of following on logically from where we are now, so the next stage of this is to be-- it's that kind of how people actually deploy this stuff. And these days, not just a cluster of servers that are using it to run a database. It's also servers across multiple data centres and multiple regions around the world. Around the country, all around the world. So that's what the cloud environment's been very easy to do, to have geographic distribution. And we are taking account of that feature in the product, or that server usage in the product. So what we're going to do is make the clustering aware of data centres and how they're organised, and allow the client to give hints about how might be the best way to serve it. So that means that you can do your reads from a server that's very close to you, with a low latency, and you can support fault tolerance across data centres when one of them goes away, or explicitly recover in a disaster recovery zone. All of these different operational scenarios. So--RVB: 16:24.512 Is that something that's coming up in the next couple of versions of Neo4j or--?AJ: 16:27.024 Yeah, yeah. So in the next couple of versions, that's the stuff that's going on. And, again, it's to be seamless all the way through the driver, so you write your application once for Neo4j on your laptop, and then it should move forward [inaudible].RVB: 16:46.117 That's very cool. I have one more question. Don't you miss the visualisation stuff that you were doing before [laughter]?AJ: 16:52.896 Yeah. So I always miss the visualisation. I try to devote my spare time to get back into it every now and then, so--RVB: 17:03.147 Very cool. Well, thank you so much for spending your time, Alistair. I mean, we want to keep these podcasts fairly short, but I'm sure we'll include a bunch of links to the documentation and the blog post that we wrote about this topic. I really appreciate you making the time, and look forward to seeing what's up next.AJ: 17:21.907 Thanks very much.RVB: 17:23.060 Thank you. Bye.
All the best
Rik
Friday, 23 December 2016
Podcast Interview with Emil Eifrem, Neo Technology
In the summer of 2015, 5-6 months after first starting this crazy podcast thing with Michael and Mark at Qcon London, I finally got my boss and friend Emil Eifrem, CEO of Neo Technology, to spend some time with me on this podcast. It was a great conversation, and I still smile thinking about the silly drumroll that we used. But just before we wrap up 2016, it felt like it was the right thing to get Emil back on the podcast, and talk about "stuff". Here's that conversation - a little longer than usual, but totally worth it.
Here's the transcript of our conversation:
Here's the transcript of our conversation:
RVB: 00:02.909 Hello everyone. My name is Rik, Rik Van Bruggen from Neo Technology. And here I am again. And I'm so excited, I can barely restrain myself. It's my “über boss” on the phone again. It's been 18 months since the last interview, and here I have him back on the podcast. Emil Eifrem. Hi, Emil.
EE: 00:21.803 Hi Rik. Thanks for finally inviting me back.
Friday, 25 November 2016
Podcast Interview with Craig Taverner, Neo Technology
The interview below was long overdue - but very much worth the wait. For the past couple of years, the Neo4j community has been brewing on a really interesting add-on capability to integrate GIS-style, spatial querying capabilities into Neo4j. It's such a great and natural fit - and one of the driving forces behind this in the community has always been this global citizen called Craig Taverner. Craig has been in the ecosystem for years - first as a community member, then as a commercial customer, and now as an employee in Neo's Swedish engineering team. So about time we had a chat:
Here's the transcript of our conversation:
Here's the transcript of our conversation:
RVB: 00:02.785 Hello everyone. My name is Rik, Rik Van Bruggen from Neo Technology, and here we are again, recording another Neo4j Graphistania podcast session. And today I'm joined by one of my colleagues actually, in the Neo4j engineering team, Craig Taverner. Hi Craig.
Tuesday, 6 October 2015
Podcast Interview with Jesus Barrasa, Neo Technology
Over the past couple of months, we have been doing a lot of work at Neo4j to try to better explain the value of Neo4j to our prospects and customers. This has been a true team effort, with lots of engineers, marketeers and sales folks participating in articulating how complex technology can be used to add true value to business processes. And as we did that, we added some really talented people to our team. One of them is the person that I am interviewing in this particular podcast episode: Jesus Barrasa. Jesus has a lot of experience with graphs and even (a particular kind of) graph databases - so going in I knew it was going to be an interesting chat. And guess what: it was. Listen on:
Here's the transcript of our conversation:
All the best
Rik
Here's the transcript of our conversation:
RVB: 00:00 Hello everyone. My name is Rik, Rik Van Bruggen from Neo and here I am again recording a Neo4j graph database podcast and my guest today is all the way in the UK. Jesus, hi Jesus.
JB: 00:13 Hi Rik. How are you?
RVB: 00:14 Hey. I'm always scared of pronouncing your name in the wrong way. I'm sorry.
JB: 00:19 You did great. You did great [laughter]. I've heard much worse than that so that was brilliant.
RVB: 00:24 Okay [chuckles]. Okay, Jesus, you just joined Neo a couple of months ago as a pre-sales engineer but you have a long-standing history with graphs. You did a PhD on the subject if I'm not mistaken, right?
JB: 00:37 That's correct, yeah. It all started probably more than ten, nearly 15 years ago, so quite a while yeah.
RVB: 00:43 Yeah.
JB: 00:43 And you're right. It all started in the semantic technology space, in the RDF space, so that was the first time I was exposed to modeling data as graphs and yeah. That's been a long story. After the PhD I did work for a company in London called Ontology where we did use graphs to best resolve problems in companies, mostly in the telecommunications sector and yeah. As you say, two months ago I joined the field engineering team in the London office.
RVB: 01:18 That's super. What was your PhD about exactly then?
JB: 01:21 Well, at the time… I did model-- I mean, I formalized a way of translating relational schemas into ontologies. Ontologies is the way you represent metadata in the RDF model. I don't know if we will have time to talk about that but yeah, it's actually an automated mapping between relational schemas and ontologies. That's what I--
RVB: 01:45 There's a lot of people looking at that I think, still today, I get that question quite often.
JB: 01:50 Absolutely.
RVB: 01:51 So another question because you've been working so long in the RDF world, what's the difference between the RDF semantic technology space and the property graph model of Neo4j - what's the key difference for you?
JB: 02:04 All right. I'd say you can answer this question from two perspectives because obviously RDF is a presentation paradigm and there's the different implementations, the triple stores that you can find in the market but as a model I think the main thing they have in common is the fact that they both represent data as a graph. And that makes them very, very close to each other. The difference I would say is RDF is simple as it can be. It's only based on the notion of URIs to identify resources or nodes if you want to establish the properties with a property graph but there's this element called triple, subject, predicate, object. And that's all the constructs that you have to model your domain. And of course I think that's the biggest difference because in the property graph model, you have nodes with properties, you have relationships with the properties and they have this brilliant thing that's this white-board friendliness, this excellent thing that's the way you conceive it, the way you model it in your head, in your whiteboard, is exactly the way it's represented and stored physically, whereas in RDF there's still this gap where you have to translate that into triples, things that may sound simple like giving a weight to a relationship so there's a connection between Rik and Jesus because we work together. You want to give weight to this relationship. That's something that's completely natural in the property graph but it's not something you can do directly on a triple, so you have to model this relationship probably as an intermediate resource. There's a bit of a gap and I think that makes it sometimes less intuitive, sometimes a bit less humane if you want. So that's the difference.
RVB: 04:06 Well [crosstalk] actually that was sort of what I was hoping you would say because I've always been told that the difference is really on the predicate, the fact that it's so difficult to qualify a predicate in the RDF model.
JB: 04:20 That's [crosstalk] exactly. Then there are other interesting things in RDF, the whole idea of being able to use the model itself to represent the metadata, the ontology and that gives you in certain cases some interesting powerful things you can do like querying all the data and the metadata at the same time. But I would say the biggest thing is more what they have in common and it's the fact that they look at data as a graph, a set of connected nodes.
RVB: 04:49 What attracted you to the graph in the first place? Why did you get into this field in the first place if you don't mind me asking?
JB: 04:55 Well, yeah, sure. I think there's two things. One was this incredible flexibility. At the time when I started we didn't talk about NOSQL. That's a concept that was coined later. It's the idea of being able to start storing data without having to model it up front. I thought that was brilliant and that gives you an incredible flexibility, this thing of schemaless model or at least implicit schema depends on how you talk about it but this flexibility was one of the things and the other one, I'd say the possibility of infering new knowledge based on the information on your graph, you identify patterns and you can enrich your graph with new knowledge based on the data elements that you have. So I think probably these two things attracted me to this world and I find them unbelievably useful and powerful.
RVB: 05:51 You know, as your couple months at Neo, what do you think is the most interesting use-cases that you've seen so far? Anything that jumps out?
JB: 06:01 Sure. Well, I'm amazed-- the thing is the great thing about Neo is how you can use and update your graph in real time, at what speed you can not only read it and query itbut also keep it up to date and how it's possible to identify fraud rings for example is one of the cases that I like the most, being able to pick up the status of your accounts, your users, their information, the transactions that they're carrying out, and at the same time be able to pick up, detect the patterns that identify a fraud ring is one of the ones that I enjoy the most.
RVB: 06:43 And do that in real time you mean?
JB: 06:44 Exactly, the real time is the key thing and that's pretty impressive yeah.
RVB: 06:48 It's not like with my experience a couple of weeks ago with my Amex card, I get a fraudulent transaction on Friday and I get a call on Monday [laughter].
JB: 06:57 Oh yeah, just in time right.
RVB: 06:59 Just in time [laughter]. Didn't really work. Very cool. So where do you see this going, Jesus, where's the industry taking us do you think?
JB: 07:09 Well, I think adoption is growing. It's amazing the number of different organizations and companies that we talk to. I can't think of a single vertical, a single sector that would not benefit from scenarios where modeling data as a graph adds incredible value, so I think adoption will definitely grow and that's one of the things, and then the other one that I think is going to be key as well is about integrating the graph with the rest of the data architecture. These days there are so many alternatives to represent data and some of them adequate for certain scenarios. I'm all for peaceful co-existence with all the other approaches. So integration I think is going to be the other important one and I think being able to expose the graphing in ways that make it easy to inter-operate with other stores, sometimes not all of the information is going to be in your graph but the extremely valuable information in your graph will need to be combined with some external information. That's one case where you will want to visualize it in different ways using BI tools, well you name it, there's so many elements now in data architectures that integration I think is going to be the other important aspect that we'll see developing in the next few years.
RVB: 08:22 There was one thing that I wanted to ask you and I obviously forgot. I'm so good at this podcasting thing [chuckles] is actually you've done a lot of work on virtualization of data, right and then [crosstalk] integrate and that links to that integration story I suppose.
JB: 08:36 Exactly. Exactly. I did work in the data integration space with a data virtualization company in the couple of years between Ontology and Neo and yes, I'm particularly interested in that and it's a way of integration data virtualization that's based on this idea of wrapping the sources and make them look as if they were relational even though they're not so they're not copying the data into centralized stores. You leave the data where it is and you define the logic to extract it and combine it and I think that was a powerful paradigm for new ways of representing data like the graph and make it easy to integrate them with other technologies and other types of stores and yes.
RVB: 09:27 So in a case like that the Neo4j would be one of the sources of virtualization? Is that what I'm--?
JB: 09:32 That's correct [crosstalk]. One essential one, that's the thing because the importance in the end is what value is there in your source? Neo can be perfectly, for example, in an MDM scenario. It can be the core. It can be your master model. And you can link it with the different rest of it and provided the detailed information about the entities but exactly, it would be one of the sources and the data virtualization will expose it in a way that's easily consumable by say for example BI tools in analytic scenarios or that's one of the--
RVB: 10:05 Are there any examples of that yet? You know, like open-source virtualization tools that integrate with Neo [crosstalk]?
JB: 10:10 Well, there's not much to be honest. There's one quite limited community edition of one of the vendors called Denodo which is the one I used to work for. There's another one from JBoss but I'd say there's not much, Rik, available in there. I mean, JBoss would be the obvious option. I'm actually now trying to work a little bit with it and try to build some integrations with Neo and yeah. That's what you can find.
RVB: 10:40 I think this is kind of like community call for help, you know [crosstalk].
JB: 10:45 Yeah [crosstalk]. I definitely really hope to be publishing something soon, at least in some idea, some small examples that can inspire people to look at these.
RVB: 10:55 That would be great. Cool. I think we're going to wrap up here. I think we like to keep these podcastsquite short and snappy but thanks a lot for sharing your perspective. I think that was very interesting although because of my limited presentation skills a little bit chaotic [laughter].
JB: 11:13 Right. No, it was great [chuckles]. Great to have this chat with you Rik.
RVB: 11:15 Thank you, Jesus. And I'll see you soon yeah.
JB: 11:18 Lovely. Cheers.
RVB: 11:19 Bye.
JB: 11:20 Bye now.Subscribing to the podcast is easy: just add the rss feed or add us in iTunes! Hope you'll enjoy it!
All the best
Rik
Friday, 5 June 2015
Podcast Interview with Tobias Lindaaker, Neo Technology
Next month, I will be celebrating my 3-year anniversary working for one of the best companies I have ever worked for: Neo Technology. It seems like I have been here a lot longer - so much has happened in those three years! I guess time truly flies when you are having fun :) ... but it makes me wonder sometimes - what was it like in the early early beginnings. What is it like for people that have been here a LOT longer to look back on the Neo4j journey?
So I decided to ask. Here's a lovely conversation with Tobias Lindaaker, employee number 1 of Neo Technology. Overall nice guy and top engineer - he's got a lot of ideas and perspectives on that journey:
Here's the transcript of our conversation:
All the best
Rik
So I decided to ask. Here's a lovely conversation with Tobias Lindaaker, employee number 1 of Neo Technology. Overall nice guy and top engineer - he's got a lot of ideas and perspectives on that journey:
Here's the transcript of our conversation:
Subscribing to the podcast is easy: just add the rss feed or add us in iTunes! Hope you'll enjoy it!RVB: Hello everyone, this is Rik - Rik van Bruggen - from Neo Technology, and here we are again, recording an episode for our Neo4j Graph Database podcast. It's a wonderful Wednesday afternoon, and together with me on the call here - Skype call - is Tobias from Neo in Malmo. Hi Tobias.TL: Hi Rik?RVB: Hi, welcome on the podcast.TL: Thank you, thank you.RVB: Tobias, I invited you to the podcast for a couple of reasons, but maybe you could start by introducing yourself, so people know who you are.TL: Yeah, so I am Tobias Lindaaker. I work as a senior developer at Neo technology, working on pretty much all things in the development of Neo4j. I've had my hand in almost all feature stuff we've released.RVB: That's pretty amazing, and that's probably a good lead into why I sort of wanted to talk to you, because you're probably one of the first engineers that Neo hired, isn't it?TL: I was the first engineer that the company actually hired.RVB: When was that?TL: That was in 2007. So September 2007, I got on the company to help our very first customers with using Neo4j.RVB: Wow, and how did you get into it? Did you know some of the founders then?TL: Yeah, me and Emil went to college together. I taught him calculus, and he thought that I was good enough with that, that he might as well give me a shot at doing this job for him.RVB: [chuckles] Any gory details that you can share with us [chuckle]? I'll always appreciate it.TL: On Emil--RVB: Yeah, of course.TL: —from the early days?RVB: No, don't go there. Don't go there.TL: So, I did not only teach him calculus, he taught me a great deal about programming, because he actually wrote code back in the days, in particular about testing, because he was really big on testing. He spent more of his time writing tests for his code than writing the actual code. He was always last handing in his lab assignments, because he spent so much time developing his tests, and his testing framework, and all of that stuff. He didn't actually get the job done, he just played around it.RVB: Do you still still remember what was the first feature or project that you were working on at the time?TL: With Neo4j?RVB: Yeah.TL: The first project was a customer project. I was working on a geospatial system. We were doing road user charging, so essentially road tolls based on GSM-based positioning. We had a algorithm or an idea for an algorithm for how to use GSM triangulation, by knowing the position of the cell towers. We had two modes for this system. One where we would collect data and store it out in Neo4j, to match what the signal profile from different cell towers looked like when driving on particular roads. We would find the profiles for roads, and then we'd use that when a user was driving in the same area, to match the signal profile from the cell towers, with the road that he was driving on. So, that was the first part of what I did for Neo4j.RVB: Super cool. So, how's it been? What's it like to work for a start-up that's been going through all this evolution in the past eight, nine years?TL: The main thing for me, is the fact that it's the start-up or it has been a start-up. Finally this year, I'm starting to see signs that we're growing up and become a mature company, but there's been lots of ups and downs getting there. There was a failure to get investments of 2008, where the company nearly went bankrupt, and we were almost out of jobs - all of us. Then of course, there's been frustrations with things that happen when you on board more people, and you get less and less influence over the company because the company grows and becomes bigger. Pretty much when we started, everyone was on the same level, because it was just a bunch of guys working on the same thing together. Then, as we've grown, I've stayed on the “working on things” level, and the people who were with me at the time have gone on to be CEO and CTO and such things, so there has been a transition, the relative influence at least in formal terms has diverged overtime. It's interesting that I get less and less time with my old friends at the company.RVB: Yeah, but I suppose there's new friends coming on board right? There's new things happening--TL: Absolutely.RVB: —and there's new exciting stuff happening with the products as well.TL: Absolutely, and in terms of the product, it's the main driving feature why I love this company. I think the product really has a lot going for it.RVB: I couldn't agree more. What do you think the future holds Tobias, both for you from personal perspective, and for the company and your job there? What do you think is in store?TL: As I said, the company is growing up now, and I think we will start seeing effects of that in the product pretty soon. In that, since the team is growing, we've got more developers now than we've ever had before. Last year we started hiring developers for real, started actually growing a development team, and this year those developers are up to speed with what we're doing, so we can start continue hiring even more. What we're starting to see now, is the ability to work on a lot more features at the same time, and even start taking risks in what features we're developing, and invest in sort of high risk-high reward type of projects, where we aren't really sure if they will pay off, but we've got enough people that we can spend a small team actually trying it out. That's really exciting both in terms of getting to work on those things, but also in terms of the potential that those features can deliver.RVB: Maybe I can finish off with one question. What is your favorite feature of Neo4j?TL: My favorite feature?RVB: Surprised you there, sorry [chuckles].TL: The … so there’s… I can either go very general and say that I really like the model of the Graph Database, because that's always been the main driving thing behind why I wanted to work on Neo4j, because I really like that model, but I'm not sure that really qualifies as a feature.RVB: No, it isn't, right.TL: So, in terms of something smaller, I'd say the Query language. Is that good enough to qualify as a feature? It's a recent thing in terms of if you compare to how long I've been with the company. The Query language hasn't been with product for as long as I have, but it's really nice to see the expressivity and usefulness of it.RVB: Yeah, I know. I couldn't agree more. Cool. Tobias, thank you so much for coming on the podcast. You know that we want to keep these things nice and short, so I appreciate it. It's been really cool having you here, and I'm going to be talking to a lot of other colleagues, and people that are working on Neo a lot less long than you have. So, it'll be interesting to compare notes in the next couple of episodes.TL: Yeah, I'm looking forward to hearing that.RVB: Absolutely. Thank you so much Tobias , and I"ll talk to you soon.TL: Thank you.RVB: Bye.
All the best
Rik
Sunday, 22 March 2015
Starting the week with a podcast interview: Dr. Jim Webber
The workweek is almost there, so what better time to publish another interview for our Neo4j Graph Database podcast. And this one is a bit special, since the interview is with one of those people that is probably as close as you can get to the forefront of graph database technology: the one and only, Dr. Jim Webber.
Here's the transcription of the interview:
Before I leave you, just two more things:
All the best - have a great week.
Here's the transcription of the interview:
RVB: Hello, everyone. My name is Rik and today I'm in sunny Amsterdam recording another podcast for our Graph Database podcast. We have a new guest which is the ever charming Mr. Jim Webber. Hey, Jim.
JW: Hello, Rick, how you doing?
RVB: Doing very well. The sun is shining and it's a beautiful day outside. People might not know you that well so why don't you introduce yourself? What's your link to the wonderful world of graph databases?
JW: My name's Jim Webber and I'm Neo4J's chief scientist. My link to graph databases goes back to about 2008 when I first started to use Neo4J and ultimately contribute to the database before I joined the Neo team about four-and-a-half years ago.
RVB: That's great. So in this podcast, I really don't have a lot of questions. The two really most important questions are what do you love most about Neo or graph databases and where do you see this going, so let's start with the start. Can you tell us a little bit about what you really love about graph databases and why it's the best thing since sliced bread?
JW: It may even be better than sliced bead. There's two answers, really, Rick. At the moment, I work on the inside of Neo4j and I've got to say that it's a joy and a privilege to do some really fascinating computer science research and development work that people then take and build amazing systems on. But that amazing systems aspect is really what got me hooked in the first place as an Neo4j user.
JW: In what feels like an eternity ago now, I was faced with a challenge of building a product catalog in a telecoms company and modeling in this product catalog things that the business users wanted, particularly up-sell - knowing what products you already had, being able to price those products as a bundle, and importantly being able to cross-sell and up-sell to you to increase your value as a customer. Back in the day, we were going to do that with a mixture of off the shelf software and customization relational databases. We thought that give or take, it would take about three years before we had up-sell functionality implemented.
JW: I bumped into Emil Eifrem, one of the founders of Neo4J, quite accidentally in the bleak depths of a sudden Swedish winter and he explained to me that really, my model was a graph. I kind of got that at a conceptual level. I knew that things depended on other things and so on. But then Emil told me about this thing called Neo4J where the graphs were represented as first class citizens in the database. I've got to say that didn't sit comfortably with me. The database for us back then, that was all about the relational database. While I understood in my Java code or whatever I was going to have a bunch of objects connected together in my database, I expected tables.
JW: Anyway, I went back to work and thought, "What was that funny named database that that Swedish guy told me about?" We gave it a go and actually within an afternoon - admittedly a long afternoon - we spiked out what it would mean to do a product catalog for telcoms and implement up-sell, and that really blew away. The first time I ran a query in Neo - and it wasn't Neo as we know it today, it was an early version of Neo - didn't have the wonderful cipher query language or any of that stuff, just really did a simple graft traversal. But when it told me given my starting product what I should buy next, blew me away. I honestly thought I'd built Skynet and then you think, "I haven't actually, I've just done a graph traversal." But from that moment on, I was hooked because we took a problem that we thought might take three years to deliver and we delivered it in order of magnitude hours. Being able to conceptualize a problem as graphs makes things that were previously intractable readily tractable and I was hooked from there on in.
RVB: That's super. So in summarizing, it's all about the model? Is that-- it's the model that most attracts you to it?
JW: So the graph model, I think, is the most expressive and pleasant model that I've ever worked with, because it matches the way that I think many of us think as humans about stuff being connected to other stuff - kind of rich, semantically-driven network. What attracts me to Neo is that it was the technology that supported that model that was leagues ahead of everything I've ever seen before, and even today is still leagues ahead because it's by far the most mature graph database available. So If I want to adopt the most expressive and straightforward and pleasurable model, and indeed in many cases the most performing model, that's graphs. If I want the technology that supports that, that's Neo.
RVB: Super. Preaching to the choir, but I couldn't agree more. The follow-up question is where is it going? As the chief scientist, you're perfectly placed to answer that question. Where would you see the technology in five years from now and what's a realistic objective there?
JW: I guess there are two levels to that answer. There's the business impact that the technology's going to have, then there's the technology itself, and I think they're both fascinating. Right now my sense is that graphs are primed for the big time. You look at all the indicators that we have, all of the metrics that the analysts are running, and graphs are by far outstripping all other database categories, even those equal in terms of growth right now. There's something that's built up and pent up and now graphs are a thing. Not just Neo4j, although Neo4j is leading the charge, but graph tech as a whole is really taking off. I think that at some point in the medium term we're going to look at graphs in the same way we look at relational today - it's just going to be the data model. I'm super confident about that.
JW: Where are we going in terms of Neo4j and its implementation technology? Well, where are we not going? There's years of computer science ahead of us, some of which is already written down. The academic community has been doing some very pioneering work there. But to boil it down to a few podcastable sound bites, I think there's a bunch of work that's going to happen around query languages. I think that cipher query language is going to go from strength to strength. I think those guys are going to figure out better and better ways of doing query planning and optimization, potentially even things like your parallel queries and distributed queries and so on.
JW: In terms of the database engine itself, it's a whole bunch of fundamental concurrent programming, algorithms, and data structure stuff where we're going to be pushing the limits in terms of performance and robustness. Indeed in terms of the distribution system stuff, which is in my background, I think we've seen a resurgence of incredibly high performance transaction protocols which all but detonate the reasons for having soft state and eventual consistency because they eliminate so many of the unavailability and un-performant characteristics of traditional 2PC. So I think we're going to see a resurgence of extremely high performant, high concurrency transaction processing and commit protocols. I'm definitely looking forward to living in that world in the next few years.
RVB: That is a super answer and I'm really looking forward to living that together with you. It's going to be an exciting ride. Thank you so much for the time, Jim, and I look forward to speak to you again.
JW: Pleasure. Thanks, Rick.
Subscribing to the podcast is easy: just add the rss feed or add us in iTunes! Hope you'll enjoy it!
Before I leave you, just two more things:
- after this interview, Jim and I spent another half hour laughing ourselves into a dent because the Doctor had misunderstood one of my closing comments (underlined and bold above): he thought - being in Amsterdam and all that - that I said "I look forward to living together with you" - instead of what I really said. Wishful on his part, probably, but it did give us a lot of giggles :) ...
- there's a reason the interview has the song below, performed by Wilco and Billie Bragg. Woody Guthrie's commentary on the "big depression" bankers seems ever so actual today. And I think Jim would like this cover.
Rik
Subscribe to:
Posts (Atom)
