You've no doubt heard that 95-percent of A-I projects fail. It’s a stat that gets thrown around as proof the technology is overhyped. Not everyone sees it that way, including our next guest. Philip Rathle is the chief technology officer of Neo4j — a graph database company that's become, in his words, an accidental piece of infrastructure for how serious AI systems actually work. And he says that 95% failure rate isn't a red flag. It's the checks and balances doing their job. I’m Jennifer Strong, and this is SHIFT.  [THEME] This episode - it’s the latest installment of our oral history project. Philip Rathle: [00:00:00] Good to be with you, Jennifer. I'm, uh, Philip Rathle, CTO at Neo4j. I've been working in data, databases, AI in some form or another, analytics for 30-plus years, and for the last decade plus, had the amazing journey of creating a new category in data and databases, accidentally building a product that's serving as the left brain to the LM right brain in AI systems Philip Rathle: There's been all this discourse about AGI, and AGI in a sense makes a pretty big assumption which turns out to be wrong in many contexts. And that assumption is that an LLM is a monolithic system that is going to develop all these capabilities. The way AI systems are coming into being is not as single monolithic systems. Philip Rathle: In fact, the MIT figure of 95% of AI projects are failing [00:01:00] is a sign that the checks and balances inside of companies are alive and well with respect to just not letting things ungated into the world that are then going to tank their reputation and their revenue and have health and human safety and ethics and regulatory and other implications. Philip Rathle: So the systems that are successfully making their way into production, i.e., the 5%, and the ones where the stakes are genuinely high, are ones where you don't have just an LLM that's doing all the AI. You have a composite of different systems. Now, this can be big models, small models, other models working together, but importantly, it's other systems like databases, knowledge layers, frameworks, tools that can make up for some of the shortcomings that LLMs have[00:02:00]  Philip Rathle: For many problems in the world, in the business, in society, in government, it's not okay to even have a 1% failure rate. Furthermore, they are a black box. We don't know why they're making decisions. I recently listened to a podcast from someone from Anthropic who explained that even inside of reasoning models, so-called reasoning models that explain their steps, those steps are just as prone to hallucinate as the actual answers. Philip Rathle: So it's not really explaining. It's, again, doing a probabilistic job of telling you what its reasoning is, [00:03:00] and it may or may not be accurate. They also lack discernment. What do I mean by discernment? Any data that they've been trained on is fair game to use for any situation, any person, and as we know, there's a lot of data inside of enterprises that is regulated by privacy laws, by ethical concerns, et cetera, where you just can't share anything with anyone at any time. Philip Rathle: And then last but not least, models only know about what's happened up until the time they were trained. So this is why in the enterprise and in government, you know, these AI systems of any sort of significance that go live not just as an LLM monolith, but as a composite. And, you know, there are some pretty big implications here. Philip Rathle: It sounds like a simple and obvious thing, but then what is AGI? Uh, you know, you're talking about artificial general intelligence with respect to some set of components, which, by the way, are different [00:04:00] from month to month as a system evolves and are different across different use cases, even in a single business, let alone across businesses. Philip Rathle: So that one then falls apart and almost presumes that once machines reach this point where they're making better decisions than human, we are in some way morally and ethically obligated to do what the machine says, which is super dangerous. Because ultimately, I mean, one of the few human freedoms we want and must retain control over is where do we apply our sense of agency and our volition, and, you know, under no circumstance do we have an obligation to turn it over, let alone to something that's non-sentient. Philip Rathle: And I, I, I will assert that intelligence does not automatically imply sentience, consciousness, life, any of that, really. It's, it's, it's a, it's a machine carrying things out. I guess that's one One, one, one [00:05:00] rant. And, and another one I'll add is the notion that AI systems must only be probabilistic. And, you know, I think there are a lot of assumptions that are founded in this, but, um, before there was probabilistic AI, there was symbolic AI, which is rules-based. Philip Rathle: And, you know, it's pretty popular these days for folks in contemporary AI to, to poo-poo rules-based AI. But I'll, I'll, I'll tell you in my day-to-day work, one of the ways that one might use the technology I work in, Neo4j, which is a... often used as an AI knowledge layer, is often used for agentic memory Inside of agentic memory, you have a need for short-term memory, which is like working memory. Philip Rathle: There are various forms of long-term memory, episodic, which remembers the different episodes that you experience. There's temporal memory. There's, let's say, ontological memory. What are the different things that are in my [00:06:00] world and remembering what those things are. And then there's process mem-memory, which is a series of steps. Philip Rathle: And there are quite a few companies that are using probabilistic AI to reverse engineer documents and emails and so on inside of an org to come up with what are the defined set of rules that we need to follow, because guess what? In this particular case, there are a defined set of rules, which by the way, might be regulated and might rep-represent best practices. Philip Rathle: So the kind of memory which I wouldn't have thought would be popular in this day and age of everything probabilistic is actually process memory, which is a deterministic set of steps. Um, likewise, it turns out that there are a lot of questions that have an exact answer where it's important to have an exact answer. Philip Rathle: And I'll, I'll give you an example of a client of ours. Um, Walmart has an AI system available to their 1.6 million employees that lets any employee look at any role in the company and understand what the different viable career paths [00:07:00] are to get there. And how do you get there? Well, if I have a knowledge base that understands all the different paths and all the different prerequisites and someone's career journey up until now, then you're basically playing a game of connect the dots. Philip Rathle: So there, there are a lot of these AI problems that can be rationalized by simply playing a game of connect the dots, whether it's from an ailment to a candidate drug through the network that is the human, uh, all the different human processes do-down to a cellular and molecular level, down to money laundering, down to an ultimate beneficial owner, all kinds of finance and regulatory use cases, which are forms of symbolic AI and which brings me to not only is AI in the enterprise already moved to something that's very much a composite system, but where that composite system has deterministic [00:08:00] as well as non-deterministic parts and is therefore neurosymbolic, which is th-this idea that, uh, dates back to the '90s where you could combine this original form of AI, which was following a set of concrete rules to this probabilistic form of AI, which is, you know, really amazed all of us, including me, and bringing the two together lets you do more, go farther, but, you know, with-without all the downsides. Philip Rathle: Right. So if, if, if I look across enterprise and government ac-activities, if I'm doing any sort of machinery maintenance, there's a zero tolerance for error. Even with a zero tolerance for error, there are sometimes mistakes, and when, when they are, they're absolutely terrible. So you need to do [00:09:00] everything you can to minimize those. Philip Rathle: The probability tolerances of an LLM are nowhere near what it takes for machine maintenance. But, but even going back to the Walmart example, someone's career path. Another example, Uber's a customer and all their configuration inside of their different cities of, you know, what, what are the requirements of having for a driver in terms of what vehicle they need for what level of service and what levels of service are available. Philip Rathle: They can't just go out and provision, you know, misprovision people and then come back later and say, you know, say, "I'm, I'm sorry we said that you could drive this speed or have this particular service, but we don't even offer service in, in this city." Or, you know, the, the LLM hallucinated it. So these are kinds of questions that are also, you know, higher stakes in the sense that multiplied by the number of people who are getting those answers, the tolerance is, is, is fairly low. Philip Rathle: Medical, that's another [00:10:00] case where you need some knowledge layer with information about that particular patient and their history, combine it with the LLM. And another one is fraud detection, money laundering, these kinds of things. And yes, there are false positives, but there's a huge cost to both false positives and false negatives. Philip Rathle: And moreover, there are certain kinds of questions that LLMs just still aren't good at answering. So they've gotten a little bit better at like math, but it's gonna be a long time, if ever, before they can do multi-level, let's say like 20 level deep in some supply chain to do a supply chain calculation or asset ownership or, you know, back to the HR hierarchy example where it's many, many, many levels. Philip Rathle: So one, one of the many kinds of areas where LLMs just aren't good at processing is anything requiring explicit context or connectivity beyond maybe one or two levels of causality. And [00:11:00] this is another area where you want to couple them with a different kind of system. And, you know, in, in, in this case, and what I live day to day is with a, with, with a graph system that Philip Rathle: Neo4j is a graph intelligence platform, the, at the core of which is a database management system. So think, and, you know, for those of you in tech, an Oracle or a Postgres or a MongoDB that basically stores and processes data both analytically and for use by applications. And what's new and different about it, and f- fun fact, it's one of the few parts of the AI stack that has an ISO [00:12:00] standard behind it, is that it represents data the way it shows up in the real world. Philip Rathle: Now, what do I mean by that? Data that shows up in the real world oftentimes as networks, networks of people, ideas, computers, spread of disease, spread of, spread of joy, uh, biology, ecology, these all operate in networks. There are also lots of kinds of systems that operate naturally as trees or hierarchies, so organizations, asset ownership, supply chain. Philip Rathle: And then you have another set of patterns that you could say are like a path through a set of choices, like a choose your own adventure, and this is things like, uh, patient journey, customer journey, et cetera. And those things are all... You can see how they don't really fit very well into tables or spreadsheets. Philip Rathle: And yet, the world has been taking this kind of data, real world, digital world, and trying to put it into spread, you know, tables which are effectively like, you know, [00:13:00] more complex form of spreadsheets. And then trying to understand context and causality in o- in order to then m- move oneself forward in the world in whatever domain you're in. Philip Rathle: And Neo4j is a system that lets you take these networks, hierarchies, and journeys and then store them as a graph, which is, you know, think not graph like a Cartesian system, a graph like nodes and relationships back to Euler, if you remember that from maybe high school or college which is a perfect generalization of these kinds of shapes that the real world moves in. Philip Rathle: So it turns out there's a new-ish ISO standard for database languages, which is the first one in close to 40 years since ISO came out with SQL as a inter-international standard, and it's called GQL, and it's more or less the language that Neo4j invented, evolved, open sourced, advocated, and that's a standard for the category now and for [00:14:00] the world. Philip Rathle: And there are a bunch of connectors and tools and agents and MCP and so on to help get data in structured and unstructured form into a graph in the first place in order to ask questions through a model so that a model can then use data that's in your knowledge graph as a tool and retrieve answers. Philip Rathle: Another fun fact is that models, as well as humans, but models are better at writing Cypher/GQL queries than they are SQL queries for complex questions. And a u-universal observation seems to be that when democratizing complex corpuses of data, um, to end users so that they can ask through a command prompt like ChatGPT, like arbitrarily complex questions, those questions get very, very intricate, and the SQL starts getting bigger and bigger and bigger, and then before too long it blows up. Philip Rathle: Like, it either doesn't run or takes forever to run. And so because this language [00:15:00] and this data model and this way of representing data is better suited for complex questions and, you know, more complex connected domains such as naturally occur in the real and digital world, models are actually better. Philip Rathle: So it, it's ended up being a very popular thing for companies like Walmart, Novo Nordisk, Uber, Adobe to take their LLMs in their application and then couple them with a graph-based knowledge layer such as we've spent the last decade building. So you could say we accidentally spent a decade plus building something that pitch perfect solves for AI hallucinations, explainability, et cetera.[00:16:00]  Philip Rathle: Lots of drug disco-- almost every pharma company is using this technique and our technology for accelerating drug discovery by a year or more over a 10-year cycle, which has hundreds of millions of dollars of impact, not even counting the human impact of getting drugs to market faster. Lots of supply chain anti-fraud. Philip Rathle: I mean, AI brings a lot of promise to the world and a lot of positivity, but then for every actor using technology for good, there are-- there's at least one other actor who's at least as smart using t- that same technology for dastardly ends. And so there are lots of ways in which we have customers who are using graph-based AI to actually combat more sophisticated kinds of schemes. Philip Rathle: There is a, in, in- including Intuit, who secures over 100 million endpoints with, with Neo4j for, you know, various kinds of [00:17:00] zero-day attacks and so on. There, there's a great saying that originated with someone at Microsoft quite a long time ago, which is, "Attackers think in graphs." Meaning attackers know that most computer systems aren't able to, um, understand it when you go through multiple intermediaries, whether that intermediary is another bank account or another corporation, uh, if you're doing maybe f- you know, funny, financial funny business or multiple computers if you're doing cyberattacks. Philip Rathle: And therefore, if defenders aren't able to understand the connected system and how things are connected, then they're not even gonna know that this stuff is happening. They, they might discover years later, and then it's too late, and then someone's walked away with hundreds of millions of dollars or with your data. Philip Rathle: So we're used by all, all, all top banks for everything from cybersecurity, customer 360, entity disambiguation to, again, so that your-- if, if your [00:18:00] LLM knows about five different people, it might give you a recommendation about one person not realizing that those five people are actually one person. It's just I have different silos. Philip Rathle: And then I'll give just one last example is That the way data has naturally grown up inside the enterprise is in silos, meaning you have different departments, each one of which has their fiefdom and their application and their data, and you have the same data copied in multiple places, and then you have certain data that's only in a given silo. Philip Rathle: And what that means is if you just build AI on top of those silos, then the capacity of an AI agent is only gonna be limited to what data it has access to inside of that silo. But the opportunity for AI is not just that. You actually want an AI agent to be able to use all of the intelligence in the company. Philip Rathle: Like, that's a huge differentiator, right? To-- [00:19:00] It's what's gonna cause y-you as a company to survive down the road. And connecting up the silos is another technique that's uniquely possible with, with a graph of let's build an enterprise knowledge graph where you, you take not all the data and recopy it, but let's say the signal, not the noise, leave the noise where it is, pull out the signal, that might be 1% to 5% of the data, and then connect it up, understand how that signal is connected up across different silos. Philip Rathle: Um, so that provides a more intelligent AI that as a byproduct is also explainable, up-to-date, and has discernment[00:20:00]  Philip Rathle: So clearly models are gonna keep getting more and more intelligent, but they'll be more intelligent at world data and still not at company-specific data. So as time goes on, there's gonna be a clearer understanding of what's the right tool for the right job and what does the AI stack look like, and that the AI stack includes knowledge graphs and AI knowledge layers. Philip Rathle: I believe next year is going to be the year of agentic memory and memory systems, and really understanding what do you need in terms of knowledge of both interactions of an individual human being with an AI companion or one or more, as well as, all right, how does that data combine across departments? Philip Rathle: Maybe there's some data that's specific to me, but there's some data that I wanna combine. So there are really interesting questions around different [00:21:00] approaches to memory, different kinds of memory, as well as the scope of memory, be it one person, one agent, one department, um, and then memory portability. Philip Rathle: Whether, you know, do I have a right as an individual to port an AI's memory of me around? So some really, really interesting problems there that are, I think, societal and where, you know, that, that are gonna push the limits of individual digital rights. Philip Rathle: So I, I actually just did a presentation internally to the company where I explained back all the different ways in which we [00:22:00] were using AI inside the company. And I think this m-may not apply to every listener, but it probably applies to most. So one is you have just general purpose chatbots where I can ask anything of anything is number one. Philip Rathle: Number two is I have general purpose AI tools, but they're narrowed to a particular purpose. So something like Figma for design or something like, you know, any, any number of coding assistants. Then you have AI applications that are designed to do a particular thing. So we, for example, have a number of AI... Philip Rathle: new AI tools that, you know, couldn't even have imagined these things existing, that do vulnerability and endpoint scanning. You know, but in, in a consumer world, this could be something to help organize your household. Someone actually just used AI to build an app to help do that, or if you're an architect, something to help you design a house or to track my [00:23:00] finances and so on. Philip Rathle: And then there's AI that you build yourselves, and we have a, a number of tools that we built internally to help us. But let's say just for an everyday consumer, think even in the first three buckets, that AI opens up the possibility of having a number of new applications, like programs that you probably use, like, you know, Intuit for, you know, various-- your taxes or keeping track of your music collection or whatever it might be. Philip Rathle: So new kinds of applications that you can use, you know, note to note-taking and so on, and then your general purpose chatbot, and then you may or may not have one particular domain that's open-ended where you're doing something like coding. And by the way, coding is getting easier and easier and more democratized. Philip Rathle: And if, you know, for anyone who hasn't tried it, I would recommend tools like Lovable or Replit for, you know, like coding for the non-coder to just, you know, type out an application and have it come up. So, you know, as part of that, just remember these are just tools. They only know what they know. They're going to be wrong, and that's okay.[00:24:00]  Philip Rathle: You have no obligation to let them decide your life for you. And just remember that, you know, the, the products that you're using, those all are built by people who are not just using an LLM and trusting it for everything. There are all kinds of checks and balances back there, so trust yourself to be your own, your own check and balance.