How to get discovered in AI search

Narrator:

Welcome to the Practical AI Podcast, where we break down the real world applications of artificial intelligence and how it's shaping the way we live, work, and create. Our goal is to help make AI technology practical, productive, and accessible to everyone. Whether you're a developer, business leader, or just curious about the tech behind the buzz, you're in the right place. Be sure to connect with us on LinkedIn, X, or Blue Sky to stay up to date with episode drops, behind the scenes content, and AI insights. You can learn more at practicalai.fm.

Narrator:

Now onto the show.

Daniel:

Welcome to another episode of the Practical AI Podcast. This is Daniel Whitenack. I am CEO at Prediction Guard, and I'm joined as always by my cohost, Benson, is a principal AI and autonomy research engineer. How are doing, Chris?

Chris:

Hey. Doing great today, Daniel. Cliff, how's it going?

Daniel:

It's it's going great. It's good to see you. You're visible to me, and and that that that's an interesting part of the topic today. An awkward segue into how how do things become visible to us on the Internet these days, which seems to be increasingly through AI platforms, AI, chat interfaces, answer engines, whatever you call them. And today, we're privileged to have with us Liam Dunne and Ben Moore, who are cofounders at Discover Labs, to talk through some of these things.

Daniel:

Welcome, Liam and Ben. Great to have you.

Liam:

Good to be here. Thank you.

Ben:

Likewise. Thanks.

Daniel:

Well, for for those, that maybe are less familiar with this topic in general around AEO, GEO, answer engine optimization, AI visibility, whatever kind of term is around this, and maybe there are differences between those terms. But for those that aren't as familiar, could could you all give us a context for kind of what those things mean? And then also, like, how how you're involved in those topics day to day, what what you're kind of doing at Discover Labs and which is kind of the context that you're doing some of the work that we'll we'll talk about.

Liam:

Cool. Yeah. So I would say and everyone's got a different opinion on this, so feel free to take mine with a pinch of salt. So I would say AI search is like the broad category. And then within that, you'd have answer engine optimization.

Liam:

Some people say, which is a EEO. Some people say GEO generative engine optimization. I view those as the same thing and I just call

Ben:

it

Liam:

AEO. Honestly, for a very simple reason, there are a lot of venture funded companies that have spent a lot of money on that term. And so I'm just gonna fly behind them and lean into it. Background of us, so how we involved with this. So we're co founders at Discover Labs, it's an organic search agency.

Liam:

So we provide end to end services. Now, organic search for us splits into two buckets. You've got like traditional SEO and you have AI search. So traditional SEO, we want our clients to be at the top of Google whenever people are searching related keywords. AI search, we want our clients to be appearing inside LLMs in a way that they want to be.

Liam:

And so we view that as like the AI search side of things.

Daniel:

Gotcha. And I I guess maybe for context, people might be familiar with, like, SEO, search engine optimization. And is it maybe help, orient us. Is, like, SEO is is that still a thing? Is this topic basically replacing that?

Daniel:

Is it kind of, how do the how do the two interact? I I mean, this is always I guess this is also a fluid thing, but if I am a company, this you know, I'm I'm a company in 2026. Like, what, what what is most important and what how are people thinking about the effort they're they're putting into maybe traditional SEO versus these topics. Do you have any thoughts there?

Liam:

I do. Feel free to but I, you know, I can waffle too much, so feel free to ask me questions. I'll go deeper. But ultimately, you won't hear me say that SEO is dead. Not an opinion I hold.

Liam:

I just think the space has grown. We're doing the same three jobs that's on page. So, you know, looking after your owned website, We're doing off page, which is building brand authority and we're doing technical, which is making sure agents and Google systems can access your website and understand information. It's the same three jobs. For me, it comes down to a matter of tactical priorities.

Liam:

And I think this is where people will like get lost. I do think there are some things like depending on what the goal is, right? Like to rank number one on Google for a commercial keyword, you're probably going to use different tactics than if you were to optimize for a surface area like ChatGeepti. Again, you're gonna be doing the same jobs. You're gonna be creating content.

Liam:

You're gonna be building the brand of your company. You're gonna be optimizing for for technical SEO, but the tactical priorities is where it changes. Like for example, what does good content actually mean from a perspective of Google versus ChatGPT? So ultimately SEO is still a thing. It's now expanded.

Liam:

We now have AEO. Those together I call organic search. We're doing the same three jobs, but the tactical priorities have just changed a bit, would say.

Chris:

Could you talk a little bit about, like, if you are, you know, a business owner out there and you're marketing, you have your website out there, you got your social, how has this changed? If you were to take a snapshot of the industry, and a long time ago I was in the digital marketing industry for a while, and that was obviously before AI came a thing. You were focused on these things minus the AI engine optimization, and so how has that changed? How is the thinking of the marketing department changing to accommodate the fact that you now have this whole way that people are going out and searching? All of us are AI first.

Chris:

That's very different from maybe fifteen years ago. How should business owners listening be thinking about that in terms of how they approach the whole thing?

Liam:

Yeah, so broad topic. So It is, sorry. I'd say No, no, no, it's all good. So I would say the first thing is buyer behavior, right? So people researching and consuming information inside LLMs has changed buyer behavior, right?

Liam:

And so the symptom of this is people will see that, hey, like where's all my organic traffic gone? And it's because that organic traffic, a big chunk of it was created by people consuming information on your website. But if people are consuming information inside an LLM, the LLM is the new website visitor. They're taking that information, they're chewing it up and they're spitting it back out to the user inside an LM and so you've now lost those clicks. As I think that's the first thing and so the implications there are, well, how do we measure that?

Liam:

And that's really challenging. Like I'm a marketer, Ben is the engineer and measurement has always been like a nightmare for marketers. And so this concept of zero click researchers has been around for a while. I think AI has just sped that up a bit. You know, we've been dealing with that on social platforms.

Liam:

I think now organic search is starting to feel that. So definitely a change in buyer behavior. You know, it's probably very similar as like when the mobile came around. I think we're seeing some changes there. And then also I kind of touched on it there.

Liam:

There is now just a new visitor to your website. So websites for the last however many decades they've been around have really been optimized around the human visitor. That's why we have like user experience. How do we make the navigate? That's why websites look the way they do today, right?

Liam:

Like how do we shape the navigation bar, like conversion rate optimization, all of these things for human visitor. Whereas now it sees agents visiting your website. And so I think there's some interesting considerations there of, well, what are they looking for and what state does your website need to be in for them to be able to access information and understand you. And so I think there's some considerations there. I don't know where this ends.

Liam:

I know there are a lot of companies out there that are building towards this. I you have like one version of your website for humans, one website version for agents. But I'd say those are like the and then obviously there's a bunch of implications like downstream of all of those things. So does that answer your question? I'm happy to go too deep deeper into any

Chris:

of that. No. It's a great start. I appreciate it kinda laid laid the landscape out there. And I know Daniel has one, but after that, got plenty more.

Daniel:

Yeah. Yeah. I I think my what triggered in my mind, Liam, as you were talking about that is is that measurement piece. One of the things that I think I appreciate about your all's approach and both from a both from a company standpoint, but also a research standpoint is you're interested in knowing actually how the mechanics work around these things and and understanding it at a deeper level. Ben, I'm I'm wondering if kinda turning to that technical side, you've published a variety of research over time and you'll have more coming.

Daniel:

Heard a little bit about that before the interview when we were talking. But, for example, like, citations or mentions in these answer engines or in AI systems. From a technical standpoint, like, obviously, I, as a human, can go into chat GPT or or Gemini or Cloud or whatever, and I can test out some prompts and see what's cited. How from a technical stamp like, what's required from the technical standpoint to actually have have a scaffolding and a mechanism that you can track your like, what does it mean to track your visibility in these systems over time? What are the relevant things?

Daniel:

Like, I mentioned a couple of things, citations, mentions. Like, what are the relevant things that you're looking at? And at a high level, like, what does it take to put the right tooling in place to actually measure and track that?

Ben:

Yeah. As Liam said, it's a big question, but it's a great one. So I would say this is probably the thing that has probably shaken a lot of the SEO industry at its core a lot in the sense that everyone kind of showed up, you know, in the past with traditional SEO and said, look, great. Let's like kind of color by, is it like paint by numbers kind of thing? Like you fill in the colors and we do this, we do that.

Ben:

Everything's great, everyone's happy. And no one really actually thinks about what's going on here. And you could go into like, essentially unless there was a Google core update, you essentially like, the rankings were essentially very deterministic and you had a set of keywords and it was like pretty finite, pretty manageable, all pretty controllable, and again, paint by numbers. And now we've kind of like gone into a system and go really deep on this as needed, but we've gone to a space where there are sources of a like non determinism or stochastic elements that people need to start modeling around. And the SEO industry, and we've seen it as a whole, doesn't have like great statisticians or like great data science literacy, just means blunt.

Ben:

And if you're any serious statistician or control engineer, you would look at the system you're probing and you would say, Well, what are the requirements on the signals I'm trying to measure here in order to bound the noise on these measurements essentially, and the quantities you're trying to measure. And then you would say, okay, well, then you'd set up your model of the system. Essentially, lot of these LLM agents are essentially gray boxes. They're not totally black boxes. There are research papers out there.

Ben:

There are hints in the network traffic and some of the, if you can decrypt the packets, there's also information there. And there's also like distilled models, so when I look at the training data, you can also figure out like some information from the distilled models. And so the thing is to set up basically an experimentation, a kind of engine or pipeline, whatever you want to call it, and then say, well, how do I bound the you mentioned there, citation rate, mention rate, share of voice? How do I bound the uncertainty on those metrics and what do I need to basically bring that uncertainty under, let's say, 5% margin of error. And to be honest, no one in the industry is really doing this, or even really thinking about it, people go on, they buy a ProFound or some other thing, and they're like, Here's my prompt, let's go away, go do this.

Ben:

And ProFound aren't interested in it either, just like to sell you something. And then everyone says, oh look, number goes up, how lovely. And then the SEO industry has like an incredible ability to point at numbers going up and claim that it was them, and then has an also incredible ability to point at a number that's going down and claim its Google core update. So, yeah, I just think in general I can go into how we think about this and how the various systems are displaced here and how we think about bounding this randomness, but in general, you need to approach it like how am I going to model this system that I know hints about and with the data we have, and then how

Daniel:

do we

Ben:

bound uncertainty on those quantities we're trying to measure for, as you mentioned there, Chris, like for the business that has whatever they're selling, you know, gardening tools, whatever it is, right? That's what people need to be getting to, and no one's really talking about that. Having said that, it does take a lot of time. But I also think we've got to the point where it's like we've got to like stock market dynamics. Like there was a change in Reddit recently and, you know, no one also really knows, is it meaningful?

Ben:

So, lots of times we have to come to our clients and say so let's say I'll give you a pro example, right? There'll be like, you'll play a CNN on the background there and someone will say like or ABC News, whatever it is and then they come down and say, Oh, Dow Jones is down 5%. Everyone freaks out. If you actually looked at the data, basing the standard deviation over the last year, you'd be like, Well, this is not statistically significant. I think a very basic check, like it's very basic, but we need to print news, so we do that.

Ben:

And the SEO community has a bit around that, And also people just want to like get eyeballs on stuff. But again, you know, a lot of people waste a lot of marketing hours and a lot of engineering time around things that aren't statistically significant. That's a very basic thing that we can check. Like, that's not that's not broad sense. But that's a very simple example.

Ben:

Anyway, that's why I think we're, like, kind of, like, state of attrition SEO where we are now and where people are freaking out around it and where we need to add.

Sponsor:

I don't know about you, but I'm sick and tired of showing up at AI or technology events and immediately realizing that they're just gonna be full of sales pitches or things that aren't useful for my day to day. I have more than enough to fill my time, and I really need to focus on bringing back value to my company and my day to day work. That's exactly what the Midwest AI Summit is about, October 15 in Indianapolis, Indiana. We we as the podcast are partnering as a media partner with the Midwest AI Summit, and there's just an amazing amount of practical value that you'll get if you attend the Midwest AI Summit. There's actually an AI engineering lounge where you can sit down with folks like me and other AI experts to get feedback on your road map, your architecture, tooling, how you're going about your AI transformation in your company.

Sponsor:

And there's amazing speakers, great food. This is one not to miss. This is gonna happen in Indianapolis, October fifteenth of this year, twenty twenty six. We already had one last year. It was an amazing success, and you're not gonna want to miss the one this year. Check it out at midwestaisummit.com, and you can get 20% off with the code Practical AI 20. So check it out at midwestaisummit.com, and use the code Practical AI 20 for 20% off registration.

Daniel:

So, Ben, I I wanna ask a follow-up question. You did a great job at kind of laying a foundation for maybe some of the things that we should be some of the things that we should be thinking about statistically as we approach this topic. I'm wondering if you could help our listeners understand, like, if I put in a prompt to one of these systems and a brand shows up or a thing that I want to show up shows up, how what are the mechanisms by which that can show up?

Daniel:

And I I know you've done, some research across, like, citations and web search versus Reddit versus training data. Before we get into any of those, like, very specifically, could you help us understand, like, what what do those like, what happens behind the hood? How might I show up mechanically in one of these systems?

Ben:

Yeah. So let me dig into it. Just before I do, would say to anyone anyone listening for sure do not run down the rabbit hole of being like, is this one prompt and I would love to rank on it. You wanna basically model like, from like, basically a topical domain standpoint and get into more of that. But basically, yeah, what happens is, essentially, your prompt is fed in, it's tokenized and then ingested by the LLM.

Ben:

And then there's essentially a reasoning stage, right? So basically, the transformer basically has embedded your prompt and now looks at all those tokens and starts generating its reasoning stage based on its training weights, and basically it has a it's running this forward pass. And there's typically two sources of randomness, and the reason I mentioned this is kind of important, it's like the kernel function. So there's like floating point error in GPUs. That's one source of randomness.

Ben:

And the second source of randomness is the actual temperature itself, like deliberate temperature randomness that generates the next token. And then the final part of this is like, it's also speaking with Liam about this, there's now fingerprinting on these LMs. So they will also nudge and create hashing functions in the background to know where the content came from. Anyway, so it goes through this reasoning phase, let's say, you say, like, a best gardening service in, let's say, Dublin and Ireland, I'll use that example. And it'll then start reasoning and say, oh, it'll have a certain set of criteria, it'll do some reasoning before it makes any queries with its retrieval engine.

Ben:

And it's exposed to a series of tools, and one of them has also got to be the retrieval engine. I'm saying a series of tools because when we get to agents, they could be exposed to a whole host of other things that you're not aware of. Then it accesses this retrieval engine, creates a set of, basically, query fan outs, you've probably heard the term, and then pulls back in basically a set of initial candidates essentially to be pulled into the context window. But before that happens, essentially, they look basically it's not always exactly the same algorithm, but DeepMind had an agree style algorithm, it's called. And basically, it's a way to avoid hallucinations.

Ben:

And so once they pulled in these, let's say, to 60 sources from various places, it's done as, let's say, another set of analysis on it, is almost across the board very similar to the AgriFrame algorithm from DeepMind. It's very similar in Anthropix models. It's obviously, it's in Gemini's models and also ChatMatee, and we've seen this in our own analysis. And also Claude moved to it recently when they moved to Fable and Fable 5.1. Essentially, it's doing now, trying to see like, well, what is the consensus so I can basically use majority voting systems so I don't make a mistake or I don't come up with a solution.

Ben:

And then it says, okay, cool, great. This is our kind of like distilled set of context we want to work with. Now, let's say generate the final response essentially from that cleaned up context. And this actually kind of ties back to like Reddit people freaking out about the final citation. And again, the way I look at like AI visibility, everyone's like, oh, we'd love to know about citations and mentions in the final, like, output.

Ben:

But there's all this other part before we got there. So I talked about, you know, the weights, how you're showing up in the weights, we did some research on this recently. How is the model reasoning? Like, what is what is it thinking? Like, when it makes core decisions around your domain or be that bottom of funnel, top of funnel, whatever it is.

Ben:

It has a criteria it's going through its internal logic, just like a human would, essentially. It's not like human, but you get my point. And then finally, you have like your retrieval engine and response, but you really need to be looking across all of those four pillars in order to dominate your domain. That's how we look at it for all of our clients.

Chris:

There are lots of things to dive into, but before we do that, I'm curious. I actually wanna go back to Liam for a second, as I'm listening to Ben explain these things, the thing in my head was kind of something that you were saying early on about how it changes behavior, and I'm curious, as a marketer, I like the ability to go back and forth between the technical and the human behavior side. How is that changing how the human, now that you have the agent going and doing this, what does that mean for the human that now has this proxy in the agent going out and doing that, and how is that changing the behavior and potentially the transactions that are following up on that, just to kind of come full circle for a moment.

Liam:

Yeah. So I guess I'll look at this from like demand and supply side. So from a demand side, the buyer, I think this is just a much better deal, right? Because if we look at search behavior pre LLM, trying to find relevant information on Google is just a nightmare, right? Because Us marketers just ruin everything.

Liam:

And you'd look for things and like you just get the list of 10 URLs that you then have to go visit and research. Right. And whereas if we compare that today to like chat based LLMs, it's more conversational. So I could be like, hey, I'm Liam. This is my company.

Liam:

It's in this industry. We're doing this revenue. We're really struggling. These are our core constraints at the moment. We're looking for this solution.

Liam:

These are the attributes of the solution, you know, we're looking for a piece of software, must have a free trial, must charge monthly, must integrate with hub spot or CRM. All of these requirements, I can front load that. And then the LLM does its thing and it comes back with personalized recommendations. Now I'm not saying that marked as a, you know, that's our job to influence those recommendations, but it's just a far better deal than like clicking through all those websites and finding the information myself. Like me personally, I just use it for everything, Date night, you know, meal prep, like buying software, everything.

Liam:

From So a demand side, I think it's a greater deal. And so therefore, I think more people are gonna take the deal. From a supply side, the people who are the vendors, I would say I probably look at this from two ways. So number one, if people are front loading all of that context, that now creates a really long tail distribution of queries that those models are using. It's no longer just best cold email software, which is what people would search on Google.

Liam:

It's all of these entities and things are kind of hang, you know, appended onto the end, right? The integrations, the business model, all of these unique requirements. So that creates this long tail distribution of queries. Now, great thing about long tails is they're less competitive. And so I think this has really been where the opportunity has been over the last year is where ultimately I would simplify it down to relevancy, be authority.

Liam:

Now there's some nuance in there but this is really where the edge was and this is where say if I'm competing against HubSpot who has dominated Google search for all the money keywords I'd love to rank for, well now I can create relevant content for those super niche queries that my buyer is searching inside an LLM. So I think there's that perspective is the long tail distribution means we can create relevant content that targets there. And then I think there's the authority consideration is, okay, well, if everyone has relevant content, how does the model then reason over who to trust? Who does it decide who to trust? Historically with traditional SEO, we look at page rank, it's how we know Google assesses authority of pages and domains.

Liam:

Again, it's it comes down to consensus. It's not as weighted. It is my understanding. You know, Google Google have publicly said this. It's not as weighted nowadays.

Liam:

What's like the LLM's version of of that? And this is what seems to be changing a lot at the moment. Right? Because marketers use a bunch of tactics. They work well.

Liam:

Everyone starts using those tactics. Law of shitty click throughs kicks in, and so then they diminish. And then we need to find the next edge or, like, you know, the these vendors, like, patch patch the vulnerability, basically. This is basically what we're doing is we're exploiting systems. You know?

Liam:

SEO is just exploiting Google systems. So I think there's like a whole thing there of like, well, how do I ensure that my brand brand's information is chosen as a candidate and the position I want is communicated back to the user inside the LLM. Like that's how I'm viewing at the moment from like a marketing perspective.

Daniel:

Yeah. That's really helpful. And I guess to that point, there can be you mentioned this Liam, or or maybe it was Ben, like, trying to get people to think more about, like, oh, here's the prompt I want to optimize around. And I know something that, like, people have talked about to me from various aspects is, oh, you need, like, a Reddit strategy to to be to to do great in AI visibility. And, we are we are chatting a bit about this before the conversation came about, but I know you all have done some research, both earlier research on, like, what actually gets cited in these l l LM systems, but then more recent research around the the weight that and the weight and the importance of kind of a Reddit strategy and how that surfaced visibly versus influencing some of that behinds behind the scenes activity that you were mentioning before, Ben.

Daniel:

Do you wanna just kinda help us understand some of the research that you've done there and maybe that Reddit piece or or other things that kinda influence that that citation that are that are relevant in in relation to Reddit.

Ben:

Sure. Yeah. So on the Reddit front, so obviously, as I mentioned there, there's obviously the the fourth pillar of AI visibility, the actual response, there's the reasoning or retrieval stage. And then we did a piece of research on the ChatBeti retrieval engine at the time, and it had allocated, like, an inset about, like, almost a third of its basic retrieval slots for Reddit. So if it could find relevant threads for for that topic or that query, it would basically inject that into the context window at the time.

Ben:

And then, obviously, as it does, like as it gets those those Reddit threads, let's say 30% of it's retrieval, let's say 60 slots, 20% of them are allocated to Reddit, most of those were actually rejected at citation time. So they basically are kinda used to ground the model and formulate its thinking, but also not actually cited in the final response. And then there's also been so that's one aspect of it. And then there's also the long form Reddit like AMAs questions and Reddit training data is also used within the training itself. I think it's maybe around it's typically around the oral HF stage where it's like a reinforcement learning from human feedback.

Ben:

And that's basically golden training data for that. And we actually see when we looked at this, there is traces of basically that is showing up in the waste, particularly from Chat and Key as well as Gemrite. When we looked at the BTD open source models, it's not like super present, you can see it's there. So with those in mind, I think, you know, as I said, people say, oh, well, quotations have dropped. It's kind of like saying, you know, it's one signal on the engine, I would say, but it's not everything.

Ben:

And ultimately, if you wanna, as I said, you wanna, basically, you know, move the narrative for your brand or own in a certain direction, you need to look at all sides. The reasoning engine, the retrieval side of it, as well as the response and the weights. So, yeah, I would say it readily is still very much a big player.

Daniel:

Yeah. And may maybe just to check my understanding here. If I'm understanding what you're saying, this would be like, oh, in that retrieval and query fan out stage, let's say I retrieve 20 sources and 10 of them are from Reddit saying a very similar thing, and maybe one of them is from TechCrunch talking about the same. Like, it confirms the 10 Reddit things. It could be that the answer engine uses those 11 sources as confirming the same thing, and that's what it's gonna talk about.

Daniel:

But it doesn't cite the Reddit things. It cites the TechCrunch thing because it's, like, I don't know, for whatever reasons behind the scenes, cites Do I have the right understanding here?

Ben:

Exactly, yeah. That's exactly it. So essentially, it's like at that, basically, once we retrieve those, that set of, let's say, 20 sources that you mentioned, and let's say 10 on Reddit, There's a series of processes that go on, like, some of that could be, like, re ranking versus the query. And then there's also some later research around basically, like, domain authority and where ChattingTee wants to send users itself. But yeah, all of those can basically compound to say, look, we've used Reddit to ground the answer here, we're confident to move forward with this one, but we actually cited it.

Ben:

That's exactly what's going on there.

Sponsor:

If you're like me, you need a good amount of help keeping your website updated, launching new websites, landing pages, etcetera. And you need an actual platform for your company, not just a builder of websites, but a platform where you can launch and continue improving your site. I'm excited to share with you about Framer, our partner, and what they're doing to actually enable this sort of work. They have agents integrated into their website platform that help streamline collaboration. They can build custom code components.

Sponsor:

They can manage CMS content and much more. They they work side by side with humans. Framer is the pro site builder that's for creatives, teams, and businesses that want a professional site and care enough to get every detail right. Learn how you can get more out of your site from a Framer specialist or get started building for free today at framer.com/practicalai for 30% off a Framer Pro annual plan. That's framer.com/practicalai for 30% off.

Sponsor:

Framer.com/practicalai. Rules and restrictions may apply.

Chris:

Ben, I want to do a quick follow-up on what you were just talking about, and that's kind of that relationship between retrieval and citation a little bit. So it seems like intuitively I might have thought that there was a positive relationship between them and that retrievals would lead to citations. You just talked about the fact that that isn't necessarily the case. Could you talk I mean, seems like that's a significant thing. I would guess that if you're a marketer out there and going with that intuitive approach, that that would be kind of a substantial pivot that you'd have to make to accommodate that.

Chris:

Could you talk a little bit about what that means? What does it mean, the fact that the retrieval and the citations don't necessarily line up that way? And Liam, I'd also love to hear what you have to say about that on the marketing side.

Ben:

So, yeah, so what I would say is, yeah, it is a big change, because typically it's like, look, again, traditional SEO would say basically, look, query an index, we get something back and it's reasonably stable minus a Google core update. And now we have essentially, yeah, multiple multiple steps, multiple query financing on. And not only that, we have, like, re ranking involved as well as, like, DLM basically critiquing the information itself in this agree style algorithm that I mentioned. So, I think basically just to kind of link back all the possible sources of that citation time we talked about like basically citation time optimization to say, like, well, how did we end up in the final citation? And then one of the things is obviously consistency, like, or a consensus.

Ben:

Like, so if I have if I have this long tail query, basically, one of the query files we've been retrieved on, what other supporting materials are out there that are likely to get ranked on that query to support it. So that's like consensus. We have a number of audits we do on this, etcetera. And then the other side of it is, so once I pull in that, like, before, so I pull in that set of sources that basically add consensus that'll help. And also you want to avoid saying, let's say there's some like anti consensus information out there, you want to avoid that building up.

Ben:

So let's say there's something about your brand where it's actually like negative, let's say reviews, and that seems to be like it's published in multiple places, that'll be a good example. That's actually one thing we've built out a lot here is AI perception. And one thing we've found is actually quite interesting is that like, if there's blank space, you can almost publish, you can almost get anything cited. So, for instance, let's say certain industries are very, very sensitive about publishing pricing. And we found this works like insanely well if you publish numbers that are even like yards directionally correct because there's blank space and the LLM doesn't want to hallucinate, essentially.

Ben:

So that's just a good concrete example of it. And similarly, competitor isn't talking about those, you can also, like, do something similar to kind of capitalize on. And the second thing is, well, even when it basically perceives that information it's got from your brand, like, how are you showing up in the weight? So it might have a very negative sentiment around your brand. And if it's looking at data from your brand, then let's say, for instance, it's quite funny when we did the model rank analysis, profound as a really negative sentiment in the score.

Ben:

Essentially that's gonna impact them. That's what I take your time. So like and that's gonna even come down to simple things like what is the name of your brand? So if your brand is like a word that is like, let's say, already existing English word in the dictionary, for instance, that can be an issue. Or if it's like has negative connotations with it, or it doesn't align with the, like, basically overall, let's say it's a security platform and it's like openly or something like that, that could have a So negative there's multiple points there, both on the retrieval side, like consensus, as well as like how the model is, like, basically the sentiment it has towards your brand, and also what.

Ben:

So when these models are trained, going back to the weight side of things, it's all about co occurrence, right? So there's this attention function basically where it's basically going along looking at all the tokens in a certain space and it's associating with you. So if I had like, let's say, guardingtools.com, what are the other entities I've related to as I've gone through the training data? And then when it comes to citation time, it's like, well, that also influences like, okay, well, this, you know, we're talking about X. So essentially, like, let's say the entity shows up that's related to you, you're more likely to get retrieved again.

Ben:

So it's all like, essentially, like an embedded form of a knowledge graph. So typically, of these indexes from Google, etcetera, are built on knowledge graphs, but now with an embedded form in this transformer layer. But you've got to try and assess that. So I'd say basically, there are three areas, well, or primarily two areas: the weights and how it's reasoning on those weights for you, and then the consensus around that. And the weights embed factual accuracy, sentiment, and co occurrence.

Ben:

That's how I think about it. But, yeah, it could be like it's not it's not easy to it's not easy to operationalize that and then, like, start winning enough for for, like, the key if you just get profound or something like that.

Daniel:

Liam, I'm I'm curious. So Ben just described kind of of course, there's this whole chain of things that influences what eventually shows up in the actual response. Everything from the the model weights to this query fan out to the agreement algorithm, the the, the naming, all of those things, co occurrence. I'm wondering, like, over time, obviously, this industry of AI visibility is evolving. Right?

Daniel:

And so at a certain point, maybe it was like, oh, you need FAQs on your website. This is how like, the easiest way to to show up. Now now you're kind of thinking across all of these stages of of the process. So how do you when when you're engaging with new with new clients, different brands, like, what is are there any generic kind of takeaways around, like, the strategy of where where you start and what what is kinda near term and and long term important for your brand so that you're hitting all of these things, but also you're able to make progress quickly? Any thoughts on that?

Liam:

Yeah. So I would just say three words and then I'll dig into them. So I think it's relevancy, consensus, consistency. Right? And all of this you could also put into buckets of like tradition, like again, same three core jobs I said at the beginning.

Liam:

So and it really depends on what your current state is. Like, do you have a website today? That's gonna be super important and like all the foundational stuff really matters. Like the I used to run a paid ads agency and I was like, man, why does some companies like really succeed with ads and others don't? There's lots of factors that go into it.

Liam:

But one of the key moments I had is messaging. Is like the companies that really succeed in marketing, they they have they have a really defined ICP. They really know what their product does and they're able to map that to clear messaging and they have really good differentiators. And I think that's really important in AI search is like, what are those terms you want to be associated with? Like, what do you actually do and how do you translate that into a language that people are going to be searching?

Liam:

And I think that's how you create that relevant content. But you can't create good content if you don't actually know who it's targeted to or what you do. And so I think that's like the relevancy piece. I think that's how you ensure that during that career fan out, these LLMs are coming across your content. Now there's we can go deeper on that.

Liam:

Like I think you shouldn't just be posting blog content, you need a website, you know, your service pages, like, you know, all of that good stuff. I think then consensus is and Ben covered this in enough detail is basically my content is like me throwing my hat in the ring. Like, hey, we have relevant stuff over here. Consensus are all the votes of confidence. Like why should we choose this company?

Liam:

Why should we trust them? Mark has always asked me like, well, hey, where should we focus our efforts? And my answer is everywhere. Really what we're talking about here is it's a budget and time constraint, right? Like if budget and time weren't a consideration then be everywhere.

Liam:

If they are a consideration, well, that's really gonna depend on like industry. We're quite bullish on places like Reddit. One of the advantages of Reddit is a lot of B2B companies don't really know how to do it well. And so just like because of that, you can gain an edge. I think also doing just traditional, you know, digital PR, I think building a presence on social channels is super important.

Liam:

If I was to, you know, gun to my head, where should I focus? I'd bet on YouTube. It's actually a native search channel. It's owned by Google. I think it's gonna be around for a while.

Liam:

I don't think you should, I'm personally not that bullish on LinkedIn as a search channel, as a social channel, absolutely. We're on there and I really believe in it but for search activities, I don't know. I think also building customer advocacy. So we operate in the SaaS industry, that's gonna be your G2s of the world, right? Getting people to say positive things about your brand on the internet is always gonna be a vote of confidence for consensus.

Liam:

And then consistency, this is kind of where it closes the loop. What I kind of opened the thought with is you need to get really clear on what you want to be known as, as an entity, as a company, right? So if you go to our website, you're gonna see we use similar words across the place because I want those to be embedded in these models and the world. And so whether it's your activities on Reddit, you're posting videos on YouTube, you're posting content on LinkedIn, even like your company LinkedIn profile, content on your website and on your blog, always presenting yourselves with a consistent message. Right?

Liam:

I think this is like and so like one of the key contrasts here is if we look at like link building and SEO, the big check-in the box was, hey, we got a link. We've got a do follow link. You know, we're gonna get some link juice back to our website. In AI search, I actually think the blurb surrounding, even if that link didn't exist, the blurb surrounding you as an entity is far more important because like who are you as a company? We want that to be really dense.

Liam:

Right? Like I want if if an LLM picked this paragraph and put it back inside to an answer, are we being positioned competitively? Like we're thinking about it from that perspective. So that's like consistency. Now we can get into like the tactics and all that stuff but it really depends on like your current state because some people we can talk about Reddit and all these tactics but some people might not even have clear messaging.

Liam:

Right? They might just have a website with a few paragraphs of text on it. Right? And so then your priority is gonna be slightly different. So it it does depend on on where you're at.

Chris:

I'm curious with with kind of that with the focus on Reddit and and the potential there with, you know, what tends to make one Reddit thread more likely to surface than another? Are some of the signals that you would have that are upvotes and comment count and user reputation, do those correlate with retrieval? And so when you're looking across different threads on Reddit, does that make a substantial difference? How do you think about that?

Liam:

Yeah, so I'm not gonna speak in absolutes because I think these things are always changing. What we have observed is engagement in terms of upvotes. We haven't seen that correlate with increased citations. What we have seen highest correlation with is the actual content itself, right? So I always like to break things into like easy ways to understand because that's how my brain works.

Liam:

But if we just look at this from the principle of user prompts LLM, LLM does query fan out and it's looking for information contained within those queries. There's lots of things and entities, right? So if I go back to that prompt example, SaaS founder, doing this in revenue. In those query fan outs, it's containing all these entities, right? And so I want those to be contained within my content on Reddit.

Liam:

And so we found very often that the passage being extracted from Reddit is like 50% down the page and the comment has like one up or two up votes and it's because the passage was relevant. Now, I think that might change over time, right? Because similar to how relevancy is super important, but authority is becoming even more important. I think these models are gonna keep similar like, you know, how Google always releases a core update is, you know, they it's a cat and mouse game between these providers and marketers who are constantly trying to catch catch up. And I know there are commercial agreements between OpenAI, Reddit, Reddit, Google.

Liam:

And so maybe things like upvotes, maybe even the history of those accounts, things like that might factor into it. I don't know the complexities of that might change over time. But yeah, just what we we've observed so far is the overlap between what the user is searching and the the content on Reddit.

Daniel:

Liam, you you started getting us towards kind of, like, how things may or may not change, how how they're evolving over time. As we get kind of close to an end here for for this conversation, I'd love to close out by just asking each of you, maybe circling back to Ben to start. Each of you, like, as you look forward, what what is kind of top of your mind either in terms of, like, oh, there's, like, a huge opportunity here for for people that are willing to to jump into it or something that's like, oh, there's a really open challenge here. We haven't figured it out yet. There's more work to be done in this area.

Daniel:

In either of those areas, what are you thinking about as you end your days or as you start out in the morning that's kinda top of your mind going into this this next phase of work?

Ben:

I'll let Liam go. He's ready to go.

Liam:

Okay. So I'd probably say two things. Off-site, like how are those third party earned sources gonna influence how I am or our clients are perceived inside LLMs. I think there's just so much change happening there. And ultimately, brand is the final moat.

Liam:

Right? You see this plastered everywhere is like and so I really think that's where things are gonna change a lot. And so I'm always thinking about that. I'm always thinking about, you know, our clients only have certain budget and time they can allocate, and they're on strict timelines. And so how do we increase the probability that the bet they make is gonna have, you know, the highest expected value?

Liam:

So I'm always thinking about that. The second area is agent accessibility. So I think everyone at the moment for the like everyone I'm speaking in such generalities, but I think everyone's been focused on discoverability, Right? Like, how do we get my content discovered? And I think there's so many edges to gain.

Liam:

Now this is this is going back to my original point. It's like walk before you can run. But we're quickly moving to a place where it's not just gonna be agents crawling your website and retrieving information. They're actually gonna be taking actions on your website. And we've already seen evidence of this.

Liam:

We have clients reporting to us that, hey, an agent, an AI agent booked a demo last week. We're seeing things like that happening. And so then that just again, websites have historically been optimized for humans. Humans, if your website takes a little bit to load or like your form, your demo form is confusing, it's okay. We can figure it out.

Liam:

Right. We can, you know, use our intelligence. We can figure it out. It's not fun and it's gonna impact conversions, but we'll figure it out. Agents, are they gonna wait around to figure it out?

Liam:

I think there's actually some standards that I'm hoping you've seen some movement here from Google. And so I think that's gonna really impact how websites look and behave. And I think that's really exciting because I don't think anyone's looking there yet. I know Ben wants to research more in this area. I'm kind of holding him back a bit because I don't think it's sexy enough just yet.

Liam:

But I see that definitely as the next frontier. And we have clients that are generating a lot of conversions, you know, tens of thousands of conversions per week. And a lot of their users are coming via command line interface. So they're saying, hey, I wanna build x. And then the agent is the one going out there and coming back like a private procurement team.

Liam:

And and like not just saying, hey, recommend these tools. They're saying, hey, I installed these tools for you and they're good to go. I think that expands the surface area that we need to look at. So, yeah, I'd say those two is is what I'm thinking about.

Daniel:

Awesome. Yeah, Ben, anything to add there?

Ben:

Yeah, definitely, as Liam said, the agents, I just think it's moving from like a, I always describe it as like a read only environment to a read write environment is where we're headed. And as Liam said at the very start of this conversation, it's just a better deal. It gets a better deal to get all the context in one place, just talk to an agent. It would be an even better deal if you didn't have to fill out the form. It would be an even better deal if you didn't have to pick a slot on the calendar, yada yada yada, get the idea.

Ben:

And so that's for sure where we're heading. I think, you know, we've seen evidence of this and if you look at the Google Web OCP program, like this is gonna it'll probably it'll be very, very slow and overnight go boom. And that's what I expect to happen. Just because it's human, it's just human behavior. It's just easier.

Ben:

It's just easier. Lowest friction path, that's where we tend to. And then I would say, I do think to date we've been very focused around basically, k, like, here's some query fan outs and here's the citations. How can I map them up? And we're kinda, like, playing this this a little bit of a most marketing teams are playing this game of, like, here's my prompt and I'm I now rank for it kinda thing.

Ben:

But I think people are gonna start focusing more on, like, the weights and then what's driving the weights longer term because the two two reasons is that well, the the market in general is, like, sophisticated. People are becoming more aware about how these systems work. The second thing is that the training cycles are actually coming down. So when we first started, was cool. Had like a new version of Chapter B key once every like nine, twelve months.

Ben:

Right? And now obviously these engines are getting trained like on a weekly and biweekly basis at the top layers. So that even, like, the top layers are being retrained and being tuned, and then the backbones are being retrained every, let's say, three months or so now. And so they are more influenceable. And if you can get influence or imprint in those tokens, it's worth, like, orders of magnitude more than being in a in a in single context window.

Ben:

Having said that, you also ultimately want both and it's not easy to get into that training data. But that's where I I see I see things heading. In terms of otherwise, outside of that, yeah, offside is gonna be a big thing. I think being able to because essentially what to the kind of a kind of overall trend is that Google is having to or all these all these agents are having to avoid, basically, like dead Internet theory. Right?

Ben:

They're trying to avoid training on, like, basic model collapse, training on their own systems. And so they're looking at all this content out there, that's part of what the fingerprinting release is. It's part regulatory. It's part it's probably a part IP play. It's part like multiple plays they have out there.

Ben:

But it's ultimately it's one of the reasons I also think they have it is they don't want to train on their own. So if you post a Reddit with, like, a generated Torrent comment, they will know that. They look up their hashing function, very simple. Boom, the probability of this being AI generated from our model is very high. And they just, they essentially, I think that's gonna ratchet up, and I think there's gonna be a lot of cleaning and basically all these like, basically the protection of human level context.

Ben:

And actually, when OpenAI started out, they have this manifesto online to they wanted to offload the verification of text to third parties. Like Trustpilot's a good example, they're in partnership with them. They'll keep trying to do that, and I think they'll keep trying to up the ante here. So, it'll get harder and harder to basically evade them. So, you need to keep climbing the edges around that.

Ben:

So, that's where one other, let's say, AI content, bot content, I guess, or bot behavior in general is gonna become, like the human signal is gonna become, very valuable there, but also emulating it well too.

Daniel:

Makes sense. Yeah. Well, I I was writing down furiously in the background a few things that that you all mentioned kind of throughout. I need to level up on my understanding things are just moving so fast. I really appreciate both of you joining us to like, with your expertise, really helping us and our listeners understand this topic.

Daniel:

I would encourage everyone listening to go check out Discovered Labs. In our show notes, we'll link some of the research that we talked about here. We'll link their research page, which is which is really great. Thank you so much, Liam and Ben, for joining us. Looking forward to having you back on the show sometime when everything is different.

Liam:

Thank you so much. Thanks for having us.

Ben:

Thanks, guys.

Liam:

Cheers. Bye bye.

Narrator:

Alright, that's our show for this week. If you haven't checked out our website, head to practicalai.fm and be sure to connect with us on LinkedIn, X, or Blue Sky. You'll see us posting insights related to the latest AI developments, and we would love for you to join the conversation. Thanks to our partner Prediction Guard for providing operational support for the show. Check them out at predictionguard.com.

Narrator:

Also, thanks to Breakmaster Cylinder for the beats and to you for listening. That's all for now, but you'll hear from us again next week.

Creators and Guests

Chris Benson
Host
Chris Benson
Cohost @ Practical AI Podcast • Principal AI / Autonomy Research Engineer specializing in fully autonomous UxS swarming with embodied intelligence.
Daniel Whitenack
Host
Daniel Whitenack
CEO @Prediction Guard & cohost @Practical AI podcast
How to get discovered in AI search
Broadcast by