Today on Sacking Growth, Dr. Augustine Fu joins host Matt Chanel to talk about Adfraud. They dig into the details of the evolution of Adfraud, how the bots have grown with marketing media, and some things that you can do to avoid wasting all of your ad spend just on robots. Hope you all enjoy.
I am really excited to bring in the Dr. Augustine Fu, who is the, who is, I would call you Dr. Fu like the, I know in the intro we had another word, I would call you the Cape Crusader of Adfraud, basically. So if, if, if bots are the bad guys running all over Gotham, I would say you are literally Batman throwing them in jail and keeping them from running them up, essentially.
That's a good one. Thanks, Matt. All right. So there we go.
So I want to talk, we're going to talk today about Adfraud and B2B, and B2B advertising. And I really wanted to bring Dr. Fu on to talk about what some of our favorite ad channels and then some of the specific ad types within it, because it's not enough that you advertise on Google or meta or LinkedIn or Reddit or use programmatic or CTV. It's also like, what are the controls that you need to put in place to combat this stuff for yourself?
On top of like having that extra layer of security, which is what I know your tool provides Dr. Fu, but you can even without that tool, can make, put controls in place to avoid wasting ad more ad spend than you otherwise would think you are due to, due to bot traffic and bots clicking on your ads. So Dr. Fu, I want to give you a moment to introduce yourself.
And then we're going to go right in to all these channels and all these ad types and kind of walk through the trap doors that exist for each of them. So go ahead, give a little bit of your background. I think it's super impressive. And also just how you hatch the idea for Fu analytics in the first place.
We went over that and prepped. I thought that was really interesting. Sure. Sounds good.
Augustine Fu here. I've been working in digital since the very beginning. So 95, 96, so early 96, I left McKinsey company and kind of jumped in with both feet into what is known as Silicon Alley here in New York, as opposed to Silicon Valley in California. So back in those days, we still needed to convince clients they needed a website.
So those were really the early days. But fast forward, many years, I've been on both the client side and American Express and the small business division doing digital marketing, as well as on the agency side at both IPG and Omnicom holding companies. My last role was chief digital officer at Omnicom's healthcare consultancy group. So we were surveying, we were a group of eight agencies serving pharma, healthcare and med device clients.
And back then, I was looking at some of the Google search campaigns that they were running. And we're seeing strange things like greater than 100% click-through rates. So how is that possible? Humans don't click on ads that much.
And the other strange thing was that we saw the clicks continue to come to the website even after the campaign was over. So there's literally no ads to click on, but we still saw the bots clicking through the site. So back then, no one could give me a straight answer about what that was, but I'm sure all of us can surmise it was bot clicks and bot traffic. So that's when I left Omnicom and started building a set of tools to help my clients, the pharma companies audit these campaigns.
Initially, I was using the tool myself, right? We would put an onsite code on their sites, figure out which portions of those clicks were bot traffic, and then help them go upstream and start blocking those bad sites from actually running their ads. So in the early days, it was more of a consultative model where I basically screenshot from the dashboard, put in a PowerPoint and made the recommendations. But by 2020, I opened up the platform and just gave it a name, Fluentalytics, so that other people could log in and do it themselves, because there's only one of me and other people can actually just look at the data and see very easily how to improve their campaigns.
So that's kind of how Fluentalytics came about. I didn't really set out to build a platform for sale, but it kind of evolved into that. And now it's being used by advertisers and agencies so that they have more detailed data with which to optimize their campaigns. So that's the last almost 30 years, some rest.
Yeah. And doing the Lord's work out there with regard to add fraud for sure. So I want to talk through the concept of add fraud a little bit, because I think there's this juxtaposition in B2B tech advertising and BV marketing really in general, which is this drive to make marketing as deterministic and exercise as possible. When in fact, a lot of what we do in marketing is probabilistic.
And what we want to do with as much of our executions as possible is increase the probability of a good outcome. And largely with what add fraud or combating add fraud does is help to lubricate that probability a little bit more, because there's really no more frustrating thing than having small or large budgets, right? And having your ads not appear in front of humans. So with that said, like, talk through a little bit the concept of bot traffic and how bots actually are getting into some of just getting into ad platforms.
And we can start with Google, because any Google is really where this sort of started because this is really the advent of digital advertising, right? Google and churches Google search, then they moved to some more of these programmatic plays. So let's start with Google and talk about like where add fraud comes from there with that particular channel, since that's where the majority of spend is happening right now for many brands. That's where really fraud is about as per base as it gets depending on that type.
Yeah, I'll boil it down to something very simple. Instead of thinking about Google as one holistic whole, you need to think about Google in two parts, the main property, Google.com, and all of the sites in the search partner network, sites and apps, which basically run their ads and get a portion of the ad revenue right on a cost per click basis. The reason we have to separate those two things out is that we all understand that humans Google things, right? So they go to Google or the app and they search for something and then they see the organic results or paid search results and they click on it.
Those are fine. Where the fraud comes in are all the outside sites and apps that also run those search ads, right? These are typically syndicated search other sites that use Google search to kind of make money. The problem is when you have an outside site, the CPC ad revenue gets split between the site and Google.
So the site now has a way to make money. You now have the motive. And then the means is basically using bots to type in the search term load the ad and also click on it because these are all CPC. So you have to click on it in order to earn the revenue.
So as you can see, if you're a fraudster, you would set up a website, copy and paste some code from Google to run on your site like AdSense. And then you would basically use bot traffic to cause the ads to load and click on them. And then you have a stream of ad revenue. So the main distinction is between the main property.
If you leave your ads only on Google.com, you're going to get it in front of humans. Now it's not to say that Google.com doesn't have any bot traffic. They do. There's a lot of bots trying to scrape their content and search results and whatever.
But at least you're avoiding the most obvious fraud, which is mostly in the search partner network or what we'll more generally call audience extension or just a partner network, right? So that's kind of how you should think of every single platform. So on Facebook, there's the ads that run on the main app and on Instagram, those are great. And then you have the audience network, Facebook audience network, which again, are all the outside sites and apps that have a whole bunch of crappy sites and apps in them.
And that's where you're going to run into the problems of those sites have both the motive and the means to rip you off as much as possible as fast as possible. So hopefully that kind of gives you the background on how to think about it. Right. And so really with like some of these, some of these, some of these terms, like in Google, there's really kind of two, two things that play here, right?
Like one is the sort of programmatic ad type where it's appearing in front of a bunch of partner sites, right? And almost anybody can get in to those partner sites because there's not like really that much of a vetting process that goes into checking them. And then on the search side, you know, bots will scrape websites. People will also upload lists too and still try to control who appears in the auctions for some of the terms that they appear for.
But even that doesn't necessarily protect you either because if you'll, you'll notice this with you have your tool loaded because you think again, going back to putting controls in place, you'll put these controls in place with audiences, but yet you'll still see UTM clicks coming from other countries or even from other places that like you are declared bots, even there's even friendly bots on top of their main issues. So talk a little bit about those kind of tensions that exist overall. And then we can talk about how to control for each of these. Sure.
So basically, I'll use an example from my clients in pharma. So a lot of them, they do upload lists, like you said, so whether it's email lists of people who have signed up for their information or lists of NPI numbers, which are national prescriber IDs. So these are typically adopted would have that so they could prescribe medications. We're talking about just to be clear, we're talking about customer matching lists and things like that.
Yes, exactly. Exactly. So when you upload the NPI numbers or the email addresses, you typically go to a databroker like LiveRAMP and have them do a cookie match process, because what you want them to do is find all the cookies that pertain to this individual doctor. That's where it starts to get messy, because what LiveRAMP will do is they will make their best guess and they'll basically say these two or three dozen cookies, we think are related to this one doctor.
So then you have this cookie pool, which are supposedly matched against that list that you uploaded. And whenever those cookies show up in programmatic, they'll show up in the bid request. Oh, it's that doctor. Let's go bid on that.
And the pharma companies are willing to pay a much, much higher CPM because they think they're targeting doctors or even specialists. Same thing happens with patients and things like that. People who have a disease state that they want to get their ads in front of. Now, that's all fine and good.
That's been the assumption that the pharma companies and the advertisers have used over the years. Same thing would be to be to upload a list of customers or prospects or whatever and they'll try to match that. The first thing to realize is that the cookie matching process is extremely dirty or not accurate. It's an approximation.
The second thing is cookies turn over. So you have to kind of re-match them and re-update or refresh your cookies, otherwise you're advertising to stale cookies. The third thing, which I don't think a lot of people understand, is how the bots can actually start impersonating the doctors or get into your target audience. It's actually very simple.
I have an article out a few days ago where all the bots have to do is harvest the cookies and start replaying them in the bid request. When they do that, they can easily see which cookies make them more money. So they literally don't know who the cookie is. They don't care if it's a doctor, they don't care if it's a patient or whatever.
They just know this cookie makes them five times higher CPM than this other cookie. So they'll keep replaying that. So in that case, we actually have seen cookies that start off as blue flip over to red because that's when the bots copy out that cookie and they're going to replay it as long as they can and as long as they make money. It's ultimately that cookie or that identifier is going to turn over either expire or get lost and they'll have to find a new one.
So in those cases, it's actually the bots that are replaying those cookies pretending to be the doctor that you're targeting and you're paying an extra CPM. And the reason we see all these crappy websites and mobile apps show up in the data is because they want the ads to load on those crappy websites and mobile apps. We call them cash outsites because they're also controlled by the fraudster. So when the ads load there, they make the money.
Okay, so I won't go into too much more detail, but just understand that there's a way for the bots to now impersonate your target audience. And that's how it can still get into kind of stealing your budgets. I want to be clear too, because you said, not click, you said, we play the UTM. So basically, what they're doing is they're finding that UTM from Google and then they just keep reloading the site, not necessarily rerunning the search, right?
Exactly. Because once they copy off that click through URL, they click through URL already has the UTM source equals this and this. All they have to do is copy that URL once and then just keep replaying it. That's kind of what we observed more than 15 years ago, right?
When I was leaving on the com, those clicks kept coming to the site even after the campaign was over. So there was literally no ad to click on to begin with, right? But that's how easy it is. And all of those are visible, right?
It's in plain English, right? It's not encoded or anything. So it's trivial for the bad guys to use copy off those URLs and replay them. That's crazy.
Okay. So let's talk about ways to control for this in Google. I'll go a little bit through my own experience using through analytics and things that I've noticed running. And then I want you to walk through really the big crux of this, which is the list inclusion and exclusion list.
But for me, when I use through analytics myself, with my clients and I'll run search campaigns, for instance, and the first thing that I always look for are the countries. What countries are my campaigns in peering in front of versus the one side actively targeting. And no matter what, no matter what, you can put as many want to target this area or this area and this area, whatever you include, you still have to very actively exclude because there's still going to be people who are going to mimic or spoof, I guess, IP addresses and actually come in from other countries and still click through on your ads, fraud, generally. So even so you ideally have to have everything UTM that really neatly, you should also ideally have a pixel per landing page or dedicated, the pixel dedicated for pay landing pages only.
That way it's more easy for you to see even in breaking down possibly by channel. Then look at the countries that you're appearing in front of, you're going to find countries that you're appearing in front of that you weren't that you did not believe you were paying for and then actively exclude them to the greatest extent possible. That's been my experience with it. Yeah, it's very simple.
Basically, Google assumes that if someone is searching a particular keyword, they're looking for that information regardless of where they are. So they always override the geo targeting, right? To serve the ads. That is crazy that they just even hit actively include.
They still are like, well, we're going to let you appear for it and we want to get paid. Yeah. So I think that's when, like you said, we just need to actively check to see, okay, now, if it's 98% going to the country you want, it's only 2%, it's not really that urgent. But if you're seeing 5 to 10 to 15% going to some other country where it's not relevant, then you just start negative targeting those particular countries to clean that up a little bit more.
So I've seen that for years and years, but it's really based on that assumption where Google just says, oh, well, if someone searched for that term, they must like that or they must be interested, so we're going to serve it anyway. So you have to go in and control it by adding negative targeting. If the other countries are a substantial portion of the clicks, right, and you basically see that by putting the food analytics tag on the landing page, because when they click through, you get to see it arrive on your site. And here's the kicker, you won't see that in Google Analytics, because if it's a bot, GA is required by the IAB to filter out all bots and spiders in the IAB list of bots and spiders.
So what happens is that they simply don't show you, and that's problematic, because then you won't even know where it's coming from and that you have a problem with your search campaign. So for most cases, the advertisers just add food analytics next to GA, not in place of, right? They just add food analytics to the same landing pages. So if there are those discrepancies, you can use food analytics to troubleshoot, because then you can actually see, because we don't discard anything.
We record everything and we also tell you what type of bot it is, right? Yellow means search crawlers, orange means declared, red means bad bots. So then when you can see where they're coming from and like, oh, here's the UTM, it clicked through from my search campaign versus my TikTok campaign, then you know which ones you need to go make some optimizations. That's the big thing for this is also anytime you're going to run something like PMax, for instance, which generally you would not recommend writ large, but if you are going to run it, 1000% must have an inclusion list or whitelist and not actually just try to lead instead with an exclusion list, correct?
I think from my understanding PMax, I don't think you can use an inclusion list because it's an automated algorithm where they decide where to put your ad, but they just, I think through community, you know, people, enough people shout it out at them, they just launched a feature where you can upload a block list into PMax. So again, like you said, the best way to do it is if you want to use PMax, that's fine. Just make sure you have something like a food analytics pixel on your landing page. So then you can see, okay, are the clicks coming through?
Are they red or blue, first of all? And also you'll start to see, okay, why am I getting these clicks from all these crappy sites that nobody's ever heard of before? Are these flashlight apps or kids coloring apps, clicking through to your site? Keep in mind, in those cases, it's not actually someone clicking through from the flashlight app to come to your site.
There are mobile apps that are designed for ad fraud and click fraud, where they have a built-in browser. And again, when they add loads, they can actually see the click through URL. And all they have to do is load that URL in the hidden browser. And that registers as a click through, even though it was done by the app, and there was not even a human behind it, right?
So when they load that landing page to you, it looks like, oh, okay, I got to click, but it was a flashlight app and no one actually meant to do it. But because they did it, they can earn the CPC revenue. So again, you can see that by having a food analytics tag on the landing page. And then if you do see that kind of fraud, you would then go back and upload a block list to tell Google P-Max not to serve when those things show up again.
Got it. I wanted to touch, I really want to touch on CTV and DSP, but I first want to touch on Reddit because I was going through this one scenario with you in Parep where I was talking about Reddit is, for some Reddit, it's a very trendy channel for B2B right now because the CPMs are very low. There's this illusion of control because of the subreddit. So there's community that you think is very active and there's also just a high level of engagement on the platform.
But one thing that I see fairly constantly on it is Reddit over delivering or reaching more people than your audience otherwise suggests the audience actually is. You explain how this, why this happens in a way that I never thought of before and this actually kind of brings a bit of an AI play into it as well. So walk through what is happening on channels like Reddit where you reach actually outpaces your audience and why that happens. Yeah, I think in that particular case, it's really a matter of the scrapers that come to Reddit.
So there's been a lot of other published reports, which I haven't corroborated myself, but they said a lot of the AI, like chat GPT and whatever, they actually went to harvest the free content off Reddit to use to train their AI. So basically, and we've known this for years, just like Wikipedia has constant amounts of bots there because they're there to steal the content and then Reddit and any other of these published open websites will have a whole bunch of bots there just to, they're scraping their content. So Reddit in particular, yes, there are humans on there using it, but there's also a whole bunch of scrapers and bots and whatever on the site. And when they're loading the page, Reddit doesn't have a mechanism to filter out the bots.
So your ads still going to get shown. And in some cases, some of them will click through because they're just clicking through on anything they see. So you will start to see that and there's going to be some percentage of red in your click through when they arrive on your site. It's kind of, you know, if you still want to advertise on Reddit, you almost have to say I'm just going to eat that 10%, 20%, whatever percent.
It's kind of like on mainstream publisher sites, there's always going to be scrapers on CNN.com, New York Times.com against to steal the content. And the filtering just doesn't work that well. They're not supposed to show an ad when it's an obvious bot. But in most cases, these bots are not obvious anymore.
They disguise themselves. They don't say their name honestly and whatever. So when the bot hits the page, add loads, right? And those are things you pay for.
So again, if you still want to advertise on those sites, the best thing to do is just make sure you have the analytics in place to see the quality of the clicks coming through to your site. And then you will know whether you have a problem or not. Right. Okay.
I want to circle to LinkedIn as well. LinkedIn is super popular for B2B. There's a lot of illusion of control on that channel for people. They think they're targeting these job titles.
They think they're appearing on the LinkedIn feed. We know what LinkedIn tends to do. They tend to auto enroll you into audience network and audience expansion. Expansion is a little bit more of just guessing who your IP is.
But audience network is a little bit more of a problem and a little bit more where product tends to get fairly pervasive. I want to talk through that because it segues into CTV. But let's start just with LinkedIn and kind of where fraud exists and is pervasive there and how to control for that. Yeah.
Again, the answer is simple. The answer is free. The answer is no cost. You don't need any other tech.
Just go uncheck the checkbox that says LinkedIn audience network. That is all. So basically, same concept. There are humans on LinkedIn.
There's also a bunch of fake accounts. You can't avoid those, right? But you want to keep your ads on the main property because that's where the humans are log in. Everything in the audience network, right?
Outside of LinkedIn, that's probably 90% fraud. Right? I do see some good sites in there. But again, even those sites like BBC.co.uk is not going to have zero bots just because they're public site and there's scrapers, whatever.
But there's a ton of sites in there, again, that have both the motive and the means to commit fraud. So you're going to be exposed to all of that. So unless you have a very specific reason to not turn off the audience network, it's always a good idea to turn that, like uncheck that checkbox so that you can avoid 90% of the obvious fraud. When your ads run only on LinkedIn and I have 15 years of doing my own LinkedIn advertising, so I can tell you for sure.
Just turn off audience network. Audience expansion, that's fine because that's just still on LinkedIn. It's just like people related to the target that you specified. That's totally fine because it still remains on LinkedIn.
But I would just say, be vigilant about that. Again, measure the clicks that come from LinkedIn to your site to see if you have a large proportion of bot traffic. And on LinkedIn, you can also upload a list, right? If you are using Audience Network, you can upload a list of sites to block.
So those are ways you can control but the easiest no-cost, free way to avoid 90% of the fraud is to just turn off audience network, the outside sites and apps. All right. I want to second to CTV on LinkedIn because it's a pretty new ad type and CTV is really hot right now in B2B. People are looking headed through like Vibe or through Mountain through LinkedIn.
Again, the illusion of control with I am going to target CFOs and use LinkedIn's CTV in order to facilitate that for my best video content. In theory, it sounds great in practice. It's actually more rife with trapdoors for fraudulent impressions. Talk a little bit about that because we were walking through this yesterday and you were talking a little bit about what are some tells on that in relation to CTV?
And we're going to use LinkedIn as an example, but this actually exists in any of these CTV platforms. So let's go through a little bit about what to be vigilant about as you embark on CTV as an ad type for yourself. Yeah. So to think of LinkedIn in this context, it's just a place where you can buy CTV ads, just like you can buy CTV ads through trade desk or you can buy CTV ads directly from Disney+, Hulu and ESPN and whatever.
So the same problems exist, not just on LinkedIn CTV, it is trivial for bad guys to fabricate CTV bid requests out of thin air. They don't need a TV, they don't need a streaming stick, they don't need a Roku box or anything. They can just fabricate the XML that looks like a CTV bid request and just declare it's this Roku app or that Roku app or ESPN or Disney+, or whatever they want. And because in that environment, we can't run JavaScript, no one else can run JavaScript for that matter.
There's very, there's much less accurate measurements, right? We rely on a pixel, we can still gather some information, we can still tell you the obvious fraud, but it's much more limited measurement. So the main problems are that the bad guys can easily fabricate those CTV bid requests. And I'll use some examples over the years to illustrate.
Maybe eight years ago, Grinder, the mobile app was one of the first to get caught fabricating CTV bid requests so that they can make more money. Because CTV, CPMs are way, way higher than display ads or video ads. So, you know, they said, oh, well, why don't I just, you know, instead of generating a bid request for a display ad, I'm just going to generate a bid request for CTV ad. They sent those into the endpoints and they were successful until they got caught.
Okay. So a mobile app, so Grinder is not the only mobile app doing that. Every single mobile app out there, if they decide to do fraud, can easily do that, right? You just fabricating the CTV bid request.
And then we've heard of examples where smart refrigerators are doing the same thing because they have a CPU. A appliances are doing this. Because all you need is like a CPU and internet connection and the smart refrigerators connected to the internet all day long, right? And then the last and probably most clever scheme, CTV fraud scheme, involved JavaScript in an ad slot.
So the JavaScript, once you get it into the ad slot, it can start fabricating CTV bid requests and sending that into the end point. So the reason I said it was clever is because it now has an automated way to disguise the IP address, right? So if a human were just sitting at their computer, they add loaded into their ad slot, it's going to be that person's IP address, right? So if you see like a whole bunch of CTV bid requests coming from a data center IP address, okay, something's fishing about that.
But now if you have millions of residential IP addresses or mobile IP addresses, because it's in the ad slot fabricating those CTV bid requests, it becomes a lot harder to catch. Okay. But all that being said, all we have to look at is the outcome, right? Where did my CTV ads run?
So I'm going to show you a few examples that we can actually see in Food Analytics where it's actually not CTV. The first and most obvious thing is when you put the Food Analytics CTV pixel in your CTV ad, if we detect that it went to a website or mobile app, you just got ripped off because that is not CTV. You're paying very high CPMs, you expect it to be on a large screen TV, but it ended up on a website or mobile app. Now the reason that happens is both fraud, but also even mainstream sellers, and I'm not going to name names right now, but even the big mainstream sellers have been tempted to do audience extension.
So remember the theme, I mean, we're talking about audience work, audience extension. Once they start mixing audience extension so that their numbers get much larger, meaning the quantities of impressions get much larger, the CPMs get blended down to be much lower. That's when you have these problems and you, the ad buyer are not getting ads on the big screen. Okay.
So that's one. The second would be wireless, where you know, your ad, if it ran on a big screen TV at home, it's usually in the living room connected to the cable modem. So it should be your cable modem IP address like Comcast, Time Warner, Spectrum, whatever, whatever. But if you see Verizon wireless, T-Mobile, AT&T, that's not running on a big screen TV, right?
It's on your wireless device like a phone or an iPad. Now in some cases, you might say, okay, I'm going to make an exception and I'm going to be okay with that. It might be Disney Plus and it's the kid on the iPad in the backseat. Hey, that's probably fine.
You can make that judgment call yourself. But if you see a large portion of your ads going to wireless IP addresses and or these crappy mobile apps, then you know you're getting ripped off because you're not getting ads on a big screen TV. So I'll kind of pause there and just say, you know, we're very strict about our definition of CTV. And it's also very easy and trivial for the bad guys to pretend to be CTVs by just fabricating those bid requests by the billions.
And that's why you need to be vigilant and measure so that you know what you're actually getting. Yeah. To me, this really gets into kind of the motivations or the tensions that exist with some of these because to your point, it's all about having the biggest audience network that you possibly can, right? Because what you're getting advertisers really is the illusion of spread, right?
You're going to get on all these different networks. And these are all the different ways your consumers consume content, right? But that doesn't necessarily mean you're getting what you pay for because you think you're paying for certain ad types, not necessarily certain placements, right? Those things should have a one to one relationship and very rarely do they actually have that.
So like on CTV, for instance, let's use LinkedIn as an example, something that we've used quite a bit at refine, you know that that brand safety list is quite long. And you're going to see multiple variations of the same channel necessarily. And some of them look like they're legitimately CTV or through a cable. And then some of them look like they're through apps or something like that, right?
So that's typically what you see when you look at these brand publisher list, these brand publisher list are extremely long for probably the amount of legitimate placements that there are. Yes. Yes. So if you're on LinkedIn, just look up the data from T vision.
So T vision is a company that has a little set up box that basically detects the sound and therefore they can detect which CTV ad ran. But based on their data, there's only like a dozen and a half. So 15 or so CTV channels that most humans watch. So if most humans are watching 15 to 20 CTV channels, right?
All the ones you've heard of like Disney plus Hulu, Paramount plus, whatever, peacock. What are the other tens of thousands of apps CTV apps that you've never heard of before? And those include wallpaper apps. Okay.
The TV is on your ads are running on a wallpaper app. Nobody's looking at the screen. And there's even more hilarious ones where there are these entertainment things for your pet, right? You leave the TV on and they're just showing entertainment to your pet and that's where your ads are going.
You think that's valuable. I mean, so some of these are just hilarious. And when you look at some of the bundle IDs, the problem is that whether it's a good seller or not, they typically hide a lot of these bad apps or fast channels and stuff like that in the bundle IDs. So you can't see it.
So again, that's why you have to actually measure it with an independent third party pixel, like the Vue and OXCD pixel. So you can actually see where the heck your ads went, because it's not going to be shown to you in the place reports and or log level data, because place reports and log level data record the site or app that was declared in the bid request. That could be different than the site or app that you're at ended up on. So same thing with CTV.
You have to differentiate what was declared in the bid request versus where your ad actually went. So you need your own independent measurement. You can't trust the reports from the platforms. Gotcha.
I want to move to DSPs. And kind of a lot of things to talk about are because CTV is programmatic really in theory, right? But DSPs, you mentioned to me when we were talking that this programmatic was really kind of the start of all of this, right? Like originally when we did when programmatic advertising kind of blew up and that's become its own enormous marketplace, right?
Trade desk, stack adapts. So there's all kinds of I don't mean to name names as if those are only two dollars. Everyone has the same symptom, right? So talk a little bit about like why you say programmatic is the start of this.
Like why that kind of became the advent of ad fraud really kind of ramping up. And yeah, and then sort of how do you control for that as you're looking at doing something like programmatic, which for a lot of companies is very attractive because again, going back to the illusion of spread and cost is I'm going to appear in front of a lot of people on a lot of different placements for a relatively low CPM. And so explain a little bit of that whole sort of marketplace tension that exists. Yeah, you know, like I started, you know, many, many years before programmatic came along.
And in the early days, you know, you would go to a Yahoo portal and that's where most people log in because whether sports scores, stock quotes, everything was there. So you would just negotiate a media buy with that publisher. And then over the years, you would go negotiate with Engadget, because Modo, some of these more niche blogs. But once programmatic came along, you kind of lost sight of who you were actually buying from, because you're no longer sitting across the table from a publisher or seller, right?
You're saying, here's a big chunk of budget and go find me as many sites as you can and run my ads on there, right? And also bid out every single ad impression. When the programmatic exchanges came along, for example, AppNex is being one of the early ones, it gave the bad guys an opportunity to onboard tens of thousands of fake websites. And at that time, they could generate those fake websites using WordPress templates and plagiarize content.
Okay, these days they're using AI to just make that process even faster. Okay, but once you have tens of thousands of fake websites in the exchange, that became a way for the bad guys to siphon lots and lots of dollars away from the advertiser. Because before the ad runs, you don't know exactly where it's going to run. And that's when the programmatic exchanges just open this Pandora's box for ad fraud.
And then started out with websites, then it went into mobile apps. Now there are tens of millions of mobile apps between the Google Play Store and the Apple Store. They're just running ads all day long. So again, like an alarm clock app is running ads in the background continuously, whether the app isn't used and whether device is on the nightstand or not.
That's a form of fraud that doesn't involve bots. But that's how they chew through a whole bunch of ad impressions. So that's kind of why I put the blame squarely on programmatic, because that was the first time when you really lost sight of who you were buying from, because you're no longer negotiating directly with the seller, the publisher. But long short short, okay, that's ancient history.
That's been done. What can advertisers do today to mitigate some of that? And you already mentioned it before, which is start using inclusion list, right? There's only a finite exclusion list, right?
In inclusion list, because there are infinite bad guys, and you can block them all, right? You don't have all day long to sit there and look at where your ad's going. So the best way to avoid all that work is to start with an inclusion list. And you should realize that there's a very small number of sites and apps that humans actually visit or even know about, right?
So I've used this question, the following question in my digital marketing class, I asked my students, name 10 websites that you use every single day, as fast as you can. They'll start rattling off a few when they get to five or six or seven. They'll slow down. They can't even name 10 websites that they use every single day.
Now, of course, there's going to be some long-till recipe site that you go look up one recipe on, and then you never visit again, right? Some of that happens, of course. So it's a random Google search. You're not intending to.
You don't remember that site. Same exact thing with mobile apps. I said, name 10 mobile apps you use every single day. They can't get to 10.
So then what are the other tens of millions of mobile apps that nobody's ever heard of? That's what makes up the majority of the fraud that your ads are ending up on. So you start with an inclusion list of a rational number of sites and apps that you've heard of before. That's a very good common sense way of gating that.
And then you still measure it, right? You still want to check. But like I said earlier, even the mainstream sites are not going to have zero bots. There's going to always be some red in there.
But at least you're avoiding the obvious fraud. And the reason you want to do that is we see these AI generated sites still. And these are, there are many, many of these new sites, recipe sites. I think the bad guys love recipe sites, because recipes don't have unbrand safe keywords like war or shooting or whatever.
So they're just generating all these recipe sites where all the pictures are AI, all the recipes are AI. And like we see in the data, this site is 33 days old, and it's selling millions of impressions. So those should have been prevented by whatever exchange, right? The SSP should have caught it, the DSP should have caught it.
But yet it's still showing up in the data because our ad ran on it. Right? So now again, once we can see that, we can just add it to a block list. But that kind of takes us back full circle to say, start with an inclusion list, because they're way, way too many bad guys to block.
You just won't be able to block them all. Right? Okay, that was great. And that's really the thing about when we talk about most of the things you can do to control Adfraud is really in your own hands.
It's just don't give too much power to the ad channel, because the ad channel will take a lot of liberties with your audience, with your placement and with your budget. If it means you pay a higher CPM, CPM is a price and not a cost we went over that yesterday. It's like, let me emphasize that. Let me emphasize that too many people talk about CPM, right?
The agencies keep talking about cost efficiency. And what they mean is like a lower CPM that misnomer has misled clients, advertisers for 10 years. So let me just reemphasize cost versus price. CPM is a price.
And you multiply that by the number of impressions you buy. Right? So let me use some simple math. 15 years ago, you bought ads at a $30 CPM.
These days you buy ads at $3, but you didn't save any money because you're buying 10 times the number of ads. So you're still spending the $30, but the price is $3. And because it's so low, you end up buying a lot of fraud, because mainstream sites like legitimate publishers can't sell you the inventory for $3. So that way, you just use that simple math to understand CPM is a price you did not save money, even if the CPM looks lower.
So now when you buy from better sites, your CPM is going to go up, but you don't need to buy that much. You can buy far less quantity. So you can still save money, because even if it's higher CPM, you're basically avoiding having to pay for impressions that were not shown to humans in the first place, all those fraudulent techniques that we talked about earlier. So for really the way to do this to control against it is inclusion list, inclusion list, inclusion list, do not use audience expansion, audience networks, avoid those like the plague, and be uber restricted if you're going to use sort of black box kind of products like PMAX.
PMAX, advantage plus, whatever. Yeah. And Metta would even touch on, but Metta is its own thing as well with advantage plus, turn that off, use the audience that you're actually targeting, don't let Metta take liberties with your audience, like their algorithm is very good. But if you give it too much liberty, because what these bots do, or what the algorithm does, is it feeds itself off of the conversion event, and so it just perpetuates fraud because it's basically the algorithm is showing them more people based on click data, or things like that.
It's showing them more bots over and over and over again. Yeah. So basically, here's the way I summarize it. It's not so much the conversion event, because that's hard.
The conversion mean the real sale. It's very hard for the algorithms to see that, because sometimes those happen months after the ad exposure. So they use something that's easy to measure, which are clicks. So the optimization algorithm, both PMAX, advantage plus, whatever black box, they use easy to detect signals like number of clicks and just assume that that correlates to performance or outcomes.
So now when you think about the bad guys, sites and apps, they always have a higher click through rate than real sites. Again, because humans don't click on ads that much, bots do. And bots are programmed to deliberately click on the ads so that the ones that ran on crappy websites and fake mobile apps appear to be performing better simply because you got more clicks. And so the algorithms will faithfully allocate more of your budget to the fake sites and fake apps.
So in fact, we've seen this happen to advertisers. More and more of their ads go to more and more fraudulent sites. So the small business owner that I helped years ago, he had been running his own Facebook campaigns for five years. At the beginning, he was seeing sales.
But then by the end of the fifth year, he was saying, all my sales went away, even though I've been getting many, many times more ad impressions and many, many times more clicks. So I asked him one question was Facebook audience turned on or off. It's on by default, right? Because Facebook wants to roll you in and you have to make you jump through all these hoops to turn it off.
You have to uncheck. Yeah, you have to uncheck. Find it and then uncheck the checkbox. So once he turned it off, his number of impressions cratered, his number of clicks cratered.
But then his actual sales started coming back. So it was all about that. So at the beginning of the five year period, 90% of his ads ran on Facebook. So it got in front of people.
By the end of the five year period, it had completely flipped. 90% of his ads were automatically sent to the audience that worked. Only 10% of his ads were running on the main Facebook app. So he didn't know that because it was all automated behind the scenes because you let it, right?
So as long as you uncheck the audience network checkbox, you can prevent most of that from happening. So to summarize, I mean, this might sound like a scary recommendation to a lot of people because the CPMs are going to go up, the clicks are going to go down. But slow is better. There's not this magical number of customers out there ready to buy yourself at any given time.
You just can't throw more money at it and expect the sales to grow. So slow and steady always is better. This is also why we've talking about this about small businesses versus large businesses. This is also why large businesses are the most apt and vulnerable for fraud because they have these incentives, right, to spend the budget, get the number of impressions possible.
And it's basically just forces them into buying a lot of fraud in order to justify their number. And I found that to be a fascinating tension that exists with larger advertisers. The very largest advertisers that I work with, like a CBG companies, they have so much budget. There's just not enough humans on Earth to generate enough impressions to use up all that budget.
But yet they have to spend it, right? In contrast, if you think about a small business owner, if they spend $100 on Facebook ads and they got nothing, right, they got zero additional sales on that, they can't spend the next $100 because they can't afford it. So small business owners tend to be a lot more vigilant because the dollars matter a lot more to them, even though it's a tiny, tiny amount, $100. When you have $100 million to spend, it's like, you tell your agency, go buy me as many ads as possible and get me the best price as possible.
And like you said, they basically incentivized the agency to go buy the crap because there's just not enough real quantity to buy from all the legitimate publishers. You have to make it in the crap. Yeah. Especially when you get the constraint of making the cheapest price possible.
Yes. That doesn't lend itself to getting in front of humans. I wanted to talk through some visual tellers on fraud because we went over this a couple. This is actually what we're going to look at who analytics now works a little bit because you talked about a couple examples.
I want to bring them up. Pixel stuffing, ads not rendering, CTV that's not CTV. There's some examples like that where even with all those controls in place, you can actually see it using your tool. And I'm sure those other tools I do similar kinds of things as well.
But I did want to bring this up just for people who want to see what this looks like and how you can tell. Show a little bit of visual tellers that I think I'm getting value, but I'm actually getting ripped off. Yeah. Let me start with something simple, which is pixel stuffing.
So normally an ad is supposed to run on a 300x250 ad unit. Those are pixel dimensions. But if you look at some of this data, the window size is 0x0. A human can't see that.
Literally it's invisible. So in this particular case, 44% of the ads went to a 0x0 window size. That's called pixel stuffing. So this campaign needed to be cleaned up.
In the second one, only 8.1% went to pixel stuffing. So it's already better, but still that's 8% of your impressions. It can't be seen. So you paid for them.
It's not valuable to you. And all we have to do is look at where this is coming from. I can tell you from experience, most of it are those crappy mobile apps like the flashlight app, or a lot of these kids games, like these tile things, the puzzle things, they, the normal way is to show an ad between levels, before you get to the next level, you've seen that. But that's just not enough volume for the apps to make money.
So what they do is while you're playing, continuously, they load as many ads in the background as possible into 0x0. So it doesn't interrupt your gameplay. So you don't actually, I mean, I know in practice here that you don't see it, but the way they deliver it even is what can't ask. And you have no prior of even seeing it.
You don't know it's even happening behind the scenes. But because your mobile device is continuously connected to the internet, it's perfectly convenient for the bad guy to keep doing that. So the next thing is as didn't render, this is actually a little hard to see. So I'm going to use a different slide to illustrate.
Okay, here. So in this case, all of these ad slots are blank. You see a little triangle right here? That's the ad choices icon, but the ad didn't load, right?
And this is because the programmatic bidding process takes so long that by the time you call the ad, and then it starts to download, right, it gets served from the ad server and starts to download into your device, it didn't make it in time to even be rendered on screen. So on the left hand side, all of these are blank. On the right hand side, I'm proud to say this is my client, George Pacific. You'll see almost every single ad loaded.
And these are small mobile banners, and these are display ads. So the file size are tiny. So if you compare that to video ads, the vast majority of video ads never rendered. Okay, let that sink in for a second.
Okay, so if a lot of these advertisers are buying video ads because they believe, oh, it's like sound emotion, it doesn't work if the ad didn't even render. Okay, so those are those are key things. What was your thing? So I want to talk about CTV.
That's not CTV because like, again, people think they're buying CTV, they're actually not and you can tell by looking at the URLs in the exactly. So when you use a CTV pixel from Food Analytics and then we detect or pan that cool math games dictionary dot com, these are not fraudulent sites. The problem here is that it was a CTV campaign. It should not have shown up on the sites.
And similarly, in the second one, it's all mobile apps, a Sudoku, audio, Mac, whatever, okay, so again, it's not CTV. And we're very strict about that. And then there was an example where it was a direct buy from a big TV platform. I won't name them.
But basically we thought it was going on big screen TVs. But once they put my pixel in, we found that 80% of the ads went to websites and mobile apps. So we went back to them and said, what the heck is going on? So long short, they were doing audience extension without disclosing it because they wanted more quantity and they wanted the CPMs to be lower.
So if anyone is buying CTV and you're getting it for $15, okay, you're not getting CTV ads, okay, it should be way, way higher CPM if it's on a legitimate publisher like Hulu, Disney Plus, Paramount Plus, Peacock, whatever. Got it. All right. That was fascinating.
We had one of the one you had, I wanted to go over with frequency capping because a lot of companies walk through and try to do frequency capping for their ads thinking that that's going to increase audience penetration and increase reach. And in fact, sometimes the ad platform just does not behave, right? Yeah, exactly. So yes, set your frequency caps, but no, don't just believe that it's happening.
You have to measure it yourself. So in fluid analytics, each fingerprint is basically an anonymous representation of a device browser combo. And you can see in this case, hundreds of impressions got shown to that fingerprint. So it all kind of depends on the time frame.
So you can say, you know, I want to look at a week, how many ad impressions shown were shown to a particular device. I want to look at a month, right? So you might have different frequency caps set, like maybe 30 per month or maybe one per day or whatever. So you have to kind of adjust the time frame in the fluid analytics platform to kind of see if this makes sense to the way you had set it up.
But as an example, when I ran my own campaigns, and again, a lot of this is me actually paying for campaigns with my own money and measuring it, because I'm a scientist, I need data. Right. So if it's not a client campaign, it's something I purchase myself. I set a lifetime frequency cap of one.
And then this is what I got. Okay, so that tells me that the DSP, I think they're trying, but I just don't think it's executed well. Right. So even though I had set the lifetime frequency cap of one, I'm still getting hundreds, if not, you know, multiple dozens of impressions shown to the same device.
Okay. So this is based on my data, the reports you get from the platform, you're going to tell you everything's fine. Don't worry about it. Right.
But again, even if it's not malicious, they sometimes don't know. And let me be very specific. I'll use the log level data example, because someone asked me about it this week, log level data and the police reports record the domain or app that was passed in the bid request. It doesn't mean that's where the ad actually went.
So just be aware, you should ask for log level data, but it's a ton of data is really hard to handle. And even then you're not going to be able to see all the bad guys sites in there, because when a fake site pretends to be CNN.com, it's going to be recorded under that row under CNN. It's just going to get mixed into the CNN row. So you can't see it.
That's why you have to run a post bid JavaScript tag or post bid pixel for CTV to actually detect where the ad went. So if you have to analytics in place, you don't have to process a lot of log level data, because even if you did and had a data scientist know how to do that, you still can't get it the truth because the log level data contains errors. Fascinating. We can we can stop sharing it real quick.
I do want to give a couple minutes, but I do want to open up. If anyone has any questions, I would definitely welcome you to bring those forward is that I thought this was incredibly interesting, just to kind of see the level of that overall. If not, like I have a couple I could certainly ask as well, and we only have a minute left, but anyway, does have a question like dropping in the comments or raise your hand. I'll bring you on live.
Okay. Maria, I believe it's Maria or Maria has her hand raised. We can bring staff if you want can come live and ask the question if that's okay. Yeah.
What's make that happen? Hi, I'm Maria. I've been following you on LinkedIn for a while, and just found this presentation really interesting. So I just wondering if one is this going to be this recording?
Is it going to be available to us? Yeah, we'll definitely get a level on YouTube for sure. And then also, is there like a sort of like tips or checklist or something available on food analytics that would help us, like people just starting to investigate this more deeply? It's not pre-made, but I can just email me afterwards or connect with me on LinkedIn.
I can send you some of these slides where the decks we can see how to start using food analytics. Awesome. Thank you. Any other questions?
Anyone else want to bring a use case up maybe? If not, that's totally cool. Because we picked time. I do have to go ahead and yeah.
All right. That's okay. All right. This was this was fantastic.
Thank you. Listening to this will get the recording out. I know stuff will clip a lot of this also. So you'll find this stuff on LinkedIn as well getting sort of playing fed out there.
But Dr. Fu, I want to thank you so much for your time. Thank you. It's got all these rounds of applause you're getting.
All right. So I can with your material. Good to see everybody. We appreciate it.
So much to connecting with you on LinkedIn. Thank you, Matt. Thank you, Stephanie, for putting on a great set program. All right.
Thank you. Have a great rest of the week. Everyone appreciate it. Yeah.
Bye.