Why self-run AI employer branding checks mislead: ten methodological problems
Asking an AI engine about your own organisation feels like evidence. It has properties that make its output unreliable in specific, predictable ways.
01 · The instinct is reasonable. The output is not.
It is a reasonable first instinct. If you want to know how your organisation is described when candidates or customers ask an AI engine, ask it yourself. Open a chat, type the question, read the answer. It takes two minutes and it feels like evidence.
It is not, and the reasons matter, because decisions are being made off these informal checks. The problem is not effort or intent. It is that the exercise has properties that make its output unreliable in specific, predictable ways. Ten of them are worth understanding before you act on anything a self-run check tells you.
02 · 1. One query is one sample, not a measurement
AI responses are not deterministic. Ask the same question twice and the answer will vary, sometimes in which organisations are named, often in the language attached to them. That variance is a feature of how the systems work, not a glitch.
A single query, then, is one draw from a distribution. It tells you something is possible, not what is typical. Treating one answer as your position is like judging a market by asking one customer. Establishing what an engine typically returns requires repeated sampling, enough runs to separate the stable pattern from the noise, and a discipline for aggregating them.
03 · 2. You are the worst-placed person to ask
Engines increasingly personalise. Your location, your account history, your previous conversations, and the fact that you have been reading about your own organisation all shape what comes back.
So the person most motivated to run the check is the person whose result is most contaminated by their own footprint. You are not seeing what a candidate in another city, with no history of asking about you, receives. You are seeing a version conditioned on being you. That is the one perspective the exercise cannot afford.
04 · 3. The prompt determines the answer
"Tell me about working at X" and "which employers in this sector treat people well" are not variations on a question. They are different instruments, and they produce results that cannot be compared.
Without a fixed, documented prompt set, every check is measuring something slightly different. Movement between checks is then uninterpretable, because you cannot tell whether the market shifted or your phrasing did. Consistency of instrument is the precondition for any comparison over time, and informal checks almost never have it.
05 · 4. Prompts written by insiders ask insider questions
There is a subtler version of the same problem. People inside an organisation phrase questions in the organisation's own language, using its terminology, its framing of its own strengths, and its sense of who its competitors are.
Candidates and customers do not. They ask broader, vaguer, more comparative questions, and they ask about things you may not consider your category at all. A prompt set built from the inside quietly tests the market you believe you are in, rather than the one people are actually asking about.
06 · 5. A result without a competitive set cannot be read
Suppose a check returns a warm description of your safety record. Good news, or meaningless?
It depends entirely on whether every competitor is described the same way. Position is comparative by definition. A characterisation shared by the whole market is table stakes, and a characterisation unique to you is an advantage, and the two look identical when you only look at yourself. Self-run checks are almost always self-focused, which strips out the only context that makes the observation actionable.
07 · 6. A snapshot cannot show a trend
The risk in this area is directional. An account of you firming up, a competitor gaining ground on a dimension you own, a strength quietly dropping out of how you are described. All of those are movements, and none of them are visible in a single look.
Because responses vary run to run, the problem compounds. Two checks a month apart may differ entirely because of sampling variance, and you have no way to tell that from a real shift. Distinguishing signal from noise requires a consistent baseline and repeated measurement over time, which ad hoc checking cannot produce by construction.
08 · 7. One engine is not the market
Different engines return different answers, drawing on different sources and weighting them differently. People check the one they personally use.
That gives you a view of a slice of the market, presented as the whole. An organisation can surface strongly in one engine and barely register in another, and the informal check will report whichever happens to be open in the browser.
09 · 8. Reading answers is not measuring framing
The most common self-run method is to read a few responses and form an impression. That is qualitative interpretation, and it does not produce a number you can track.
Measuring how you are framed requires classifying language consistently, against defined dimensions, applying the same criteria every time and across every competitor. Without that, "it seemed fairly positive" is doing the work, and it will be applied more generously to your own organisation than to your competitors, without anyone intending it.
10 · 9. Motivated reading is unavoidable, not a character flaw
Nobody runs a check on their own organisation neutrally. You know what you hope to see, you know which claims you have invested in, and you have a professional stake in the answer.
That shows up in ordinary ways. Ambiguous phrasing reads more favourably. An unflattering result feels like a fluke worth re-running, while a flattering one feels like confirmation. Absence of a strength gets attributed to the engine rather than the evidence. None of this requires bad faith, it is how interpretation works when you have something at stake, and it is precisely why measurement of your own position should not depend on your own judgment.
11 · 10. It is a monitoring problem dressed as a project
The final issue is durability. A check answers a question once, and the thing it measures keeps moving.
Treated as a project, this becomes an activity someone does when they remember, with a different prompt, in a different engine, on a different day, producing results that cannot be compared to the last attempt. The effort recurs indefinitely and never compounds into a picture. What makes the data valuable is continuity, and continuity is the first thing informal effort loses.
12 · What a defensible approach requires
The pattern across all ten is the same. What feels like the simple version of the exercise is missing the properties that make measurement mean anything.
A defensible read on how you surface needs repeated sampling to handle variance, a neutral vantage point free of your own footprint, a fixed and externally phrased prompt set, the competitive set measured alongside you, coverage across engines, consistent classification of language rather than impression, and continuity over time so movement is legible.
That is a methodology, not a task. The distinction matters because the informal version does not produce a weaker version of the same answer. It produces a different kind of output, an anecdote, which is fine for curiosity and unfit for decisions that carry budget.
13 · The point in one line
The problem with checking it yourself is not that it is hard. It is that it is easy, and the result looks enough like data to act on, while lacking every property that would make acting on it wise.
