Skip to content
Behavioural Analytics Review

Index  ·  Context

Why UEBA Deployments Fail

The failure modes are consistent and most are decided before the product is installed. Read backwards, they give the sequence that works.

Analysis

A large share of UEBA deployments end as an expensive log: running, licensed, producing scores nobody reviews. The causes recur.

Bought before the data was ready

The dominant cause. A product is purchased, connected to whatever logs exist, and expected to produce findings.

But the analytics sit on top of entity resolution, normalisation and coverage — none of which the product supplies. Connected to fragmented identities and inconsistent schemas, it produces confident nonsense.

The tell: nobody can say what percentage of log identifiers resolve to a known entity, six months in.

No review capacity

The licence was budgeted; the analyst was not.

A queue nobody reviews is not detection. This is the most predictable failure in the field and the easiest to foresee: ask who will look at the output daily, by name, before signing.

Thresholds raised until quiet

Volume becomes unbearable, so thresholds rise. Noise falls, detection falls with it, and the system now surfaces only the most extreme activity — which was mostly already covered by rules.

The alternative — fixing data, excluding approved workflows, adding context, correlating — takes a quarter and preserves both.

Deployed across everything at once

Every source, every population, simultaneously. The result is unattributable noise: nobody can say which source or which detector is generating the volume, so nothing can be tuned.

Nobody owns it after go-live

The project team disperses. Baselines age, sources break silently, exclusions accumulate, models drift. Two years later the deployment reflects an environment that no longer exists.

Behavioural analytics is a maintained system, not an installed one, and maintenance is what gets cut.

The score was treated as a finding

An analyst escalates on a number. HR acts. The person's representative asks what 87 means. Nobody can answer.

One such case sets a programme back years, because after it HR stops accepting referrals.

Governance settled after the first case

Who could look, who approved, what was retained — none written down. The first serious investigation exposes it, and the exposure is worse than the incident.

Legal involved after purchase

Discovering a works council requirement, an assessment obligation or an unlawful monitoring configuration after deployment means unwinding it or negotiating from weakness.

Baselined over a compromise

The training period contained an attacker. Their behaviour is now normal and will never be flagged. Nothing in the output reveals this, and the only mitigations — a hunt before baselining, comparison of baselines against peers — are cheap and rarely done.

Measured by alert count

The number falls during tuning, which is success, and reads as declining value. Or it rises with coverage and reads as deterioration. Either way it drives decisions in the wrong direction.

The sequence that works

Reading the failures backwards:

Fix entity resolution and measure it. Normalise to a published schema and test the parsers. Ingest one source, verify it, then add the next. Run silent for a full baseline window. Hunt before you trust the baseline. Enumerate approved workflows and exclude them explicitly. Add enrichment before adjusting thresholds. Set the threshold from review capacity. Capture adjudication from day one. Agree governance with legal and HR before the first case. Tell employees. Name an owner for the year after go-live. Report process metrics, not alert counts.

None of that is technically difficult. All of it takes longer than a procurement timeline assumes, which is why it usually does not happen.

A pre-purchase readiness check

Ten questions. A programme that cannot answer most of them is not ready to buy, and buying anyway is how the failures above happen.

What percentage of our log identifiers resolve to a known entity? If unknown, that is the first project.

Which entity populations do we most need covered? Named, in priority order.

Which sources can we actually deliver, at what retention?

Who will review the output daily? By name, with allocated hours.

What is our realistic review capacity? Measured, not estimated.

Who owns this in year two?

What is our lawful basis, and have we consulted where required?

What does our escalation path to HR and legal look like?

How will we capture adjudication?

How will we know if it is working?

Answering these takes a week and reliably changes either the timeline or the scope.

Common false positives

Signals that a deployment is failing, which are easy to misread as success:

A quiet queue, which may be good tuning or a dead data source.

Falling alert counts, which read as improvement and may be narrowing coverage.

High analyst throughput, which may mean fast dismissal rather than efficient triage.

No complaints from the business, which may mean nobody has been contacted about anything.

Renewal without discussion, which measures inertia rather than value.

A clean audit, which examines whether the process was followed rather than whether it detects anything.

Distinguishing these from genuine health requires the process metrics, which is why alert counts alone are worse than no reporting.

Blind spots and assumptions

That failure is visible. A failed deployment looks identical to a quiet one.

That the vendor will say. Their measure of success is renewal.

That a restart is cheap. Rebaselining, retuning and re-establishing credibility with HR each take a quarter.

That the next product will be different. Almost every failure listed above is environmental rather than product-specific.