Running a UEBA Proof of Concept
Most evaluations are demonstrations on vendor data with a predetermined outcome. How to run one that is capable of failing.
Every product in this category demonstrates impressively, because the demonstration uses prepared data in a clean environment with the interesting events already present. A proof of concept is worth running only if it is structured to be capable of failing.
Decide the question first
Write down, before any vendor contact, what you are trying to establish. Not "evaluate UEBA solutions" — something answerable.
Examples that work: can this surface a compromised credential scenario in our environment within 24 hours? Can we reduce our privileged account review burden while maintaining coverage? Can our team operate this with the staffing we actually have?
A question with a yes or no answer, decided by evidence, agreed before the vendor is engaged.
Structure
Use your data. Not a sample, not synthetic. Real logs from your environment, at real volume, with your entity resolution problems intact. Vendors who cannot do this in a reasonable time are telling you something about deployment effort.
Run long enough to baseline. Thirty days minimum, sixty is better. A two-week evaluation measures the cold-start experience, which every product fails.
Include your worst data. The source with inconsistent timestamps, the population with shared accounts, the systems with partial coverage. That is your actual environment, and how the product handles it is the finding.
Have your analysts triage. Not the vendor's engineers. The question is whether your team can interpret the output, and vendor-assisted triage answers a different question.
Run a red team exercise during the window. The single most informative thing you can do. Authorised, scoped, with the security team unaware of the timing. Whether the product surfaced it, and how quickly, is worth more than every other measure combined.
What to measure
Volume at a usable threshold. Not the raw count — the count at a threshold your team could actually review.
Precision on that volume. Your analysts adjudicate; you count.
Time to triage. How long does one alert take with the context the product provides. This determines operating cost more than anything else.
Coverage. Which of your entity populations got usable baselines, and which did not.
Explainability. Can an analyst state why an entity scored highly, in a sentence, without vendor help? Test this explicitly on ten alerts.
Effort to integrate. Track your own hours honestly. This is the number most often underestimated and it usually exceeds the licence in year one.
Effort to tune. How much work to get from initial output to reviewable volume.
Questions to put to the vendor
Show me an entity's baseline. If you cannot see what the model considers normal for a person, you cannot explain a finding or defend it.
Show me the per-feature contributions for a score.
What happens on a role change?
How do you handle shared accounts? An honest answer is that they cannot be modelled well.
How is the score computed? Vague answers here predict opacity everywhere.
What personal data is processed and where? You will need this for the impact assessment.
Can we restrict what our own analysts see? Metadata review with content access as a separate logged permission.
What is the retraining cadence and are we notified?
What a failed evaluation looks like
Worth recognising in advance:
The vendor drives the triage. The evaluation runs two weeks. It uses a subset of clean data. Nobody records integration hours. Success is declared because the product surfaced something interesting, without asking whether the queue would be reviewable at scale, or whether anything was missed.
That evaluation always succeeds, which is why it is worthless.
A scoring sheet
Decide the weights before the first demonstration, so that the evaluation cannot be steered afterwards.
Coverage of your priority populations — do the entities you most need modelled actually get usable baselines.
Reviewable volume at a threshold your team can sustain.
Precision on that volume, adjudicated by your analysts.
Explainability, tested on ten alerts: can an analyst state the reason in a sentence.
Integration effort, measured in your hours, not the vendor's estimate.
Operating effort, estimated from the tuning work the evaluation required.
Governance fit — access controls, audit, retention, exclusions for protected channels.
Weight these according to your constraints, sum them, and hold to the result. Evaluations drift toward whichever product demonstrated most impressively unless the criteria were fixed in advance.
Common false positives
Evaluation artefacts that make a product look better or worse than it is:
A short window, which measures the cold-start experience rather than steady state.
Clean data, where the messy sources were excluded to make integration quicker.
Vendor-assisted triage, which answers whether their engineers can interpret the output.
Attention inflation, where the team engages with the evaluation queue in a way they will not sustain.
Ordering effects, where the second product benefits from everything learned during the first.
A quiet period, where nothing interesting happened during the evaluation and the absence of findings is read as either good precision or poor recall depending on preference.
Blind spots and assumptions
That the evaluation environment resembles production. Effort scales non-linearly with sources and entity count.
That the demonstration data reflects yours. It reflects an environment chosen to look good.
That your team's capacity during a POC reflects normal. People give an evaluation attention they will not sustain.
That comparing two products is fair. Whichever ran second benefits from what you learned during the first.
More in this section