Buying UEBA: Questions That Matter
Every product demonstrates well. These questions surface what happens in month six and what your team can do with the output.
Evaluations in this category are unusually easy for vendors to win, because the demonstration environment is clean and the interesting events are already present. These questions are the ones whose answers differ between products.
On the analytics
Show me an entity's baseline. What does the model consider normal for this person? If you cannot see it, you cannot explain a finding or defend it. This is the single most revealing question in the list.
Show me the per-feature contributions for this score.
How is the score computed? Vague answers here predict opacity throughout.
Which features are used, and can we change them? A fixed feature set chosen for a reference environment is a significant constraint.
How is correlation between features handled in scoring? One action producing five correlated deviations should not inflate a score fivefold. The quality of this answer indicates how carefully the model was built.
What is the baseline window, and is it configurable per entity type? Service accounts and humans need different windows.
What happens on a role change? The correct answer involves a trigger and a reset, not "the rolling window adapts eventually".
How do you handle shared accounts? The honest answer is that they cannot be modelled well. A vendor claiming otherwise is not being straight.
How do you handle entities with no history? Peer comparison, and whether baseline maturity is exposed to the analyst.
On the data
Which of our sources can you consume, and with what effort? Get specifics per source, not a logo slide.
Do you do entity resolution, or do we? And if you do, how, and can we correct it?
What happens when a source stops sending? If the answer is not "we alert on absence", that gap is yours to fill.
What schema do you use? OCSF or ECS support means portability.
Can we export the normalised data and the model outputs? Ask before signing.
On operations
What alert volume should we expect at a reviewable threshold, for our size? Then ask for a reference customer of similar scale and ask them the same question.
How long from install to tuned output? Multiply the answer by two.
Who operates this day to day, and for how many hours? Detection engineering, review, exclusions, drift. This exceeds the licence cost over three years in most deployments.
How are exclusions created, documented, reviewed and expired?
How is adjudication captured, and can we export the labels with the feature vectors? Without this you cannot measure precision or improve anything.
How often is the model retrained, and are we notified? Managed retraining that shifts your score distribution without warning is a real operational problem.
On governance
What personal data is processed and where? Needed for the impact assessment.
Can we restrict what our own analysts see? Metadata review with content access as a separate, logged permission.
Is access to individual data audited, and can that audit go to a different team?
Can we exclude protected communication channels — legal, occupational health, whistleblowing, employee representatives — from profiling?
What do you retain, for how long, and can we set it?
What can we not turn off? Some products bundle capabilities you do not want and cannot disable.
What to weight lightly
Detection rate claims. Measured against the vendor's own simulated data.
Machine learning as a differentiator without an explanation of inputs, outputs and behaviour under role change.
Feature count. The longest lists are frequently the hardest to operate.
Analyst positioning. Reflects market presence more than fit.
Before the first demonstration
Write down the three entity populations you most need covered, the sources you can realistically deliver, and the review capacity you actually have. Evaluate against that document.
Vendors will expand your requirements during the process — it is their job — and the written scope from before the first call is what keeps the evaluation honest.
Structuring the commercial side
The technical evaluation is only half of it, and the commercial terms determine what happens when the deployment disappoints.
Price the whole cost, not the licence. Integration hours, storage, operating time, and the analyst you will need. Over three years the licence is frequently the smaller number.
Get the volume estimate in writing. Expected alert volume at a reviewable threshold, for your size and sources. If the vendor will not commit to one, that is informative.
Negotiate an exit. Data export in a usable format, including normalised events, baselines where possible, and adjudication labels with feature vectors.
Tie payment to milestones that mean something: entity resolution above a threshold, reviewable volume achieved, first tuned quarter completed.
Ask about retraining notification and put it in the contract if the answer matters to you.
Get a reference customer of similar size and sector, and ask them the operating questions rather than the capability ones.
Common false positives
Evaluation signals that mislead in predictable directions:
An impressive demonstration, which reflects prepared data.
A long feature list, which frequently correlates with operational difficulty rather than capability.
Analyst rankings, which measure market presence.
A successful small pilot, where effort scaled linearly and will not.
Enthusiastic engineering during evaluation, which is a pre-sales resource that disappears at signature.
Case studies, which describe the deployments that went well.
The counterweight to all of these is the written scope from before the first call, and the operating questions put to a reference customer who has lived with the product for two years.
Blind spots and assumptions
That the demonstration data resembles yours. It was chosen not to.
That integration effort scales linearly. It does not, particularly with entity resolution.
That you can switch later. Baselines, tuning and exclusions do not port.
That the module in your existing suite is equivalent. It may consume only that vendor's telemetry, which is a much narrower deployment than it appears.