Feedback Loops That Improve Detection
Most deployments are static after month three. The mechanisms that let one improve are cheap and depend on capturing adjudications.
A UEBA deployment either improves or decays; it does not hold steady. The difference is whether analyst judgement flows back into the system, and most deployments discard it entirely.
The loops worth building
Adjudication to precision measurement. Every closure categorised, stored with the feature vector. This is the foundation — without it, no other loop is possible and no quality claim is defensible.
Adjudication to detector retirement. Per-detector precision over a rolling window. Detectors with no true positives in six months are consuming attention and providing coverage on paper only.
Adjudication to feature selection. Which features appear in confirmed findings versus dismissals. Directly informs what to keep.
Exclusions to process improvement. Every exclusion written for an approved workflow is a statement that a business process looks like an attack. A rising count in one area is worth investigating as a process question, not just a tuning task.
Investigations to detection engineering. After any real case, ask what would have caught it earlier and build that. This is the highest-value loop and it requires someone to own it after the case closes, when everyone is tired.
Red team to coverage measurement. An authorised exercise produces ground truth. What was surfaced, what was not, and how long it took.
Making adjudication capture work
Small fixed category set. Five or six. Free text does not aggregate.
Mandatory at closure. Optional fields are empty within a month.
Retain the feature vector. The label is worthless without what the model saw. Products routinely store the outcome and discard the input, which forecloses every downstream loop.
Make it fast. If categorising adds thirty seconds to every triage, it will be done carelessly. Two clicks.
The trap: training on your own dismissals
Learning from analyst decisions creates a self-confirming loop.
Analysts dismiss a category. The model deprioritises it. Those alerts stop appearing. Nobody ever revisits the judgement, because there is nothing to revisit.
If the original dismissals were wrong, that error is now permanent and invisible. The system has learned your blind spots and made them structural.
The countermeasure is a random sample that bypasses all learned prioritisation and is reviewed regardless of score. Five percent of capacity, and it is the only reliable way to discover what the ordering is hiding.
Also: re-examine a sample of dismissed alerts quarterly with fresh eyes. Use learned models to reorder, never to suppress. Watch for any category whose volume has fallen to zero — solved and invisible look identical from the inside.
Calibration between analysts
Two analysts will adjudicate the same alert differently. Without correction, the labels are noise.
Run a calibration session quarterly. Ten real alerts, everyone adjudicates independently, then compare. Disagreement identifies where criteria are unclear.
Track dismissal rates per analyst. Large divergence is a criteria problem, not a performance problem, and treating it as the latter guarantees nobody reports it again.
The review that closes the loop
Quarterly, an hour:
Which detectors produced findings, and which produced nothing? What did the random sample surface that the queue missed? Which exclusions expired and were renewed without examination? What did the last investigation reveal that is not yet detected? Has precision moved, and in which direction?
Most deployments never hold this meeting, which is why they look the same in year three as in month three.
The quarterly review agenda
One hour, four people, six questions. This is the meeting that separates a deployment that improves from one that decays.
Which detectors produced confirmed findings this quarter, and which produced none? Retire or rebuild the second group.
What did the random sample surface that the main queue did not? If the answer is anything, the ordering needs work.
Which exclusions expired, and were they renewed deliberately or by default?
What did our last real case reveal that we still cannot detect?
Has precision moved, and why?
What changed in the environment that we have not accounted for? Migrations, reorganisations, new systems, new populations.
Minute the answers and the actions. The value compounds: by the fourth meeting the pattern of what works in your environment is visible in a way it never is alert by alert.
Common false positives
Feedback mechanisms introduce their own errors, worth watching for:
Suppression learned from rushed dismissals, entrenching decisions made under queue pressure.
Categories collapsing as analysts default to the first option in the list.
Precision improving because coverage narrowed, which reads as progress and is not.
Detectors retired for producing nothing, when the cause was a broken data source rather than an absence of activity.
Exclusions renewed without examination, accumulating into gaps nobody agreed to.
Each of these is invisible without the random sample and the quarterly look.
Blind spots and assumptions
That analyst judgement is ground truth. It is the best available, and dismissals contain missed incidents.
That the loop is automatic. Every one of these requires someone to own it, and ownership is what gets cut when the team is busy.
That improvement is visible. These loops raise precision slowly. Without measurement, the work looks like maintenance and gets deprioritised.
That a vendor's learning does this for you. Vendor models trained across customers do not encode your environment, your workflows or your adjudications.
More in this section
Needs the same data