Skip to content
Behavioural Analytics Review

Index  ·  Governance

Bias in Behavioural Models

A model scoring deviation from the majority will score the minority higher, permanently, for reasons unrelated to security risk.

Analysis  ·  Needs: Model output, HR

Behavioural analytics identifies people who differ from a norm. Some people differ from the norm for reasons that have nothing to do with security and everything to do with their circumstances.

The system cannot tell the difference. Unless someone looks, it will flag the same individuals indefinitely.

Where it enters

Working pattern features. Hour of day is the clearest case. Someone working evenings for childcare, someone in a different timezone, someone observing a different working week, a shift worker, someone whose disability affects when they work. All differ from a global baseline on a feature that carries almost no security signal for a distributed workforce.

Location features. Country, connection origin, travel patterns. These correlate with nationality, with family abroad, with immigration status.

Peer comparison. Deviating from a group is the mechanism. A person who is the only one of something in their group — the only remote worker, the only part-time member, the only one in that region — deviates structurally.

Language and communication features. Volume and pattern of written communication vary with first language and with role.

Device and technology features. Older hardware, assistive technology, non-standard configurations. These correlate with seniority, with accommodation, and sometimes with pay.

Training data. If historical investigations concentrated on a population, any model learning from adjudications reproduces that concentration.

Why it is not fixed by removing protected attributes

The obvious response — do not feed the model gender, ethnicity, age or disability — is necessary and insufficient.

Behavioural features are proxies. Working hours proxy for caring responsibilities. Location proxies for nationality. Device configuration proxies for accommodation. A model with no protected attributes can still produce output that correlates strongly with them, through features that look purely technical.

Removing the attribute removes your ability to measure the disparity while leaving the disparity in place.

What to actually do

Measure the distribution of elevated scores. Which individuals appear repeatedly? Which populations are over-represented relative to their share of the workforce?

This requires holding demographic data for the analysis, which is a tension worth resolving deliberately with your privacy function rather than avoiding. Aggregate analysis by a restricted group, with results reported as distributions rather than individuals, is usually the workable arrangement.

Look at repeat individuals specifically. Someone in the top decile every week for six months is almost certainly there for structural reasons. That is a finding about the model.

Audit features for proxy behaviour. For each feature, ask what it is really measuring. Hour of day is the one to examine first.

Set expectations for peer group minimums. Small groups make anyone unusual make everyone an outlier, and minorities are structurally more likely to be in small groups.

Give analysts the ability to record structural explanations persistently, so the same benign anomaly is not re-investigated monthly.

The remote work example

Worth stating because it happened at scale.

Models trained before 2020 encoded office-hours, office-location behaviour as normal. When the workforce moved home, everyone deviated. Organisations that retrained absorbed the change; organisations that did not spent a year flagging their entire staff.

The subtler version persists: in a hybrid workforce, the people who work remotely most consistently deviate most from a baseline dominated by office attendance. That is a bias against a population, and the population correlates with caring responsibilities and disability.

Governance

Someone outside the security team should review the distribution. Annual is adequate. The security team is not well placed to audit itself here.

Record the review. Whether a disparity was found, what was concluded, what changed.

Treat a persistent disparity as a defect with an owner and a date, not as an observation.

Where a feature is both biased and weakly predictive, remove it. Hour of day for distributed workforces is usually both.

Running a disparity review

An annual exercise, conducted by someone outside the security team, producing a short written finding.

Take a year of elevated scores. Aggregate by the populations you can lawfully analyse — location, working pattern, employment type, and protected characteristics where your privacy function agrees the analysis is appropriate.

Compare representation. Is any population over-represented among elevated scores relative to its share of the workforce?

Identify repeat individuals. Anyone appearing in the top decile persistently over months. This list is usually short and usually explicable by circumstance rather than behaviour.

Trace the features responsible. For the over-represented groups, which features drive their scores? The answer is frequently one temporal or location feature.

Decide and record. Remove the feature, adjust the grouping, exclude the population from that detector, or document why the disparity is justified.

Report the finding as distributions, never as individuals, and keep the analysis separate from operational access.

Common false positives

Structural, recurring flags that reflect circumstance rather than risk:

Non-standard working hours for caring responsibilities, disability, religious observance or shift patterns.

Cross-timezone work, where normal local hours look nocturnal centrally.

Connections from abroad for staff with family overseas or on long visits.

Assistive technology and non-standard configurations producing unusual endpoint telemetry.

Part-time patterns, which differ from a full-time peer baseline on nearly every volume feature.

Sole specialists, who deviate from any group they are placed in.

Blind spots and assumptions

That the model is neutral because it is mathematical. It reproduces whatever the data contains.

That nobody will notice. People notice being questioned repeatedly, they discuss it, and the pattern becomes visible to everyone except the team producing it.

That this is a compliance concern only. A model that flags the same innocent people continuously is also a bad detection system, because it wastes the review capacity that is the binding constraint.

That fairness and accuracy conflict. Here they usually align: features that are biased are frequently the ones carrying least security signal.