Fighting & Violence Detection: Public Safety Alerts
Most violence does not begin with violence. It begins with a conversation that gets louder, a step that closes the distance. Security teams describe the same window: twenty to ninety seconds between the first visible escalation and the first contact, and almost nobody watching a wall of cameras at that moment. By the time anyone reviews the footage, it is an incident report instead of an intervention.
Fighting detection is built for that window. It reads motion and posture on your existing streams and alerts while the situation develops, so a guard can walk in before it becomes an assault. It does not identify anyone. It reports that an altercation appears to be happening, at a named camera, at a named time.
What the Model Actually Reads
The rule is not looking for a face, a weapon or a uniform. It looks for a motion signature that is hard to produce by accident: tracked people converging and staying in contact, with high-frequency limb movement and a rapid shift in posture, upright to crouched or horizontal. Underneath is generic object detection extended by temporal analysis.
Three parameters decide whether the rule stays useful or gets muted in week two. Sensitivity sets how violent the motion must look: too low and a handshake trips it, too high and a shove does not. Dwell requires the pattern to persist for several seconds, the best filter against sport and horseplay. Region limits it to where an altercation would matter. Tuning it against real footage is most of the work, and the part generic video analytics packages leave to the buyer.
Where Escalation Happens First
The deployments that justify themselves fastest are rarely the ones people expect. In education, corridors and stairwells between lessons are blind spots by design. In retail, the queue line and returns desk generate more aggression per square metre than anywhere else. In smart city work, transport interchanges, night-bus stops and underpasses carry late-night risk no patrol schedule covers evenly. In logistics hubs, shift handovers and gate queues are the recurring flashpoints.
All four share the same property: the camera is already there, pointed at the right place, recording nothing useful until reviewed afterwards.
Pairing Violence Detection With Neighbouring Rules
A fight is rarely the first thing that happens. Stacking rules turns one alert into a sequence you can act on earlier: crowd gathering detection catches the group forming before anything turns physical, loitering detection flags the person waiting in the wrong place for ten minutes, and intrusion detection covers the yard that should have been empty after hours.
After the event, two more rules matter for the report. Fall detection establishes whether someone went down and stayed down, which changes the medical response as much as the security one, with dedicated work on elderly fall detection for care settings. Abandoned object detection catches the bag left in a scuffle. On integrity, camera tampering detection and camera displacement detection tell you when a camera was turned or covered, often why an incident has no usable footage.
The Honest Boundaries
Three limits are worth stating before you buy. The rule reads motion and posture, not identity, so it will not tell you who was involved. It degrades with occlusion: a fight inside a dense crowd or behind a pillar is genuinely hard, because the model cannot see the pattern it looks for. It is an alerting tool, not evidence, it gets a person to the right place in time and what follows is a human decision. Low-light performance depends on your cameras, since the appliance analyses the stream it is given.
For the background on how this works, see what is fall detection.
Frequently Asked Questions
How high is the false alarm rate?
It depends almost entirely on dwell and sensitivity tuning rather than on the model. A rule set to trigger on one second of abrupt motion will fire on horseplay, sport and busy counters. Requiring the pattern to persist for several seconds, and restricting the region to places where an altercation is plausible, removes most of it. We tune against your own footage rather than shipping a default and hoping.
Can it trigger a crowd gathering alert as well?
Yes, and it usually should. A common configuration runs both rules on the same camera: crowd gathering raises an early, lower-priority notice when a group forms and stays, and fighting detection escalates to an immediate alert if the motion signature turns violent. You can route the two to different recipients, so the early one does not train people to ignore the urgent one.
Does it identify the people involved?
No. Fighting detection classifies an action, not a person. It reports that a physical altercation appears to be occurring at a specific camera and time, and it does not attempt to recognise faces or attach identities. Sites that need identity-based access or watchlist matching handle that with separate rules and separate retention policies.
Do we need to replace our cameras?
No. The rule runs on your existing RTSP or ONVIF streams on an appliance inside your network, from 2 to 128 channels and 1 to 256 TOPS, with no per-camera licence and no video leaving your site. It works offline, which matters for remote sites with unreliable backhaul.
Tell Us Your Scenario
Every site tunes this rule differently, and the gap between an alert that gets acted on and one that gets muted is thirty minutes of configuration against real footage. Send us your camera layout and the behaviour you need to catch, and we will tell you which of the 198+ algorithms fit and what to expect on your own video. Tell Us Your Scenario, or start with a Starter Kit preloaded with tested configurations and running on your own cameras in a single afternoon. Where the detection you need already exists in the algorithm library, the first demo takes about 7 days; new logic is quoted with the scope.

