In an embodiment, automated analysis of video data for determination of human behavior includes providing a programmable device that segments a video stream into a plurality of discrete individual frame image primitives which are combined into a visual event that may encompass an activity of concern as a function of a hypothesis. The visual event is optimized by setting a binary variable to true or false as a function of one or more constraints. The optimized visual event is processed in view of associated non-video transaction data and the binary variable by associating the optimized visual event with a logged transaction if associable, issuing an alert if the binary variable is true and the optimized visual event is not associable with the logged transaction, and dropping the optimized visual event if the binary variable is false and the optimized visual event is not associable.