• No se han encontrado resultados

In practice, self-occlusion and occlusions between different moving objects or between moving objects and the background are inevitable [98], and multiple camera systems offer promising methods to reduce ambiguities owing to

occlusion. Multiple cameras have been used to choose the best view considering occlusion or to estimate 3D information of each object for coping with occlusion [99-103]. The use of multiple cameras, however, requires complex computation to match identical objects from different cameras or to calibrate the cameras for 3D information.

There are several studies that propose ways to tackle the occlusion problems using a single camera by handling regional merges of multiple objects. They deal with another similar problem whereby a single object can split into multiple regions which yield separate measurements. The splitting may result from crossing occlusions or errors in background subtraction, and it can be generated despite the use of good background subtraction techniques [104]. Their methods can be largely divided into use and non-use of colour information.

A. Use of Colour

The system of Chen and colleagues [105] counts pedestrians, passing through a gate or a door, with a zenithal video camera, as shown in Figure 14a. Hue saturation intensity (HIS) colour histograms are used to distinguish one pedestrian from another in a multi-people heap because the hue component is intimately related to the way in which humans perceive colour. As the colour label can become ineffective in identification in the case of multiple pedestrians wearing same-coloured clothing, the two overlapping boxes in Figure 14b, which bound identical colour patterns in adjacent images, are judged as the same person. In order to analyse merging or splitting cases within the door area, area changes of moving regions are checked, on the basis of the fundamental cases of merging or splitting in Figure 14c. This study is not, however, concerned with a single person splitting into multiple regions owing to errors in background subtraction.

(a) (b)

(c)

Figure 14. (a) People Counting System, (b) its Tracking and (b) Basic Cases of Merge-Split [105]

Conversely, the approach of Medioni and colleagues [106] does not cope with multiple objects merging into one region, but only with a single object splitting into multiple regions. Their work involves detection and tracking of moving objects and analysis of their trajectories to recognise the behaviour of the moving objects. In order to extract the correct trajectory of each object, aperture problems, which can split a single object into multiple regions, are handled by measurement of the grey-level similarity between a moving region at one frame and a set of regions at the next frame in its neighbourhood. The size of this neighbourhood is estimated from the object motion amplitude, and the matches of moving objects between consecutive frames are represented by nodes and edges, as illustrated in Figure 15.

Figure 15. Detected Regions and Associated Graph [106]

The vehicle tracking system of Song and Nevatia [107] for street surveillance is based on an appearance model of a colour histogram to detect both multiple objects merging into one region (Figure 16a) and a single object splitting into multiple regions (Figure 16b). Apart from the colour histogram, each of the blobs, detected as moving vehicles, is modelled as a rectangle to predict its new position and check overlaps between predicted rectangles and observed rectangles for blob association over successive frames.

(a) (b) Figure 17. (a) Partial Occlusion and (b) Crowding [108]

Guha and colleagues [108] defined six qualitative occlusion primitives, based on the well-known cognitive assumption of persistence, under which objects continue to exist even when hidden from the view. The primitives are isolated,

partial occlusion, crowding, disappear, enter, and exit, which respectively

indicate blobs separated and fully visible, blobs separated but partially invisible (Figure 17a), blobs merged into one region (Figure 17b), blobs detected previously but no longer visible, new blobs with no relation to previous blobs, and blobs disappearing at the scene boundary. In order to recognise these occlusion primitives, each agent is characterised by its occupied pixel set, weighted colour distribution, and the trajectory of the minimum bounding rectangle of the pixel set. The agent is also associated with detected foreground blobs, based on the colour distribution and predicted agent position from the trajectory. All the occlusion primitive notations on each agent are recorded in the history for further recognition.

The method of McKenna and colleagues [104] tracks people through mutual occlusions when they form groups and separate from one another, as presented in Figure 18. In order to overcome the problem of a person splitting into multiple regions, the conditions of multiple regions to form a single person are defined to be in close proximity, to have overlapped projections onto the x-axis, and to have

a total area larger than a threshold. In order to track people consistently when they enter and leave people groups, a colour model is built and adapted for each person being tracked. The tracker based on the colour information can fail in tracking of each person, however, when two people clothed in a very similar manner form a group and subsequently separate, for example.

Figure 18. People in Groups [104]

Figure 19. Process of Prediction and Matching [109]

B. Non-Use of Colour

when the objects are predicted to merge. The blue ellipse of a broken line in Figure 19 is an example of the new synthesised blob. The real segmented blob is matched with the objects separately and also the synthesised blob as shown in Figure 19 by use of a geometric shape-matching algorithm. These association methods work well as long as position and motion of target objects are predictable.

Figure 20. Target-Measurement Association [110]

Joo and Chellappa [110] proposed a multiple-hypothesis approach to tracking multiple objects by handling objects which enter or exit the view or regionally merge or split, as well as by detecting split fragments of a single object owing to limitations in background subtraction. For those kinds of objects, a single target (tN) may need to be associated with multiple measurements (mD), and multiple targets with a single measurement, as shown in Figure 20. The multiple- hypothesis tracking considers a set of feasible hypotheses, regarding joint associations between targets and their measurements. The centre coordinates, the bounding box size, and the velocity of each target are defined as its state at every frame, and the position is predicted and compared with the real measurements for the best match.

C. Summary of Studies in Handling of Regional Merges and Splits

The reviewed studies regarding handling of regional merges and splits are summarised in Table 3 to present the major cues and drawbacks.

Table 3. Summary of Studies in Handling of Regional Merges and Splits

Classification First Author (Year) of Related Studies

Cues for Handling of Regional Merges and Splits

Drawback

Chen (2006) HIS Colour Histogram

+ Box Overlap

Medioni (2001) Grey-Level Similarity

+ Neighbourhood Size

Song (2007) Colour Histogram

+ Predicted Position

Guha (2006) Colour Distribution

+ Predicted Position Use of Colour

McKenna (2000) Colour Model

+ x-Projection Overlap + Area Size

The use of colour information can cause confusion in the case of a single object wearing multiple colours or multiple objects with similar colours in a group.

Kumar (2006) Prediction and Match of

Shape and Position Non-Use of

Colour

Joo (2007) Prediction and Match of

Position

The association of prediction and real measurements can work when position and motion of targets are predictable.

Colour information is employed in many studies to identify each of the multiple objects which appear together in one image region, to detect regional fragments of a single object based on proximity, or to consistently track individuals despite regional merges and splits. As the sole use of colour information can incur

limit the range of searching for the identical object in the next frame within the area where the object will possibly be.

The studies, which do not use any colour information, predict the shape or the position of each object and detect the closest match with the real measurements. The association can work only when the motion of each object is correctly estimated and its proper position in the next frame is predictable.

As the use of colour similarity can confuse tracking of individuals, the work in this thesis employs position information commonly used in existing methods of handling regional merges and splits, as described in Section 4.3.4A.

(a) (b) (c) Figure 21. Pictorial Structures from Videos of (a) Zebra, (b) Tiger, and (c) Giraffe

[111]

Documento similar