DeepBroilerTrack: End-to-End Automatic Tracking of Multiple Broilers Using Multiple Cameras

Hao Vo1, Thinh Phan1, Michael T. Kidd1, James Mason2, Santiago Avendano2, and Ngan Le1

1 University of Arkansas · 2 Cobb-Vantress, Inc

Video Demonstration

The demo visualizes multi-chicken tracking using bounding boxes with corresponding identity labels. Unified ground-plane points map each detected chicken into a shared spatial reference, enabling consistent tracking across camera views.

Abstract

Reliable tracking of individual broilers in commercial poultry facilities is essential for large-scale welfare monitoring, behavior analysis, and precision breeding. Multi-camera systems have emerged to address challenges posed by occlusions, dense flock layouts, and strong appearance similarity. However, existing approaches often rely on hand-crafted association heuristics and cascade-style two-stage designs, which struggle to generalize in crowded and occlusion-heavy environments. To overcome these shortcomings, we propose DeepBroilerTrack, an end-to-end multi-camera tracking framework unifying early-stage feature fusion within a Bird’s-Eye View (BEV) representation. Our method leverages a camera-dropout augmentation strategy to improve cross-view generalization, aggregates multi-camera features into a unified BEV space for robust detection and temporal association, and projects BEV tracklets back to individual views through a learned ID assignment module to maintain identity consistency across all cameras. Experiments on a real-world broiler dataset demonstrate that our method significantly improves tracking stability and identity (ID) consistency over existing multi-camera baselines, offering a practical and scalable solution for intelligent poultry management in challenging production settings.

System overview

Overview of the DeepBroilerTrack pipeline
Overview of our proposed DeepBroilerTrack pipeline, consisting of eight modules: ➀ Image augmentation, ➁ Image encoder, ➂ Image feature enhancement, ➃ Cross-view aggregation, ➄ BEV decoder, ➅ BEV tracking, ➆ Image object detection, and ➇ ID assignment. The raw multi-camera images first undergo ➀ Image augmentation, including geometric perturbations and randomized camera dropout. The augmented frames are processed by the ➁ Image encoder to extract initial feature representations, which are refined through the ➂ Image feature enhancement module. The per-view features are then fused in the ➃ Cross-view aggregation module to construct a unified BEV representation, which is decoded by the ➄ BEV decoder to localize objects. Temporal consistency is enforced in the ➅ BEV tracking module, generating global BEV tracklets. In parallel, each frame is processed by ➆ Image object detection to yield 2D bounding boxes, and the ➇ ID assignment module associates them with BEV tracklets, producing ID-consistent trajectories across all camera views.

Experimental setup

Six-camera setup for DeepBroilerTrack
The setup uses six synchronized Reolink 842A cameras around a 4 × 6 ft pen: two overhead cameras provide full top-down coverage, while four corner-mounted cameras capture side views. A feeder and drinking line are positioned near the center of the pen. Videos are recorded at 1920 × 1080 resolution and 15 fps, with synchronization errors below 80 ms.

Qualitative results

Qualitative tracking results from DeepBroilerTrack
Qualitative results of our tracking framework over time, with the colors are the objects’s id and colored lines are corresponding trajectory of that object.

Quantitative results

Table 1 compares the proposed early-fusion BEV framework with the previous late-stage fusion approach. The updated method substantially reduces identity switches, demonstrating improved identity consistency and tracking stability. Although MOTA and DetA decrease slightly because of BEV projection errors, the unified spatial representation reduces occlusion and perspective-related identity fragmentation.

DeepBroilerTrack Table 1 quantitative results

Table 2 summarizes the four-fold cross-validation results. The proposed method achieves an average MOTA of 98.00%, MODA of 96.49%, Precision of 99.85%, Recall of 98.23%, and only 2.5 identity switches. These results indicate highly accurate target detection with strong identity preservation across challenging multi-camera conditions.

DeepBroilerTrack Table 2 quantitative results

Table 3 presents the per-fold comparison with existing multi-camera tracking methods. The results demonstrate that the proposed BEV-based approach delivers consistently strong tracking performance across the evaluation folds, validating the robustness of the early-fusion and multi-view geometric association strategy.

DeepBroilerTrack Table 3 quantitative results

BibTeX

@article{vo2026deepbroilertrack,
  title={DeepBroilerTrack: End-to-End Automatic Tracking of Multiple Broilers Using Multiple Cameras},
  author={Vo, Hao and Phan, Thinh and Kidd, Michael T and Mason, James and Avendano, Santiago and Le, Ngan},
  journal={Smart Agricultural Technology},
  pages={102163},
  year={2026},
  publisher={Elsevier}
}