CarcassFormer: An End-to-end Transformer-based Framework for Simultaneous Localization, Segmentation and Classification of Poultry Carcass Defect

Minh Tran1, Sang Truong1, Arthur F. A. Fernandes2, Michael T. Kidd1, and Ngan Le1

1 University of Arkansas · 2 Cobb-Vantress, Inc

Video Demonstration

The video demonstrations showcase the two main components of CarcassFormer. The left video presents real-time carcass inspection, detecting and counting wings, feathers, feathers on skin, exposed flesh, and the total number of processed carcasses. The right video demonstrates cut-part classification, identifying chicken components such as wings, breasts, tenders, skin, drumsticks, carcasses, and butts.

Abstract

In the food industry, assessing the quality of poultry carcasses during processing is a crucial step. This study proposes an effective approach for automating the assessment of carcass quality without requiring skilled labor or inspector involvement. The proposed system is based on machine learning (ML) and computer vision (CV) techniques, enabling automated defect detection and carcass quality assessment. To this end, an end-to-end framework called CarcassFormer is introduced. It is built upon a Transformer-based architecture designed to effectively extract visual representations while simultaneously detecting, segmenting, and classifying poultry carcass defects. Our proposed framework is capable of analyzing imperfections resulting from production and transport welfare issues, as well as processing plant stunner, scalder, picker, and other equipment malfunctions.

To benchmark the framework, a dataset of 7,321 images was initially acquired, which contained both single and multiple carcasses per image. In this study, the performance of the CarcassFormer system is compared with other state-of-the-art (SOTA) approaches for classification, detection, and segmentation tasks. Through extensive quantitative experiments, our framework consistently outperforms existing methods, demonstrating remarkable improvements across various evaluation metrics such as AP, AP@50, and AP@75. Furthermore, the qualitative results highlight the strengths of CarcassFormer in capturing fine details, including feathers, and accurately localizing and segmenting carcasses with high precision. To facilitate further research and collaboration, the source code and trained models will be made publicly available upon acceptance.

System overview

Flowchart of the proposed CarcassFormer framework
Top: Overall flowchart of our proposed CarcassFormer consisting of 4 components: 1. network backbone; 2. pixel decoder; 3. mask-attention transformer decoder; 4. instance mask class prediction. Bottom: details of third component mask-attention transformer decoder.

Experimental setup

Camera setup for poultry carcass data collection
Camera setup for data collection. A black curtain is hung behind the shackle to provide a certain contrast to the carcasses. A camera is placed to capture the carcasses within the black curtain.

Dataset statistics

CarcassFormer dataset statistics

Qualitative results

Performance comparison of poultry carcass defect segmentation methods
Performance comparison (A) Mask R-CNN [He et al. (2016)], (B) Mask2Former [Cheng et al. (2022)] and (C) our CarcassFormer on the defect where single carcass with feathers. In the Segmentation column, notable parts with feathers were highlighted. Compare with Mask R-CNN and Mask2Former, our CarcassFormer can localize carcass with more accurate bounding box and segment carcass with more details on feathers.

Quantitative results

CarcassFormer quantitative results

Table 1 summarizes CarcassFormer’s performance on images containing a single carcass. The model achieves excellent results across all three backbones, with ResNet-34 providing the strongest overall performance: 97.70% detection AP and 99.22% segmentation AP. Classification is also highly accurate, reaching 98.02% AP for normal carcasses and 97.38% for defective carcasses.

Table 2 presents the more challenging multiple-carcass scenario, where overlapping objects and occlusions reduce detection performance. ResNet-50 achieves the best detection AP of 90.45% and segmentation AP of 98.96%, while maintaining strong classification results for both normal and defective carcasses. Despite the increased complexity, every configuration remains above 85% AP across the main evaluation metrics.

These results are particularly promising because CarcassFormer performs detection, defect classification, and instance segmentation within a single framework. Its strong performance on both isolated and overlapping carcasses demonstrates the feasibility of automating carcass inspection in realistic processing environments, where multiple objects may appear simultaneously and precise defect localization is required.

BibTeX

@article{tran2024carcassformer,
  title={CarcassFormer: an end-to-end transformer-based framework for simultaneous localization, segmentation and classification of poultry carcass defect},
  author={Tran, Minh and Truong, Sang and Fernandes, Arthur FA and Kidd, Michael T and Le, Ngan},
  journal={Poultry Science},
  volume={103},
  number={8},
  pages={103765},
  year={2024},
  publisher={Elsevier}
}