Abstract
Object recognition for industrial bin picking traditionally requires large annotated datasets and retraining when new parts are introduced. The sim-to-real gap further limits approaches that rely on synthetic training data alone. In this paper, we present an initial investigation into a threestage pipeline that combines Grounding DINO (GroundingDINO) for open-vocabulary detection, the Segment Anything Model 2 (SAM2) for instance segmentation, and a lightweight Multi-Layer Perceptron (MLP) classifier operating on DINOv2 (self-supervised vision transformer) features. Only the classifier requires training on synthetic CAD renders; detection and segmentation rely entirely on the zero-shot capabilities of foundation models. This design isolates the sim-to-real gap to a single lightweight component operating on robust, pre-trained features. We evaluate our approach on the challenging TLESS dataset, which features textureless, visually similar industrial objects with frequent symmetries, and achieve 70.0% class-level F1 on real images using only synthetic training data. We further show that DINOv2 features transfer more robustly from synthetic to real domains than alternative backbones such as SigLIP (Sigmoid Loss for Language Image Pre-training), and that incorporating realistic distractor objects during training improves classifier discrimination.