To main content

Combining Open-Vocabulary Detection with Foundation Models for Object Recognition in an Industrial Bin Picking Context

Abstract

Object recognition for industrial bin picking traditionally requires large annotated datasets and retraining when new parts are introduced. The sim-to-real gap further limits approaches that rely on synthetic training data alone. In this paper, we present an initial investigation into a threestage pipeline that combines Grounding DINO (GroundingDINO) for open-vocabulary detection, the Segment Anything Model 2 (SAM2) for instance segmentation, and a lightweight Multi-Layer Perceptron (MLP) classifier operating on DINOv2 (self-supervised vision transformer) features. Only the classifier requires training on synthetic CAD renders; detection and segmentation rely entirely on the zero-shot capabilities of foundation models. This design isolates the sim-to-real gap to a single lightweight component operating on robust, pre-trained features. We evaluate our approach on the challenging TLESS dataset, which features textureless, visually similar industrial objects with frequent symmetries, and achieve 70.0% class-level F1 on real images using only synthetic training data. We further show that DINOv2 features transfer more robustly from synthetic to real domains than alternative backbones such as SigLIP (Sigmoid Loss for Language Image Pre-training), and that incorporating realistic distractor objects during training improves classifier discrimination.

Category

Academic article

Language

English

Author(s)

Affiliation

  • SINTEF Industry / SINTEF Manufacturing
  • University of Inland Norway

Date

01.01.2026

Year

2026

Published in

Proceedings - European Council for Modelling and Simulation (ECMS)

ISSN

2522-2414

Volume

40

Issue

1

Page(s)

379 - 388

View this publication at Norwegian Research Information Repository