Knowledge · Robotics

When is 3D bin picking needed?

Many production tasks are simple but monotonous for people: taking a part from a bin, placing it on a conveyor, or loading it into a CNC machine. A robot arm or cobot can take over that repetitive work. The difficulty is that a bin is not an orderly supply from a robot’s point of view. Every part lies differently, parts cover one another, and the scene changes after every pick. 3D bin picking helps the robot see where a part is and decide whether it can be picked and moved safely.

Robot gripper above randomly piled metal parts
The camera may see a part clearly and the pick can still fail. A magnet only needs to touch a suitable surface; a finger gripper must reach around the part and close. The robot arm must then carry the part out of the bin without a collision.

The short answer

When is 2D enough and when is 3D needed?

Do not start with the most elaborate 3D solution by default. If parts lie at a known height with the same face upward, a straightforward 2D camera may be enough. Only when height, tilt and overlap keep changing does the robot cell need to perceive and plan more for itself.

From fixed position to random pile

Not every pick from a container is random bin picking

The bin itself says little about how difficult the task is. When every part has a fixed position in a tray, the robot can repeat almost the same movement. When parts lie loose and overlap, the system must decide again for every pick what is visible, reachable and safe to grasp.

PresentationWhat is known?Usually suitable approachMain limitation
Fixed tray or fixturePosition, height and orientation are predefined.Robot program with presence verification; vision may not be required.Limited flexibility across product changes.
Separated on a planeKnown height, limited orientations, no overlap.2D camera for X, Y and in-plane rotation.Fails when parts tilt or cover one another.
Layered or semi-orderedSame face upward, but height or limited tilt varies.2.5D height map or 3D guidance with constrained pose freedom.A depth image contains visible surfaces only.
Randomly piledPosition and orientation are unknown; parts overlap.3D perception, pose or grasp estimation, and collision-aware planning.Each pick changes the scene; reacquisition is normally required.
Entangled or deformableShape and physical relationships may change during the pick.3D plus a specialised gripper, active rearrangement or mechanical singulation.Direct bulk picking may be technically or economically unattractive.

Why known object height helps so much

On a flat support with known part height, a 2D detection can often be converted directly into a fixed pick Z. The number of possible orientations is reduced as well. Once parts stack, tilt or vary in shape, this simplification disappears: actual height, surface normal and gripper clearance must be recovered from the scene.

The complete chain

A bin-picking system must answer six questions

A person handles these steps almost without thinking. For a robot, each one must be measurable and verifiable. That is how a promising demonstration becomes a cell that keeps working throughout a production shift.

01

What is visible?

The camera measures colour, intensity and depth, including missing or unreliable pixels.

02

Which part is this?

Software separates instances, identifies the type, and estimates pose or possible grasp points directly.

03

Where can the gripper make contact?

Flatness, edges, holes, material, centre of mass and clearance determine whether a grasp is stable.

04

Can the robot reach it?

Reach, joint limits, singularities, bin walls and full tool geometry constrain the choice.

05

Can the part leave the bin?

Approach, gripper closure, lift and extraction must be collision-free with the part attached.

06

Did the pick actually succeed?

Vacuum pressure, finger position, force sensing or a verification image confirms the pick and drives recovery.

Complicating factors

Why 3D bin picking is difficult in the real world

A person naturally picks the topmost free part and leaves an interlocked one alone. A robot must derive those choices from camera data, geometry and explicit rules. The combination of gloss, overlap, bin walls and limited room to move is what makes production harder than a tidy demonstration.

Occlusion and physical contact

The camera measures visible surfaces only. A hidden edge or underside is inferred from CAD or learned priors, not directly observed. The target may also be trapped beneath a neighbour.

Gloss, black materials and transparency

Gloss can redirect projected light or saturate a sensor. Dark material may return too little light, while transparent parts transmit or refract it. No depth technology solves every material universally.

Symmetry and weak geometry

A cylinder may have several equivalent orientations. That is harmless if every orientation is acceptable, but critical when a hidden key must later be aligned.

Bin walls, corners and depletion

Walls create shadows, reflections and restricted tool access. A fixed camera also observes a full and nearly empty bin from different ranges and angles.

Nesting, hooking and double picks

Rings, clips, springs and sheet parts can interlock. A correct pose does not prove that exactly one part will separate.

Calibration and the tolerance chain

Errors in camera intrinsics, camera-to-robot calibration, robot accuracy and the tool centre point accumulate. Accurate depth alone therefore does not guarantee an accurate pick.

Pose found. What next?

The camera suggests a promising pick; the robot still has to execute it safely

The camera may describe exactly where a part lies, but the part is not picked yet. A magnet or vacuum cup mainly needs a clear contact surface. A finger gripper must reach around the part and have room to close. The robot then needs a posture and path that carry both the gripper and the held part safely past the bin wall and neighbouring parts.

Vision and grasp planning

  • Detection, segmentation and pose estimation
  • Generation of several possible grasp points
  • Accounting for uncertain depth and occluded surfaces
  • Selection by contact quality and clearance

Robot and motion planning

  • Inverse kinematics and joint-limit checks
  • Avoidance of self-collision, walls and neighbours
  • Planning approach, lift, extraction, transport and placement
  • Adapting speed and acceleration to grip and payload

A practical example

The camera finds a correctly posed part against a side wall. A top suction cup can reach the surface, but the wrist strikes the rim during approach. The correct system response is not “pose found, therefore pick,” but to choose another grasp, another robot configuration, another view or another part.

Reducing complexity deliberately

What makes an application demonstrably easier?

Known height or support plane

Often reduces the problem from six to three degrees of freedom and permits a fixed pick height.

A single layer

Less occlusion, lower collision risk and often several picks from one acquisition.

Rigid, matte parts

Usually provide more stable depth and retain their geometry during detection and gripping.

Large accessible grasp surfaces

More candidate grasps and greater tolerance of measurement and calibration errors.

A shallow, accessible bin

Better lines of sight and more room for the wrist, tool and extraction motion.

Mechanical pre-processing

A vibration plate, conveyor, step feeder or intermediate station can reduce overlap before vision begins.

Important: mechanical singulation is not failed vision. When parts entangle, are difficult to measure or demand a very short cycle, simple mechanics may be the most robust and economical system solution.

Solution types

Choose the architecture before the camera

2D with known pick height

For flat, separated parts. Fast, explainable and often easier to maintain.

2.5D height map

For layers, boxes or top surfaces with height variation. A height map contains one visible depth per image point, not hidden geometry.

CAD-based 6D pose

Fits known 3D geometry to the measured point cloud. Strong for fixed industrial parts when CAD, calibration and visible geometry agree sufficiently.

AI for segmentation or direct grasps

Can help with variable appearance, instance separation and grasp estimation. Training must represent actual materials, lighting and pile states.

Hybrid approach

Combines, for example, AI segmentation with CAD pose, depth verification and explicit collision checking. This is often more diagnosable than a fully end-to-end model.

Singulation before the robot

Spreads or meters parts so the robot can use 2D or simpler 3D. Particularly relevant for small, interlocking or high-throughput parts.

Measuring depth

3D technologies fail in different ways

The best method follows from working distance, measurement volume, material, required detail and motion during capture. Test not only average accuracy, but also how much valid depth remains on critical grasp surfaces.

Stereo and active stereo

Calculate depth from image differences between cameras. Projected texture helps on uniform surfaces; repetitive patterns, gloss and occlusion remain difficult.

Structured light

Projects a known pattern and measures its deformation. It can provide detailed depth in controlled conditions, while shadows, ambient light and movement during pattern capture can interfere.

Time of flight

Determines distance from the travel time or phase of active light. It produces dense depth quickly, but multipath, saturation and weak returns can invalidate pixels.

Fixed camera

Simpler calibration and often shorter cycles when the complete bin remains within range and field of view. In deep bins, capture quality changes between full and empty states.

Robot-mounted camera

Allows alternate viewpoints and more consistent working distance. The trade-offs are extra motion, stationary capture, payload, protection and cable routing.

Learn how 3D cameras measure depth →

From demonstration to production

A feasibility test should seek out the difficult picks

Do not test only a few easy parts placed neatly on top. Use actual parts and the real bin, including full corners, glossy surfaces, interlocked parts and an almost empty bin. That reveals where the robot needs help or a different approach before the investment is made.

Test at least

  • Full, half-full and nearly empty bins
  • Parts against every wall and in every corner
  • Different production lots, wear, oil, dust and colour variation
  • All stable orientations and unfavourable overlaps
  • Maximum robot reach and worst-case payload
  • Failed grasp, double pick, shifted pile and no-detection scenarios
  • Complete depletion of several bins without manual rearrangement

Measure the complete cell

  • First-attempt success: confirmed single-part picks per attempt
  • Pick availability: scans with at least one reachable, safe grasp
  • Placement success: correctly placed parts per requested cycle
  • Double-pick and drop rate: errors per successful grasp
  • Cycle time: median and relevant slow percentiles, including rescans
  • Intervention rate: manual actions per bin, hour or thousand parts
  • Depletion performance: agreed residual quantity without intervention

Optimise time per correctly placed pick

The fastest isolated robot move is not automatically the best strategy. A slightly slower grasp with a high success probability can deliver more production than an aggressive grasp that often causes a rescan, drop or operator intervention.

Available 3D hardware

The camera must fit the measurement volume, not the other way around

The Hikrobot MV-DB range includes RGB-D models for tasks such as singulation, volume and pose measurement. Compare each model's measurement range, clearance distance, field of view, depth accuracy and scan rate against both the full and empty bin.

Hikrobot MV-DB RGB-D camera
Hikrobot MV-DB

RGB-D cameras for depth, shape and pose

Use the product page to compare models and specifications. Material behaviour, viewing angle, bin walls and gripper access remain part of the practical test.

View MV-DB models →

Compare all 3D cameras and measurement methods →

Frequently asked questions

Practical questions about 3D bin picking

Is a 3D camera always needed for robot picking?
No. For flat, separated presentation with known height, 2D guidance may be enough. 3D becomes relevant when height, tilt, overlap or collision space changes per pick.
Is object detection the same as bin picking?
No. Detection supplies an object hypothesis or pose. Bin picking also needs a feasible grasp, inverse kinematics, collision checking, an executable trajectory, pick verification and recovery behaviour.
Can AI solve difficult surfaces?
AI can improve segmentation, pose or missing geometry, but it does not make unmeasured surfaces physically visible. Gloss, transparency and occlusion still require practical testing, alternate viewpoints or mechanical simplification.
Must the system scan after every pick?
Usually for a random pile, because neighbouring parts may shift. Stable layers or separated products can sometimes allow several picks from one acquisition.
When is mechanical singulation better?
When parts strongly hook or nest, are difficult to measure optically, provide little gripper clearance, or when the cycle target cannot be met with repeated 3D acquisition and planning.
What is needed for an initial assessment?
Representative parts, CAD where available, bin dimensions, fill height, material and surface condition, required output orientation, robot and gripper data, cycle target and permitted intervention rate.

Continue within Sedeco

From repetitive manual work to a robot cell that keeps running

An initial assessment does not have to be complicated. A few representative parts, the real bin and the required cycle time are usually enough to identify a promising approach and decide what should be tested first.