How AI 2D-to-3D conversion works
The process has two steps. First, a depth-estimation model analyzes the single 2D photo and predicts a depth map — essentially a grayscale image where brighter pixels are closer to the camera.
Second, that depth map is used to synthesize a second, shifted viewpoint of the same scene, filling in the small gaps that appear at object edges using AI inpainting.
Why it isn't perfect
AI depth estimation is a prediction, not a measurement — it can misjudge depth on reflective surfaces, transparent objects, or scenes with unusual lighting.
Inpainted edges around foreground objects (the 'disocclusion' gaps) are the most common giveaway that a photo was converted rather than shot in native stereo.
A practical workflow
- Start with a high-resolution, sharply focused source photo — noise and blur confuse depth models
- Prefer photos with clear foreground/midground/background separation over flat, single-plane compositions
- Keep the synthesized depth shift modest; aggressive settings exaggerate artifacts
- Preview the result on the actual target headset or display before finalizing, since screens vary in how forgiving they are of small errors
When to convert vs when to shoot natively
AI conversion is ideal for a personal archive of old photos, or when a native stereo rig genuinely isn't available. For new content, shooting native stereo pairs still produces cleaner, more reliable depth.
See it in action at 3dstreaming.org