Depth → Image (ControlNet)

Upload a photo - we estimate its depth map (MiDaS) and use that 3D structure to control a generated image. Keeps the spatial layout while changing everything else.

We estimate a MiDaS depth map from this. Images with clear foreground/background separation work best.
Mofuthu 0.7 Stricter
~1,200 tokens (SDXL × 1.2 ControlNet)
Bo_lemo

U rata Free.ai? Reka ho ba lelapa la hao!

Litlhaku tsa leqephe lena