← Back to Research / Technical Whitepaper
3D Neural Models DOI: 10.6084/m9.figshare.33416719 ↗ Published March 2026

Sketch to 3D: Fast surface and depth synthesis from rough hand drawings

Authors: Jabed Bhuiyan & Kites 3D Vision Research Lab ‱ [email protected]
📄 Download Full PDF (4 Pages) View on Figshare (DOI) ↗ Try draw3d.online ↗

Executive Summary

Translating coarse 2D line sketches into 3D geometry traditionally requires hours of manual 3D modeling or 2–6 minutes of iterative Score Distillation Sampling (SDS). We developed a single-pass, multi-scale latent conditioning framework that predicts continuous surface normal fields, continuous depth maps, and physically based rendering (PBR) materials in under 3.9 seconds on consumer GPU clusters.

3.84s
Inference Latency
0.946
Normal Cosine Fidelity
91.4%
Artist Study Preference
40,000+
Creators Served in Draw3D

1. The Challenge of Sketch-to-3D Modeling

Hand sketching is the fastest way for industrial designers, architects, and digital artists to capture spatial ideas. However, hand drawings are inherently ambiguous: strokes have variable thickness, contours are often incomplete or self-overlapping, and there is zero explicit depth or lighting data.

Prior generative approaches relied on 3D NeRF optimization (e.g., DreamFusion), which requires running hundreds of diffusion denoising steps. This takes 3–5 minutes per asset, breaking creative flow and making live, interactive design exploration impossible.

2. Our Technical Approach

We reformulated the 3D synthesis pipeline into two synchronized neural streams operating in a single forward pass:

Stroke-Aware Attention Encoder (SAE)

A hierarchical tokenizer extracting multi-resolution stroke contours (1/4 to 1/32 scale) with sinusoidal position embeddings, remaining invariant to pen jitter and line gaps.

Dual-Stream Geometric Alignment

Jointly estimates metric depth and surface normal vectors while enforcing differential tangent constraints, preventing surface distortion or hollow geometry.

Mathematical Formulation:
L_total = λ_cos L_cos(N, N*) + λ_depth L_L1(D, D*) + λ_align (1 − ⟹N, n(D)⟩)

3. Quantitative Benchmarks

Evaluated across 1,200 diverse sketch-geometry pairs on the Objaverse-Sketch and QuickDraw-3D datasets:

Method Normal Cosine (↑) Depth RMSE (↓) Inference Time VRAM
DreamFusion (SDS) 0.742 0.184 240.0s 18.2 GB
Zero-1-to-3 + NeRF 0.789 0.142 180.0s 14.5 GB
TripoSR (Feedforward) 0.821 0.116 6.8s 11.0 GB
Our Framework (Draw3D) 0.946 0.052 3.84s 6.2 GB

4. Production Deployment on Draw3D

This architecture serves as the core rendering backbone of draw3d.online. Using TensorRT FP16 quantization with custom fused CUDA operators, production latency was reduced from 7.2s to 3.84s while cutting GPU compute costs by 62%.

Cite This Work & Download Paper

You can download the full academic PDF or copy the BibTeX citation for academic publications:

@article{bhuiyan2026sketch3d,
  title={Sketch to 3D: Fast Surface and Depth Synthesis from Rough Hand Drawings},
  author={Bhuiyan, Jabed and Kites 3D Vision Research Lab},
  journal={Kites Technical Reports},
  volume={TR-2026-01},
  year={2026},
  doi={10.6084/m9.figshare.33416719},
  url={https://doi.org/10.6084/m9.figshare.33416719}
}
INITIATE PROJECT

Automate your business with bespoke AI systems.

From autonomous agents and intelligent CRM bots to full-stack custom web applications — we engineer scalable software for the future.

Plan Your AI Project Explore AI Services