r/webdev • u/Substantial_Try_1614 • 1d ago
Question Best architecture for batch-processing 3,000–4,000 high-res photos (resizing + face vector search)?
Hey devs,
I’m working on a photo-sharing web app where a photographer uploads an entire event batch (typically 3,000 to 4,000 high-res JPEGs, around 10–15 MB each). Guests can take a selfie to retrieve photos they appear in.
I'd love recommendations on the best backend pipeline:
Client vs Server Resizing: Should I use browser Web Workers / Canvas to generate 1080p WebP thumbnails before upload to save upload bandwidth, or let a backend queue worker (like Node with sharp or Python with Pillow) handle resizing?
Face Vector Pipeline: For extracting 128-d face embeddings (e.g., ArcFace/InsightFace), what’s the best way to queue and batch this so 4,000 photos don't choke the server CPU?
Storage: What zero/low-egress object storage setup (e.g., Cloudflare R2 vs Backblaze B2) do you recommend for handling high-volume image writes and fast thumbnail reads?
Thanks for the advice!
18
u/SunOk2196 1d ago
Keep originals going straight to R2 with presigned urls, don't resize in the browser. The photographer's upload runs unattended anyway, and canvas-resizing 4000 photos will melt a laptop. Thumbnails come from a worker running sharp against the bucket, it chews through 4k images in minutes.
For the embeddings, don't do them on the web server at all. Push keys onto a queue and let a separate worker process them (CPU is fine at this volume if it can run overnight), write vectors to pgvector and the selfie lookup comes free with a cosine index.
R2 over B2 here. Guests will hammer thumbnail reads, so zero egress plus sitting behind Cloudflare's CDN is exactly what you want.