
Training datasets for facial recognition models consist of controlled, front-facing images captured under ideal conditions. Production environments introduce variability across pose, lighting, camera sensors, and demographics that these datasets fail to capture. Models inherit these blind spots. Building systems that generalise requires datasets constructed around real-world variability from the very beginning.
PROJECT OVERVIEW
| Category | Specification |
|---|---|
| Model | Facial Recognition |
| Volume | 120,000 Images across 5,000 users |
| Image Types | Current Selfies + Historical Images spanning up to 10 years |
| Regions | South Asia · Africa · Latin America · Southeast & East Asia · North America & Europe |
| Variation Captured | Lighting, pose, expression, glasses, hairstyle, facial hair, accessories, indoor/outdoor, ageing |
| Consent | Collected for Facial Recognition Models |
Wide Demographic Spread, Historical Depth, and the QC Problem
Four challenges that made this project an important milestone for Humyn Labs:
Finding users with historical images. Current selfies are easy to collect. Finding 5,000 people who have 10+ images of themselves spread across the last decade, with usable metadata, EXIF data, and location information intact, is a different constraint entirely. Most people don't archive their photos with model training in mind. Images get compressed, stripped of metadata, and shuffled across devices and apps over the years. Recruiting users who actually had this kind of historical record, and could surface it in a technically usable form, was the first constraint that shaped everything else.
Demographic spread at scale. A facial recognition model trained on a narrow demographic profile will perform unevenly across the populations it's deployed on. The corpus had to represent meaningful variation across geography, age, and gender. Coordinating that spread across five global regions simultaneously, while maintaining collection standards at each site, is the next one.
Consent. Users were being asked to share not just a selfie, but a decade of their face, images tied to real locations, real dates, real changes in how they look.
Quality. Collecting at this volume across five regions, bad input is inevitable, blurry images, wrong subjects, missing metadata, and poor posture that defeats the purpose. The standard approach is post-collection QC: review everything after the fact and discard what doesn't meet spec. The problem with that is waste, wasted user time, wasted collection effort, and a pipeline that stays slow.
Instead, a purpose-built QC tool was developed that gave users real-time feedback on their uploads, catching and correcting bad images at the point of submission. It reached 94% accuracy on feedback quality, eliminating bad data before it entered the pipeline rather than cleaning it up afterwards.
120,000 Images. 5,000 Users. No Junk in the Pipeline.
The completed dataset delivered 120,000 images across 5,000 users spanning five global regions, with the within-person variation across time, condition, and appearance that a robust facial recognition model requires.
Collecting facial imagery at this scale required more than infrastructure; it required trust. In high-income regions, especially, biometric data sharing carries significant privacy sensitivity, making participation unlikely without trusted intermediaries. A community-led approach enabled informed consent, clear communication of data usage, and sustained contributor engagement.
The collection specifications itself was technically demanding, requiring participants to capture images across controlled and variable conditions while maintaining strict quality standards. Local leaders played a critical role in translating technical requirements into actionable takeaways, ensuring fidelity without loss of dataset integrity.
This approach made it possible to collect high-consent, high-variance biometric data at a production scale, resulting in a dataset suitable for real-world facial recognition deployment.