
Aurum at SThree: Testing Recruitment AI with Synthetic Candidates
How I built Aurum, the 25-agent synthetic-data system used to test and improve SThree's live recruitment platform.
I write about what I build, what the evidence shows and where the difficult parts remain.
Independent research paper · July 2026 · Publication pending
Eleven core deep-research architectures, 90 queries and a three-judge reliability audit. The main comparison holds GPT-4o and the tool layer fixed. Two local 7B systems are compared separately.

How I built Aurum, the 25-agent synthetic-data system used to test and improve SThree's live recruitment platform.

How I used synthetic data and AstroGAN while building pose estimation for ELSA-M, and why closing a measured domain gap was only part of the problem.

How I replaced a per-skill language-model validator with retrieval, fixed filters, calibration and a 7 kB classifier head, cutting operating cost by about 180 times while reaching about 96% measured precision.

How I restored and animated old family photographs for my gran's 85th birthday without letting the technology overwhelm the people in them.