A studio that turns an ambiguous model problem into a dataset you can train on and an evaluation you can defend — under one accountable owner.
You hand us the problem and the standard of done. We design the task, source the right people, run production and quality control, and return the result with evidence it meets the bar. Not a labor marketplace. Not another annotation tool. One team that owns the outcome.
A labeling request that sounds simple hides unresolved research: what to judge, which cases are ambiguous, how many ratings, which metric, who's qualified. Get it wrong and you get a big dataset no one trusts — after the deadline.
Tools and compute stay modular. What we own is the operating system that makes the output trustworthy — and repeatable release after release.
Turn an unclear request into a rubric, gold set, pilot, metric plan and price.
Fixed fee · 2–4 wksCollection, annotation, review and adjudication — sold by the accepted unit.
Per accepted unitA product-specific benchmark, human ratings, failure clusters, release report.
Fixed / retainerPairwise, pointwise, critiques and corrections for post-training.
Versioned batchesA vetted pod around a named domain lead, delivered as managed output.
Capacity retainerSFT / preference runs on your approved infrastructure, with evaluation.
MilestonesDiagnose a failing dataset or vendor — defects, fraud, drift — then fix what matters.
Diagnostic + fixLand cheap, prove the standard, grow into recurring eval and model-improvement data.
LTV pathVisual models turn subjective quality into a measurable training problem — where a standard can be owned. We translate creative judgment into a repeatable rubric instead of one generic beauty score.
A training or eval milestone is blocked by human-data capacity or a scarce expert domain.
A release needs a trustworthy benchmark, domain adaptation or continuous evaluation.
Internal hiring is too slow, or no team spans data, ML and operations at once.
The model must be judged by people with verifiable professional competence.
The parts most buyers underestimate are exactly what we own: the standard, the statistics and the controls that make a number trustworthy.
Known-answer items screen every rater on live skill, refreshed against leakage.
Cohen, Fleiss, Krippendorff by design — a low score means the rubric is broken, not the raters.
Dawid–Skene weights each rater by a confusion matrix and returns per-item confidence.
Wilson-score intervals, not point estimates — e.g. 0.86 [0.79, 0.91].
Identity, impossible-speed, answer patterns and audits route work to review, never silently define truth.
Source, instruction, rater, reviewer and model version on every accepted artifact.
The client sees one standard, whatever the source of the people. Access to workers is only the starting inventory; knowing who can produce which judgment, at what quality, is the asset.
Clear, high-volume tasks under task-specific onboarding and hidden gold.
Modality, policy or product-domain work with paid calibration and review sampling.
Verified professionals — medicine, law, finance, design — for consequential judgments.
Owns and maintains the standard: gold, edge cases, reviewer calibration, final disputes.
Scale, Surge and Mercor sell the frontier — sales-led, no public price, minimums reported in the millions. The team that needs 200 expert-graded evaluations this week has nowhere to go. That's our door.
A fixed-fee, 2–4 week engagement that ends in an approved protocol and a pilot readout — the bridge from consulting to repeatable production.
Fixed feeRecurring data and model evaluation, sold by the accepted unit under a frozen spec, with a weekly evidence pack.
Per accepted unitThe same benchmark every release; compute passes through transparently, so capital stays off our balance sheet.
Monthly retainerWe guarantee accepted work against the agreed specification. Model uplift is a controlled, statistically adequate milestone — never a blanket promise.
A technical services studio: data & evaluation design, a managed tiered workforce, statistical quality control and ML integration — starting with visual generative AI.