AI that survives contact with a real workflow.
Labelled training data, and the practical AI built on top of it.
- Turnaround
- Pilot batch in a week, builds run two to six weeks
- Revisions
- Until it is right. Not a counter.
- You keep
- Every source file, in your software.
- Starts with
- An email and whatever you already have.
About this service
Most AI projects die somewhere between the demo and the rollout. The demo works on five clean examples. Then it meets real data, real edge cases, and a team who have no particular reason to change how they already work. We build for the second part: something narrow and genuinely useful, scoped against a task you can measure before and after.
Underneath almost all of it is the training data, and that is where the real problem usually sits. Most annotation failures are not effort failures, they are agreement failures: two labellers look at the same ambiguous frame, make different calls, and the inconsistency goes into your training set where it is very hard to find later.
So we start with the guideline rather than the labelling. A small pilot, every case where reasonable people disagreed, the rule that settles it, and only then scale. Bounding boxes, segmentation masks, keypoints, classification, transcription and entity tagging, and the assistants, extraction pipelines and drafting tools built on top. We will also tell you when AI is the wrong tool, which a good share of the time it is.
02What you get
- The labelled dataset in your schema and format, not ours
- A written annotation guideline, including the edge cases and how each was decided
- An inter-annotator agreement score for every batch, and a sample you can audit independently
- A working tool your team can open, not a notebook and a slide
- The evaluation set we tested against, so you can check it again later
- Documentation of what it does badly, because everything does something badly
03How it usually goes
- 1
The first session is about the task, not the technology: who does it now, how long it takes, and what happens when it goes wrong.
- 2
We label a small pilot and deliberately hunt for the cases that break the guideline, then settle those with you. This is what determines whether the whole set is usable.
- 3
We scale to the agreed volume with a second pass on a fixed share of every batch, and build the narrow version of the tool on top.
- 4
We measure against the baseline, fix what real use exposes, and hand over the data, the code and the limits in writing.
04Is it for you
A good fit if you are
- Teams training or fine-tuning a model who have data but no labels
- Companies whose first annotation vendor returned something inconsistent
- Teams with a repetitive, document-heavy task and the volume to justify fixing it
Probably not, if you want
- Training a foundation model from scratch. Almost nobody needs this, and we are not the studio for it.
- Labelling that needs a licensed professional to be valid, such as clinical diagnosis from scans.
- Anything where a confident wrong answer causes real harm with no human deciding.
06Questions
Does our data get used to train someone else's model?
No. We work through APIs with training disabled, or self-hosted where the data cannot leave your environment. Which of the two we use is agreed with you at the start, in writing.
What accuracy can you commit to?
An agreement threshold agreed with you after the pilot, because the achievable number depends entirely on how ambiguous your task is. Anyone quoting a percentage before seeing your data is guessing.
What if AI turns out not to be the answer?
Then we say so, usually within the first week, and you pay for the week rather than the project. That has happened, and we would rather it kept happening than build something useless.
Who owns what you build?
You do. Data, code, prompts and evaluation sets, in your repository. The same rule as everything else we make.
Send us what you have.
A rough deck, a Loom, a paragraph. We will reply within a business day with questions, a fixed quote and a start date.
