Human Data
Human experience is what turns information into intelligence.
License expert-written datasets or work with us to create new human data for a capability your model needs to learn. Most core tasks, rubrics and reference answers are authored without LLM assistance.
Off-the-shelf data
Curated collections, ready to license.
Our catalogue provides licensed training and evaluation data for difficult model capabilities. Every release states its provenance, permitted use, composition, quality process and limits.
Clinical Triage Reasoning - UK
Doctor-authored pathways transformed into auditable graphs, decision turns, interactive trajectories and correction data for sequential UK clinical triage.
- Modalities
- Clinical graphs · dialogue · structured actions
- Designed for
- Action tuning, interactive reasoning, safety evaluation and RL
Bespoke human data
Need data that does not exist yet?
We work with domain experts to create new tasks, rubrics, demonstrations and evaluation sets without LLM drafting.
Request bespoke human dataData strategy
Define the model behaviour, evidence and evaluation plan before collection begins.
Expert recruitment
Find and qualify people whose judgement represents the work the model must learn.
Task production
Have qualified people turn real workflows into new tasks, demonstrations, comparisons and adversarial cases without LLM drafting.
Quality systems
Use clear rubrics, calibration, review and disagreement analysis to protect signal quality.
Why human data
Preference is not the same as judgement.
Ask enough people which response they prefer and you can learn what looks helpful. But a longer answer, cleaner formatting or a well-placed emoji can win a preference vote while the underlying judgement is still wrong. Edwin Chen has explained this failure mode in public leaderboards.
Critical work needs people who can recognise whether an answer is correct, unsafe or incomplete. We begin with the capability, then recruit people qualified to write the tasks, define the standard and judge the result.
Our default is to create the core tasks, rubrics and reference answers without LLM drafting. Models may help with mechanical conversion or quality-control suggestions. They do not write the exam and mark their own paper.
Dataset standard
Useful data needs a clear record of how it was made.
- Core tasks, rubrics and reference answers authored without LLM assistance
- Documented provenance, consent and permitted use
- Contributors selected for relevant expertise
- Human verification against explicit rubrics
- Train, validation and private evaluation splits
- Versioned collection and quality-control methods
- Known limitations recorded in a data card