Logarithms Labs
Training frontier models for critical domains.
“The fox knows many things, but the hedgehog knows one big thing.”
Archilochus
Frontier models are remarkable generalists. They can write, code, search and reason across fields. But broad ability becomes uneven when the work turns specialised and the cost of a plausible mistake becomes real.
Critical domains demand more than a fluent answer. A model must understand the task, use the right tools, recognise uncertainty and produce work that experts can inspect and defend.
That capability is trained. It comes from realistic tasks, rewards that preserve professional judgement, environments where actions have consequences, and human data that records how experts decide.
For critical-domain work, qualified people write the core tasks, rubrics and reference answers without LLM drafting. The model should not create the standard it will later learn from or be measured against.
Logarithms Labs builds those training systems. We work at the uneven edge of model capability: the domains where models are almost useful, but not yet reliable enough to trust.
Our work begins where general capability becomes unreliable. We build the expert data, evaluations, environments and training systems that make frontier models useful in critical domains.
The training system
Teach the work, not the appearance of it.
A model learns inside a controlled world: it receives a task, acts through tools, changes state, receives expert-aligned reward and is tested against cases it has not seen.
Featured guide
The Art of Benchmarking
A practical, end-to-end guide to deciding what to measure, writing representative tasks, validating graders and publishing results people can trust.
Read the guideThe Art of
Benchmarking
How to build AI evaluations you can trustResearch