Interpretability at systems scale
Making the internals of large models observable in production, not just in notebooks — tracing behavior back to mechanism.
- Feature attribution
- Activation tooling
- Eval harnesses
an independent research lab
An independent lab working at the seam between machine-learning systems and the people who have to rely on them.
Northlight exists to do a narrow thing well: build the infrastructure that makes advanced machine-learning systems legible. Not another model, but the instruments around it — the ways we measure, interpret, and hold such systems accountable.
We work in small cycles, publish what we learn, and ship the tooling as we go. The bet is simple: the field is long on capability and short on understanding, and the gap is an engineering problem as much as a scientific one.
Three threads, one aim: the instruments that make machine intelligence legible.
Making the internals of large models observable in production, not just in notebooks — tracing behavior back to mechanism.
Control surfaces for long-horizon agents: how to keep intent, constraints, and human oversight intact as autonomy grows.
The unglamorous plumbing — provenance, reproducibility, and measurement — that lets anyone verify a claim about a model.
We collaborate with researchers, fund small tools, and hire rarely but deliberately. If the work above is yours too, get in touch.