The gap between what frontier AI can do and what anyone understands about it is the most interesting place in technology. That's where we work.
fig. 01 — sampled trajectory
Research areas
How frontier models reason, generalize, and fail under real-world conditions — studied through small, careful experiments.
Honest, narrow evals for the capabilities that matter in practice, built to be reproducible rather than impressive.
How people actually work with AI systems day to day, and what good tools look like when the novelty wears off.
What it takes for an agent to stay coherent across hours of real work — memory, recovery, and knowing when to stop and ask.
How base models become useful ones — what fine-tuning adds, what it erases, and how to tell the difference.
The most valuable training data left isn't on the internet. We capture it — real people doing real work in the physical world, with consent and provenance built in.