Long-horizon reasoning
Planning across many dependent steps. Maintaining context, recovering from mistakes, and carrying work through to completion.
A research lab for long-horizon AI
We build interactive environments for agents that
reason, act, and adapt. Our focus is the work that
unfolds over many steps.
Many useful problems cannot be solved in a single response. They require an agent to make a plan, act on an environment, interpret feedback, and revise its approach—while keeping a larger objective in view.
We study the environments and evaluations that make this kind of capability possible to develop and measure. Our central question: how do we build agents that remain effective as tasks get longer, state gets richer, and decisions compound?
Planning across many dependent steps. Maintaining context, recovering from mistakes, and carrying work through to completion.
Worlds with persistent state and meaningful consequences, where agents must observe, act, and adapt as a task unfolds.
Measuring what an agent accomplishes. Testing behavior in the environment, beyond the plausibility of its final response.
A benchmark for structurally correct, visually convincing 3D assets generated by frontier models. We examine whether generated objects satisfy their prompts through executable checks and geometric verification.
Our research takes shape in engineering tasks, spatial evaluations, and generation trajectories. Each offers a different view of how agents reason, build, and improve.
Substantial engineering problems that demand sustained reasoning, implementation, testing, and refinement. Tasks span compilers, systems programming, scientific computing, machine learning, and visual reconstruction.
Evaluated through task-specific tests, reference behavior, held-out data, and correctness-gated performance.
Three.js Asset Generation (3JAG) evaluates whether generated objects execute, render correctly, and satisfy the spatial relationships in a brief.
Deterministic measurements connect each verdict to the generated artifact.
Read the benchmark researchThe process behind a working game: creator requests, iterative code changes, and playable versions from the first build to the final result.
A view into how agents respond to feedback and develop an artifact over time.
Instaplay turns ideas into playable worlds. Games brought code, visual reasoning, and dynamic environments together—and gave our research a starting point.
Research at the intersection of agents, environments, and evaluation.
Discuss research & data