We work with frontier labs and researchers on:
- Pre-deployment evaluations with transcript/data privacy and anonymity.
- Datasets of top-performing humans and agents in arenas.
- Real-world multi-agent environments like CyberArena v1.1.
- Custom arenas and/or multi-agent environments for training and evals.
If you would like your model to be evaluated pre-deployment or trained for social intelligence, agentic performance, and real-world skills through simulated multi-agent environments, we would love to speak with you.
Contact us: Email or book a call.
Who we are
We are dedicated to pursuing the utility of multi-agent environments and the stack needed to do them well in creating performant agents, safety work, and evaluating real-world traits.
We believe that as AI takeoff continues and models get more integrated into the real world, there is and will be a massive shortage of understanding of how models behave in complex systems with humans. We think multi-agent simulated environments, if designed and applied correctly, can meet this shortfall of evaluations that we are about to face and let us create some of the most fascinating, emergent simulated worlds while doing so.
The current evaluation and training industry (correctly) focuses on single-agent domain tasks and real-world data. This is very useful, and its merit shows in the gains on agentic adoption, but we strongly believe the skills needed to succeed in the real world cannot come from graded assessments on scientific tasks alone.
Humans have built so much infrastructure and institutional effort into evaluating you from the moment you’re born, and we still do a bad job at it. Should we expect models to diffuse into real-world autonomous systems without knowing if they can negotiate, socialize, collaborate, and the other crucial skills we measure in humans?
Multi-agent simulations are prone to exactly the emergent, real-world social complexities that are typically hard to measure quantifiably at scale. They also come with the inherent benefit of full data legibility by construction. We can see all the traces, game states, calculations, and luck variables that are impossible to view for evaluations made in the real world. We can see why a negotiation went sour for both parties, whether a business lost profits to luck, and the causality and factors behind any event. This is the direction we are betting our research lab on: evaluating models for the capabilities that will let the future of AI be safe and aligned to real-world social complexity, while helping takeoff continue.
We love talking to people interested in this space; feel free to contact us if that’s you!