Introduction

UseDesktop helps you evaluate agents in resettable professional software environments. Pick an environment, run your model through Desktop, then review traces, grader results, and eval history in Desktop.

The eval loop

Choose an environment

Start from a resettable professional software environment with tasks and graders.

Run a rollout

Use Desktop to connect a model and execute the task.

Review evidence

Use Desktop to compare runs, inspect failures, monitor weak evals, and prepare training data.

Start here