Completed, Yonsei University, June to September 2026 (co-author)
Finds which step of a failed multi-agent run went wrong from a small open model’s uncertainty about each step, with no training and no API calls.
When several AI agents work together and the job fails, someone has to work out which step actually broke it. Reading the whole transcript with a very large model is expensive, and training a dedicated model needs many labelled failures.
SURF uses a fixed, open-weight model as an observer. It scores each step by how likely it is to be fine, and keeps that score instead of rounding it to yes or no. A step rated 51% and one rated 91% both round to “yes” but are very different. Ranking the steps by score points to the shakiest one, with no training and no labelled failures.
- Tested on Who&When, a public set of failed multi-agent runs where humans marked the culprit step.
- Beat judging by a closed frontier model by up to 10.3 percentage points, with no API calls.
- The gain comes from keeping the continuous score: it beats the same observer’s yes/no answer in 51 of 56 settings.
- Holds across different ways of picking the step and different uncertainty measures. A single quick “is this step true?” probability was the strongest and cheapest signal.


