GPT-6.1 Sol's First Hard Test: A Clean Merge, Too Little Scope Discipline

Martin Rau · · 6 Min. Lesezeit

After my first impressions of GPT-6.1 Sol, I wanted to see how the model handles a task where it’s easy to slip. I gave Sol 6.1 an awkward merge and put Claude Opus 5.5 next to it as the reviewer. Short version: the merge itself is well done. Most of the runtime, though, went into test failures that existed before and had nothing to do with the merge.

The task

Ticket T-2042 was meant to merge six-day-old work into the target branch: the optimistic UI implementation from T-1561. Optimistic UI means the interface shows a change immediately and only reconciles it with the server afterwards. It makes the app feel noticeably faster, but it touches a lot of places in the code.

In those six days the target had changed heavily, around 600 files. That left 12 conflicting files, places where both sides had changed the same code differently and someone has to decide which version holds.

Six days sounds like little. In my setup it’s a lot: several agents work in parallel in their own Git worktrees, separate working copies of the same repo, and merge back continuously. Six days are closer to six weeks in a conventional development environment. That’s what made it a good test: a merge this size needs an understanding of both sides, not just resolving conflict markers.

How I reviewed it

Sol 6.1 ran as the implementer, Opus 5.5 as the reviewer. I deliberately use a model from a different provider for reviews, because it finds different mistakes than the model that wrote the code. Opus 5.5 read the merge diff, classified each commit that came after it, and re-ran the test runs.

What went well

The quality of the merge convinced me. The conflicts are resolved cleanly, in a way that keeps both sides: the optimistic markers from the old branch and the new target trail from the current state, plus the AGC and Cloud areas that had grown in the meantime. That’s the hard part of a merge like this. The easy way out would be to take one side and silently lose the other.

The review was clear: typecheck green, around 1,900 targeted tests in the affected areas green. For a merge this size, roughly an hour of work would have been reasonable. Sol 6.1 delivered that part.

Where the time went

After the merge, the agent spent about two more hours on four additional commits:

  • Three flaky tests, tests that pass one run and fail the next without any change to the code.
  • One real bug in the PlanPanel, a component unrelated to the merge.

None of these failures came from the merge. They were there before. On top of that, there were probably several runs of the full test suite on the VM, and each of those runs costs time and compute.

Did the agent go off track? Partly. The fixes are correct, Opus 5.5 confirmed that. The PlanPanel bug was even a genuine find. But the task was “merge T-1561 into the target”, not “clean up the test suite”. Sol 6.1 fixed failures instead of reporting them.

Why that’s a problem even when the fixes are right

You could argue: two hours for four correct fixes, where’s the harm? The harm sits elsewhere.

The task gets blurry. A merge commit should contain exactly what the merge needs. Once unrelated fixes mix in, the branch gets harder to review and harder to roll back if something does go wrong.

Costs become unpredictable. I hand out tickets with an idea of what they’ll cost. When an agent needs a multiple of the planned time because it tidies up along the way, the math no longer works. For a model like Sol 6.1, whose main argument is the price per task, that weighs twice as much.

Parallel work collides. In my setup several agents work at the same time. If one fixes something on the side that another is working on, it creates exactly the conflicts I split tickets to avoid.

What I take away

Sol 6.1 means well. It sees a red test and wants it green, even when it isn’t part of its task. That’s a likeable trait, but an expensive one in an agent setup. I have two ways to handle it.

Prompt the scope harder. From now on I state explicitly in the task what’s in scope and what isn’t. For Sol runs I add a fixed block:

Scope: only the merge of T-1561 into the target.
Failures that existed before the merge (red or flaky tests,
bugs in untouched files): report, don't fix.
Tests: only the affected areas, the full suite at most once at the end.

Use the find anyway. That Sol 6.1 spotted the flaky tests and the PlanPanel bug is valuable. I want the find, not the fix in the wrong ticket. Reported failures become their own card and get handled properly there.

The pattern isn’t new. Guardrails for agents, fixed limits on what an agent may do, are a topic of their own; I cover them in more depth in the lexicon article on guardrails for agents. What’s new is how clearly Sol 6.1 needs that limit. It fits OpenAI’s own numbers: the release report shows more “unwanted persistence” for Sol 6.1 than for Astra, the tendency to keep going beyond the actual task. I noted that figure in my first impressions. Now I’ve seen it in practice.

Where I place Sol 6.1 now

Sol 6.1 also runs in Codex, OpenAI’s coding environment, and takes over the role that specialised coding models like GPT-5.3-Codex used to have. After this test, here’s how I see it:

  • For clearly bounded, difficult coding tasks Sol 6.1 is strong. The merge was at a level I’d otherwise expect from Opus, at a fraction of the cost.
  • For open-ended tasks without a hard boundary I don’t use it. Without a clear scope it keeps working further than I want.
  • As the reviewer Opus 5.5 stays my pick. Classifying the four extra commits as “correct, but not requested” was exactly the kind of judgement I need from a reviewer.

Conclusion

Thorough? Clearly. Disciplined? Not yet. Sol 6.1 resolved a difficult merge cleanly and found real bugs along the way. But it didn’t distinguish between what it was supposed to do and what it could do. A clear scope in the task fixes that. Without one, you pay for work you never ordered.

Discover more

Topic overview